DittoFS uses a flexible configuration approach with named, reusable stores. This allows different shares to use completely different backends, or multiple shares can efficiently share the same store instances.
DittoFS has no OpenTelemetry / distributed-tracing subsystem. Observability is
provided by an opt-in Prometheus /metrics endpoint on a dedicated listener. See
Metrics (Prometheus) below for the full configuration.
DittoFS uses a control plane database to store persistent configuration for users, groups, shares, and permissions. This enables dynamic management via CLI commands and REST API without restarting the server.
database:
# Database type: sqlite (single-node) or postgres (HA-capable)
Bind interface. Loopback-only by default (secure-by-default for single-host). Set to 0.0.0.0 for multi-host / Kubernetes (then front it with TLS termination — see TLS and bind address)
port
8080
HTTP/HTTPS port for API endpoints
read_timeout
10s
Maximum duration to read request
write_timeout
10s
Maximum duration to write response
idle_timeout
60s
Maximum idle time for keep-alive
drain_stall_timeout
5m
Inactivity bound for POST /api/v1/system/drain-uploads. The drain has no total time cap — a multi-GiB flush runs as long as it keeps making progress — and is aborted (504) only if no upload completes within this window while bytes are still unsynced (the remote stalled). A drain with nothing left to upload is never aborted: no attempt can conclude, so a flat counter says nothing about liveness. Mirrors rclone’s --timeout
require_initial_password_change
true
Force the bootstrap admin user to change its password on first login. Set to false to opt out (automated/test deployments). Also skipped when DITTOFS_ADMIN_INITIAL_PASSWORD is set
pprof
false
Expose Go /debug/pprof/* profiling endpoints
pprof_mutex_rate
100 (when pprof: true; else 0)
Mutex contention sampling, 1 per N events. Applied only when pprof: true; unset/0 then falls back to 100. Without it /debug/pprof/mutex is header-only. Disable profiling via pprof: false, not by zeroing this
pprof_block_rate_ns
1000000 (when pprof: true; else 0)
Block profiling rate in ns, 1 sample per N ns blocked. Applied only when pprof: true; unset/0 then falls back to 1000000. Without it /debug/pprof/block is header-only. Disable profiling via pprof: false, not by zeroing this
JWT Configuration Options:
Option
Default
Description
secret
(required)
HMAC signing key (min 32 chars)
access_token_duration
15m
Access token lifetime
refresh_token_duration
168h
Refresh token lifetime (7 days)
Security Note: The JWT secret should be kept confidential. Use the DITTOFS_CONTROLPLANE_SECRET environment variable in production to avoid storing secrets in config files.
The control plane API carries admin logins, the dfsctl remote password login, operator credentials, and JWTs. By default the server binds to 127.0.0.1 (loopback only) so a fresh dfs start is not reachable off-host. For any deployment that must accept connections from another machine — multi-host, Kubernetes — set host: 0.0.0.0 and protect the listener with TLS.
DittoFS offers native, file-based TLS as a secure-by-default floor. It is intentionally thin: DittoFS loads certificate files (and an optional client CA for mTLS) — it is not a certificate authority, does not generate self-signed certificates, and does not do ACME, issuance, renewal, or rotation. Certificate lifecycle is left to your platform (cert-manager, a mounted Kubernetes Secret, Vault, your PKI). When the platform rewrites the files on disk, DittoFS hot-reloads them with no restart.
Path to the PEM server certificate (or chain). Both cert_file and key_file must be set to enable HTTPS; setting one without the other is a fatal config error
key_file
(unset)
Path to the PEM private key for cert_file
client_ca
(unset)
Path to a PEM CA bundle. When set, the server requires and verifies a client certificate signed by one of these CAs (mutual TLS). Requires cert_file/key_file
min_version
1.2
Minimum negotiated TLS version: "1.2" or "1.3"
When cert_file/key_file are unset, the server serves plain HTTP exactly as before (back-compatible). When set, it serves HTTPS; the files are read and parsed at startup, so a missing or malformed certificate fails fast with a clear error.
Recommended deployment model: terminate TLS for the edge at an ingress / service mesh / reverse proxy (NGINX), and use DittoFS native TLS (or mTLS via client_ca) as the secure floor for non-Kubernetes hosts and direct dfsctl access. See docs/SECURITY.md and docs/DEPLOYMENT.md. For Kubernetes, the operator renders host: 0.0.0.0 automatically so the API Service can reach the pod; see docs/DEPLOYMENT.md.
Every share has exactly one block store — the durable home for its content —
plus an on-disk journal that absorbs writes in front of it. Block stores are
created with dfsctl store block … and attached to a share with
dfsctl share create/edit --block-store <name>. Valid types are s3 and
memory; a block store has no kind, and the old local/remote split is gone.
The journal is not a store you create. It is provisioned automatically under
one server-level root, blockstore.journal.path, with each share getting its own
subdirectory (<path>/shares/<share-name>/journal/). The block store lives in
pkg/block/engine/ and composes the journal, the block store, the unified
CAS-keyed in-memory Cache, a syncer (async journal-to-block-store transfer), and
a garbage collector.
The journal (pkg/block/journal/) is an append-only, log-structured write-back
tier. Writes append to it and are acknowledged locally; a background flush pass
packs dirty ranges into packed blocks (blocks/<id>) and offloads them to the
block store. Pre-v0.16 {payloadID}/block-{idx} layouts must be converted with
dittofs ≤ v0.21 (its migrate-to-cas command) before the server will start.
Its knobs are server-level config, not per-share store JSON. They used to live
in a per-share local block store config blob; that store is gone and every key
below moved into the top-level blockstore.journal block:
backpressure_max_wait: 60s# how long a write stalls before ErrDiskFull
Key
Type
Default
Description
path
string
<state dir>/blocks
Root holding every share’s journal. Must be absolute (a leading ~ is expanded).
chunk_size
int
0 (built-in profile)
FastCDC minimum chunk size in bytes — see Chunk size.
chunk_max
int
0 (derived)
Overrides the derived maximum chunk size.
dirty_expire
duration
30s
Dirty-age fsync ceiling — see below. Negative disables it; values under 1 s are clamped with a warning.
max_log_bytes
int
deduced (25% of RAM, floor 1 GiB)
Append-log pressure budget. AppendWrite stalls once the buffered total exceeds it.
backpressure_max_wait
duration
60s
How long a write blocks waiting for the syncer to drain before returning ErrDiskFull.
Env-var mapping follows the dot-path convention, e.g.
DITTOFS_BLOCKSTORE_JOURNAL_MAX_LOG_BYTES.
A config file still carrying the old blockstore.local.* keys is refused at
startup, naming both spellings — see Upgrading.
Because these knobs are now server-level, a node mixing a VM-image share with a
general-purpose one gets a single chunking profile for both.
There is no metadata.RollupStore backend requirement any more — the journal
owns its own local state; a metadata backend only needs the block-record and
synced-hash contracts (see implementing stores).
The journal’s ceiling is a per-share setting, --journal-size on
dfsctl share create / share edit (journal_size in the share JSON). It
replaces the old --local-store-size flag and the deleted
blockstore.journal.default_remote_cache_size key.
Unset — the share has no configured ceiling. The journal still sizes a soft
default off the volume’s free space when it opens (80% of what is free at that
moment), and only degrades to genuinely unbounded growth if the free-space probe
fails, which it warns about. This is a real change from the old behaviour, which
deduced a ceiling from RAM (25%) or capped a remote-backed share at 10 GiB: a
fast writer against a slow uploader can now fill much more of the volume before
anything pushes back.
Set — the value is the ceiling eviction reclaims against.
Eviction can only reclaim blocks already offloaded to the block store.
Anything else is still the only copy of those bytes, so dropping it would lose
them. The consequence is easy to get wrong: a cap reached with nothing yet
offloaded cannot be honoured by eviction at all — the write path falls straight
through to backpressure, stalls for up to
blockstore.journal.backpressure_max_wait, and then returns ErrLocalStoreFull
(surfaced as disk-full to the protocol). That is the correct response, not dead
code: a healthy block store lets the flush pass drain dirty bytes, eviction frees
space, and the writer proceeds.
A stall logs its cause so the next move is unambiguous:
WARN journal local store full: nothing evictable, backpressuring writes
unsynced_bytes > 0 means the flush pass is behind and the wait will clear;
eviction_suspended means the block store is unhealthy (fix the outage);
eviction_pinned means a retention policy is holding the blocks (change the
policy).
Durability is a per-store property: whether bytes a store has accepted survive
a daemon crash / restart. Each store resolves an effective durable flag at
construction — a type default that an operator may override.
Store
Default durable
journal (every share)
true — bytes are on disk; un-offloaded chunks are never evicted, survive restart, and offload asynchronously
s3 block store
true — durable object storage
memory block store
false — test/dev fixture, lost on restart
Override the default per store by adding a durable bool to the store’s config
JSON — {"durable": false} for a journal on a volatile tmpfs mount, or
{"durable": true} to deliberately treat a memory store as durable in a test/dev
setup. A non-bool value is ignored with a startup warning (the type default
stands). The effective values are surfaced as Local Durable / Remote Durable
in dfsctl store block stats.
CLOSE/COMMIT semantics. SMB CLOSE and NFS COMMIT (and the NFSv3 stable-WRITE
path) call the engine flush. A hard flush error (I/O fault, block-store
rejection, metadata error) is always surfaced to the client. Beyond that, what
a COMMIT waits for is the per-share commit acknowledgement, commit_ack
(default journal):
commit_ack
CLOSE/COMMIT behavior
journal (default)
Acknowledge once the write is durable in the share’s journal. The offload to the block store stays fully asynchronous and observable via the unsynced-bytes metric / Pending Remote (bytes). Ordinary NFS/POSIX writes never EIO.
block-store
Acknowledge only when the data is on a durable store: committed := localDurable || (Finalized && remoteDurable). Every commit waits for an upload.
dfsctl share show <name> prints the setting in force as a Commit Ack row.
See Durability for what each one survives and the throughput
cost — and note that the metadata-commit relaxation is a separate, independent
axis, not a third value of this one.
Under commit_ack: block-store the COMMIT drives the offload inline — it flushes
the file’s dirty ranges into packed blocks, uploads them, and only then returns.
What happens when that cannot finish depends on whether the journal is durable:
Durable journal (the normal case): an unhealthy block store makes the flush
return its soft, non-finalized result, and the COMMIT is acked anyway — the
bytes already survive a restart, and the syncer keeps retrying.
Volatile local tier (a journal explicitly marked {"durable": false}, or
the in-memory store used by tests): nothing crash-safe holds the bytes, so the
COMMIT returns a transient I/O error (NFS3ERR_IO / NFS4ERR_IO / SMB
STATUS_UNEXPECTED_IO_ERROR) and the client re-drives.
(NFS unstable WRITE is unaffected either way: it still returns UNSTABLE and
defers durability to a later COMMIT.)
dfsctl store block stats also shows Pending Remote (bytes) — the headline
data-at-risk gauge (journal bytes not yet offloaded) — which is the way to observe
the async offload backlog under the default policy.
DittoFS splits file data into content-defined (FastCDC) chunks. A chunk is the
unit of dedup, of the local cache, and — critically — of a read fetch: a
small random read pulls the whole chunk covering its offset. The default chunk
floor is ~1 MiB, so a 4 KiB random read into a large file amplifies to ~1 MiB
(~256×). That is fine for sequential and archival workloads but poor for random
I/O (VM images, databases).
chunk_size (bytes) lowers the FastCDC minimum for a share, shrinking the read
unit. Effective average chunk size ≈ chunk_size; a hard ceiling is derived
(8× chunk_size) unless you set chunk_max explicitly.
Random-access node: ~128 KiB chunks (≈8× less read amplification) — set
blockstore.journal.chunk_size: 131072 in the server config. This is a
server-level knob, so it applies to every share on the node.
Setting
Effective avg
4 KiB random-read amplification
Trade-off
default (unset)
~1 MiB
~256×
best dedup, fewest manifest rows
65536 (64 KiB)
~94 KiB
~16×
~16× more FileChunk rows; ~0 extra dedup loss on VM/DB
131072 (128 KiB)
~159 KiB
~32×
~8× more manifest rows
262144 (256 KiB)
~287 KiB
~64×
~4× more manifest rows
Notes:
S3 object count is unchanged. Chunks are packed into ~16 MiB block objects
regardless of chunk_size, so smaller chunks do not create more/smaller
S3 objects — only more FileChunk manifest rows. Writes/uploads keep their
full throughput.
Write-time only. Reads never re-chunk — the manifest records each chunk’s
boundaries — so changing chunk_size affects only newly written data, and old
data stays readable. Dedup is not restored across a change (different
boundaries → different hashes), but on VM/DB images dedup is already ~0.
An invalid chunk_size / chunk_max combination is warned about and dropped,
leaving the built-in profile in force — a bad profile would otherwise cut every
newly written block to the wrong size, and reads never re-chunk, so the damage
would outlive the misconfiguration.
A client that writes and never issues an NFS COMMIT/FILE_SYNC or an SMB
FLUSH/CLOSE never asks the server for durability. A background loop fsyncs
each journal shard still holding uncommitted records once per interval, so those
writes reach the device within roughly that window instead of waiting for the
shard’s next 256 MiB segment rotation.
The default is 30s; tighten it with blockstore.journal.dirty_expire: 5s in the
server config, or set a negative value to disable the loop. It used to be a
per-share dirty_expire_seconds key in the local block store’s config JSON.
Default 30 s, on for every share; the loop costs an idle store nothing and
never runs on the ack path.
Negative disables it, leaving the client’s own fsync and segment rotation as
the only durability points — the loss window is then unbounded in time.
Values below 1 s are clamped with a warning.
This is a ceiling on the loss window, not a durability guarantee: only a
returned COMMIT/FLUSH says the bytes are on the device. See
Durability.
The CAS write path uses an async syncer and a fail-closed mark-sweep
garbage collector. The syncer’s sizing (claim timeout, etc.) is not an
operator-facing config section — it is auto-deduced from system resources at
startup and constructed in code; there is no syncer: config block (a stale
syncer: section in a config file is tolerated but ignored, logged as an
unknown key). Upload concurrency is adaptive by default — see below.
When you mirror a share to its block store (S3 or memory), DittoFS uploads
CAS chunks concurrently. Uploads are network-latency bound, not CPU bound:
a single PUT to a remote region sustains only a few MiB/s, so throughput scales
with the number of concurrent uploads until the uplink saturates. The right
number depends on the link, not the host — a CPU-derived default throttled fast
links and over-subscribed slow ones.
By default DittoFS discovers the right concurrency itself. It starts
conservative and ramps the number of in-flight uploads up while delivered
throughput (goodput) keeps rising, settling at the point where opening more
connections stops helping. It backs off only on upload errors or a real
throughput collapse — never on the latency rise that healthy concurrency itself
causes. No configuration is required to saturate the uplink.
To pin a fixed concurrency instead (disabling auto-tuning), set
--parallel-uploads N on the remote:
Terminal window
dfsctlstoreblockadd--namer1--types3\
--bucket…--region…--endpoint…\
--parallel-uploads32# fixed window of 32; 0 (default) = adaptive
dfsctl store block edit r1 --parallel-uploads 0 returns a remote
to adaptive mode. Observe the live window via the Prometheus gauge
dittofs_datapath_upload_window (target concurrency) alongside
dittofs_datapath_uploads_inflight (actual in-flight uploads); see
Metrics.
The mark-sweep GC is the one tunable surface, configured via the top-level
gc: server-config section:
gc:
grace_period: 1h# Objects whose LastModified is newer than
# (snapshot - grace_period) are NEVER
# deleted. Default 1h. Values in (0, 5m)
# are REJECTED at config load; values in
# [5m, 10m) are accepted but emit a
# warning. The cushion protects in-flight
# uploads whose metadata-txn lands after
# the snapshot.
dry_run_sample_size: 1000# Maximum candidate keys reported in
# --dry-run mode. Default 1000.
compaction_live_ratio: 0# Reclaim dead bytes from partially-dead
# blocks. After each sweep, a block whose
# live bytes / object size is below this
# ratio is repacked (live chunks only) and
# the old block deleted. Must be in [0, 1];
# 0 (default) disables compaction. A value
# like 0.5 compacts a block once it is more
# than half dead.
auto_enabled: true# Run background GC automatically so you
# don't have to invoke the CLI. Default
# true. Set false to require manual
# `dfsctl store block gc`.
auto_interval: 15m# Period between background GC runs.
# Default 15m. Values in (0, 1m) are
# REJECTED. Ignored when auto_enabled is
# false.
Tuning guidance:
Background GC is on by default (auto_enabled: true, every
auto_interval) and reclaims orphaned blocks on both the local
and remote tiers. Disable it (auto_enabled: false) only if you want
to drive GC entirely on demand or via external scheduling.
You can still run GC on demand at any time:
dfsctl store block gc <share> (add --dry-run to preview, capped by
gc.dry_run_sample_size; add --reconcile to also reap rows leaked by
older versions).
gc.grace_period MUST be longer than your worst-case
metadata-commit latency after a successful PUT. The default 1h is
comfortable for any commit path that completes in seconds.
A share’s manifest can end up disagreeing with its own files’ block lists:
a file claims bytes in a range that no manifest row covers. At read time
that is indistinguishable from a legitimate sparse hole — absent rows are
how sparse files are represented — so the read returns zeros and reports
success. Nothing on the data path can report it. The background scan is
what finds it.
integrity:
auto_enabled: true# Run the structural manifest scan
# automatically. Default true. Set false to
# require manual `dfsctl store check`.
auto_interval: 24h# Period between scans. Default 24h. Values
# in (0, 5m) are REJECTED — a scan is a full
# metadata walk of every share. Ignored when
# auto_enabled is false.
The scan is metadata-only: it fetches no block and touches no remote
object, so it costs a metadata walk regardless of how much data a share
holds and generates no S3 egress. It writes nothing — it neither plans
nor applies repairs. Shares are scanned one at a time.
Findings surface in three places:
dfsctl share show <name> reports when the share was last scanned and
what was found. A share with damaged payloads reads as degraded even
when every subsystem probe is healthy.
Prometheus: dittofs_integrity_last_scan_timestamp_seconds,
dittofs_integrity_last_scan_failed,
dittofs_integrity_last_scan_duration_seconds,
dittofs_integrity_files_scanned, and
dittofs_integrity_findings{kind="..."} broken out by
payloads_with_findings, damaged_payloads,
claimed_uncovered_ranges, unplaceable_rows and unknown_hash_rows.
The last-scan timestamp is 0 both for a share never scanned and for one
whose last scan failed — a failed scan completes nothing, so it has no
time to report. dittofs_integrity_last_scan_failed is what separates
them: 0 with a zero timestamp means the scan has not run yet, 1 means
it is running and erroring. Without that second series a scanner failing
on every tick would look exactly like one that was never switched on, and
stay that way indefinitely.
The server log, at WARN, for any share with damaged payloads.
Run dfsctl store check <share> for the per-file detail behind a finding,
and --repair to act on it.
A share’s journal is a cache in front of its block store: once a range is
offloaded it becomes evictable, and an evicted range is served by fetching it
back. Reads of such a range fail while the block store is unreachable. Whether a
box would keep serving through an outage therefore depends on how much of its
data is currently block-store-only, and that number moves with every eviction and
every warm.
The server reports it per share:
dfsctl share show <name> prints Offline Safe: yes, or
no (12.4 GiB remote-only across 431 ranges).
Prometheus: dittofs_offline_safe{share} (1/0),
dittofs_offline_remote_only_bytes{share} and
dittofs_offline_remote_only_ranges{share}.
Zero remote-only bytes is a provably offline-safe share. To get there, warm
the share and stop it evicting again:
Terminal window
dfsctlsharewarm/archive
dfsctlshareedit/archive--retentionpin
The measurement never guesses. Three cases report unknown rather than a
number, because a zero would read as “provably safe” for exactly the shares
whose data is most likely to be remote-only:
the share’s journal does not track residency (the in-memory backend),
the journal has not been seeded from the manifest yet — it holds no
record of ranges that live only in the block store, so they would count as
absent rather than remote-only,
the block store is closed,
the residency scan did not finish inside the request’s deadline.
An unknown share reports dittofs_offline_safe = 0 but publishes no byte
counts, so a dashboard cannot mistake it for a clean fully-local share.
A share whose block store was swapped out after it had already evicted is the
awkward case: the evicted ranges stay recorded in the journal and are replayed
from its cold log on the next open, but there is no longer anything to fetch them
from, so they never serve. Those shares report a non-zero remote-only figure
rather than being waved through as safe.
The figure is bytes, not blocks. The journal tracks byte ranges, which
split and merge independently of manifest chunk rows; a block count would
need a metadata walk to produce and would not answer “how much would break
offline” any more precisely.
Note this is read availability only. Offline writes already work — writes
land in the journal and drain when the block store returns.
The schedule restarts from zero on server start, so a box restarted more
often than auto_interval never completes a scan. Lower the interval on a
host that reboots frequently.
The on-disk journal is a temporary write-through tier, not the durable copy —
every chunk is offloaded to the block store and may be evicted locally once
offloaded. To stop a fast writer with a slow/lagging uploader from filling the
host volume, the journal is bounded and writes apply graceful, observable
backpressure when it fills:
Bounded journal. The ceiling is the per-share --journal-size. With none
set, the journal claims a soft default of 80% of the volume’s free space at open
time — see Journal size and eviction for the full
rule and for why eviction can only reclaim already-offloaded blocks. The cap is
enforced lazily, on the write/flush path (it evicts offloaded segments to
make room for new writes); it
is not a background reaper, so on an idle or read-only workload the
resident local tier is not shrunk toward the cap. To reclaim local disk — or
to force cold, remote-served reads for read-path benchmarking — evict on
demand with dfsctl store block evict (drops the read buffer and drains
resident offloaded blocks; never drops not-yet-uploaded data). Read-miss
volume (the read-amplification signal) is observable via the
dittofs_datapath_block_range_read_bytes_total metric.
Backpressure stall. When the journal is full and every resident chunk is
still unsynced, a write stalls waiting for the syncer to drain to the
block store and free space, rather than failing. The stall is bounded by
blockstore.journal.backpressure_max_wait (default 60s).
Hard failure only when the block store cannot drain. If it is
unhealthy (genuinely unreachable, not merely slow) or the backpressure
window is exceeded, the write fails with disk-full
(NFS3ERR_NOSPC / NFS4ERR_NOSPC / SMB STATUS_DISK_FULL) instead of
silently filling the disk.
Diagnosing a stall. A stall warns once with the state that explains it, so a
stalled writer is never a mystery — see
Journal size and eviction for how to read the
fields:
WARN journal local store full: nothing evictable, backpressuring writes
default_remote_cache_size was deleted along with its “10 GiB when a remote is
configured, deduced size otherwise” conditional; the per-share --journal-size
is now the whole of the sizing policy.
Prometheus metrics for cache pressure / unsynced bytes are tracked
separately (server-wide instrumentation, issue #1188); today the signal is
the structured logs above.
The recycle bin is configured per share via dfsctl share create /
dfsctl share edit (or the REST share create/update body), not the
server config file. When enabled, deleting a file or directory moves it
into a visible #recycle directory at the share root instead of
destroying it; it can be restored over the mount or with dfsctl trash.
Setting (dfsctl flag)
REST field
Type
Default
Meaning
--enable-trash
trash_enabled
bool
false
Turn the per-share recycle bin on or off. Disabling it auto-empties the bin (permanently deletes its contents).
--trash-retention-days
trash_retention_days
int
0
Auto-purge bin entries older than N days. 0 = keep forever.
--trash-restrict-empty-to-admin
trash_restrict_to_admin
bool
false
Restrict emptying the bin to admins. Users may still restore.
--trash-max-size
trash_max_bytes
int64 (bytes)
0
Cap total bytes held in the bin; over-cap evicts oldest-first. 0 = unbounded.
--trash-exclude
trash_exclude_patterns
glob (repeatable)
(none)
Deletions matching any glob bypass the bin and are removed immediately.
A background reaper enforces trash_retention_days and
trash_max_bytes on an hourly interval. Deletes of items already
inside#recycle are permanent, and in-place truncate/overwrite of a
file’s content is not recycled (only unlink and replace-overwrite are).
dfsctl share show <name> displays the active trash configuration.
Terminal window
# Enable the bin with a 30-day retention and a 10 GiB cap
A remote block store may compress block payloads before upload and
decompress on download. The plaintext BLAKE3 hash remains the CAS key,
so dedup and GC are unaffected. The decorator is per-remote: every
share that references the remote inherits its compression policy.
Add a compression block to the remote store’s config JSON when
creating it:
Algorithm: "zstd" or "lz4". Defaults to zstd when the compression block is present but algo is omitted.
Notes:
Absence of the compression block means no wrapping — zero behavior
change for existing remotes.
Per-block adaptive: if the compressed body is not strictly smaller
than the plaintext, the decorator stores the raw plaintext with no
header. Incompressible payloads (random data, already-compressed
media) cost only the encoder pass, no on-wire expansion.
GetRange on a framed block decompresses the full block before
slicing — there is no random access into a compressed body. Read
paths that consume whole CDC chunks are unaffected.
The policy is captured at remote-store creation; restart the share
after editing the config to switch algorithms. Mixed framed and raw
blocks coexist within one remote and the reader auto-detects via the
5-byte DFCMP magic prefix.
A remote block store may also encrypt block payloads before upload using
client-side envelope encryption. Compression (when enabled) runs
before encryption — encrypted bytes are incompressible by design.
See ENCRYPTION.md for the full threat model and design.
Add an encryption block to the remote store’s config JSON:
The passphrase that unlocks a local key file is read from the
DITTOFS_ENCRYPTION_PASSPHRASE environment variable — never the config
file or command line.
The s3 remote store talks the AWS S3 API, so any S3-compatible object
store works — set a custom endpoint (and credentials) and DittoFS connects
to it instead of AWS. The store reads exactly these config keys (see
pkg/block/remote/s3/store.go and the factory in
pkg/controlplane/runtime/shares/service.go):
Key
Required
Default
Notes
bucket
yes
—
Bucket name. Must already exist; DittoFS does not create it.
access_key_id
yes
—
S3 access key. For GCS use an HMAC key, not a service-account JSON.
secret_access_key
yes
—
S3 secret key.
region
no
us-east-1
Some providers ignore it but the SDK still requires a value; the default is sent when omitted.
endpoint
no (AWS) / yes (others)
AWS
Service URL. Scheme optional — https:// is prepended when absent.
force_path_style
no
auto
Auto-enabled whenever endpoint is set. Set explicitly to false to opt back into virtual-hosted-style for providers that require it (e.g. GCS).
prefix
no
—
Key prefix prepended to every block (e.g. dittofs/). End it with /.
allow_private_endpoint
no
false
Required to point endpoint at a loopback or private-network address (MinIO, LocalStack, self-hosted RGW). See the SSRF note below.
Path-style addressing (endpoint.example.com/bucket/key) is the safe
default for non-AWS providers because virtual-hosted style
(bucket.endpoint.example.com/key) needs wildcard DNS and TLS SANs that
most S3-compatible gateways do not provide. DittoFS therefore flips
force_path_style on automatically the moment you set a custom endpoint;
the only providers below that need it turned back off are those that
require virtual-hosted style (GCS).
Private endpoints (MinIO, LocalStack, self-hosted RGW). DittoFS rejects
an endpoint that resolves to a loopback or private-network address
(127.0.0.0/8, 10/8, 172.16/12, 192.168/16, link-local, ULA) as an
SSRF guard; the cloud metadata address (169.254.169.254) is always blocked.
Self-hosted gateways normally live on exactly those networks, so add
"allow_private_endpoint": true to their --config to permit them, as the
MinIO and Ceph recipes below do.
Credentials live in the store’s own config (the --config blob below, or the
equivalent --access-key / --secret-key flags) — they are not read from the
DITTOFS_* server-config environment. Each recipe is a
dfsctl store block add invocation; attach the resulting store to a
share with dfsctl share create … --block-store <name>.
These are configured exactly like the verified ones but have not been run
against a live account in CI — they are documented from each provider’s S3
compatibility guide. The per-provider column flags the one gotcha that bites.
Provider
Endpoint
Region
force_path_style
Gotcha
Cubbit DS3 ⭐ (DittoFS sponsor)
https://s3.cubbit.eu
eu-west-1
auto-on
Geo-distributed, S3-compatible object storage from Cubbit. Create an S3 access key/secret in the DS3 console; the bucket lives in your assigned region.
Google Cloud Storage (XML/HMAC)
https://storage.googleapis.com
us-east-1
set false
Use an HMAC key (access_key_id/secret_access_key), not a service-account JSON. GCS ignores region (any non-empty value works), so send the us-east-1 default. GCS wants virtual-hosted style, so override the auto path-style default to false.
Backblaze B2
https://s3.us-west-004.backblazeb2.com
us-west-004
auto-on
Endpoint embeds the region (s3.<region>.backblazeb2.com); region must match it. Use an application key, not the master key.
Wasabi
https://s3.us-east-1.wasabisys.com
us-east-1
auto-on
Region is in the hostname; mismatched region causes auth failures.
DigitalOcean Spaces
https://nyc3.digitaloceanspaces.com
us-east-1
auto-on
Endpoint is the datacenter (<region>.digitaloceanspaces.com); send region: us-east-1 (Spaces ignores it but the SDK requires a value).
Alibaba Cloud OSS
https://oss-us-west-1.aliyuncs.com
us-west-1
auto-on
Region is encoded in the endpoint host; use the matching OSS region.
Tencent Cloud COS
https://cos.ap-guangzhou.myqcloud.com
ap-guangzhou
auto-on
Bucket name must include the AppID suffix (name-1250000000); endpoint carries the region.
All of the above accept the same optional knobs as AWS S3 —
prefix, compression, encryption, and durable (see the preceding
subsections) — because they share the single s3 store implementation.
Before this release a share carried two block stores — a local one (type fs
or memory) and an optional remote one — and both the config file and the
control-plane schema were shaped around that split. The first start after the
upgrade migrates them, and refuses to start in six cases where guessing would
silently change what an operator configured. Each refusal is recoverable; none of
them lose data.
1. A config file still using blockstore.local.*
config uses renamed keys; update them and restart:
Unknown keys are normally warned about and ignored, but a renamed one cannot
ride that policy: the operator believes the setting is live, an equivalent exists
under the new name, and dropping it silently changes runtime behaviour.
Action: rename the section to blockstore.journal (every key beneath it keeps
its name) and restart.
2. Shares whose journals do not live under the configured root
shares do not live under the configured journal root
share "/archive" stores its data in: /srv/dittofs/archive
Set blockstore.journal.path to /srv/dittofs/archive and restart.
A share’s journal holds the only copy of every byte not yet offloaded, so a root
that disagrees with where the data actually is would serve a share whose bytes are
somewhere else — reads would return zeros for everything written before the
change, with nothing failing. Relocating the directories automatically would mean
moving that only copy during startup, where a partial move cannot be undone.
Action: when every share agrees on one directory, the error names it — set
blockstore.journal.path to that directory and restart. When shares are spread
across several directories, one root cannot express that: move them under a single
parent (preserving the shares/<name>/ level), point blockstore.journal.path at
it, and restart.
3. Both size columns present
shares table has both local_store_size and journal_size;
copy the intended values into journal_size, drop local_store_size, and restart
An earlier upgrade added the new column without moving the values across. Which
one is authoritative is not recoverable from the schema, and choosing wrong
silently changes every share’s size ceiling.
Action: copy the values you intend to keep into journal_size, drop
local_store_size, and restart.
4. Both store-id columns present
shares table has both remote_block_store_id and block_store_id;
copy the intended values into block_store_id, drop the old column, and restart
Same shape as above, with a worse failure mode: picking the empty column would
leave every share without a block store.
Action: copy the values into block_store_id, drop
remote_block_store_id, and restart.
5. A share left with no block store
share has no block store after upgrade
"/archive" has no block store
These shares kept their data in a local block store and never had a
remote one. Every share now needs a block store, and the local store
they used cannot become one — pick an s3 or memory block store for
each, with the server stopped:
UPDATE shares SET block_store_id = '<block store id>' WHERE name = '<share>';
A share that kept its data in a local store and never had a remote one reaches
the new model with nothing to bind to: the rename in case 4 hands it an empty
block_store_id. Carrying the old local id across is not a repair — that tier
was typically an fs store, a type that exists only locally, so the share would
come up bound to a block store that cannot be built. The choice of a real store
is the operator’s.
Action: pick an s3 or memory block store for each named share and set
block_store_id to its id, with the server stopped, then restart. The old
column is left in place until every share is bound, so the record of which local
store each one used is still there to consult.
6. Two block stores sharing a name
two block stores share a name
"archive" is used by more than one block store
Block stores no longer have a kind, so their names must be unique.
Rename one of each pair with `dfsctl store block edit`, then restart.
Names used to be unique per (name, kind), so a local archive and a remote
archive could coexist. With the kind column gone they collide. Renaming one
automatically would break the operator’s scripts at some later point; deleting one
would orphan any share referencing it.
Action: give one store of each colliding pair a different name, then restart.
dfsctl store block edit <name> --name <new name> renames a store, but it talks
to a running server and this collision stops the server from starting — so break
the tie in the control-plane database first, with the server stopped:
UPDATE block_store_configs SETname='archive-local'WHERE id ='<the row to rename>';
Pick the row by id; SELECT id, name, kind, type FROM block_store_configs shows
which is which while the kind column still exists. A share normally holds the
store’s UUID, but older rows written through the REST update path hold the name
instead, so check
SELECT name, block_store_id FROM shares afterwards and repoint any share that
was matching on the old name.
shares.local_block_store_id → dropped, after its durability / writeback /
require_durable_commit settings are carried onto the share as commit_ack and
relaxed_metadata_commit (see Durability)
block_store_configs.kind → dropped, after the collision check above
blockstore.journal.default_remote_cache_size → deleted; per-share
--journal-size is now the whole of the sizing policy
Metadata store instances are created through dfsctl; the server config file
contains the global BadgerDB cache settings below. Filesystem capabilities and
limits are supplied by the metadata store and protocol implementation, not by a
config-file capabilities block.
The BadgerDB metadata engine keeps two in-memory caches that dominate read
performance under concurrent NFS/SMB load:
the block cache — decompressed LSM-tree data blocks, and
the index cache — the block-offset indices used to locate keys.
Badger’s own defaults are tiny (256 MiB block, index cache disabled), which
thrashes on a busy server over a large directory tree. The symptom in the logs
is Block cache might be too small ... hit-ratio: 0.26 ... sets-rejected; every
cold lookup then walks the LSM tree from disk, which also widens the window for
the dedup transaction-conflict race and the local-cache write-path backpressure
stall.
By default both sizes auto-scale with the memory available to the process,
so no tuning is required:
Cache
Fraction of available RAM
Floor
Ceiling
block
15 %
512 MiB
4 GiB
index
7.5 %
256 MiB
2 GiB
The fractions are deliberately conservative because the same process also holds
the local journal cache, the metadata working set, and read buffers. The available-memory
figure is the cgroup limit inside a container, or physical RAM otherwise (same
detection used for block-store sizing). Examples:
4 GiB host → ~614 MiB block / ~307 MiB index (the floors don’t bind).
To override the auto-sizing, set the sizes explicitly (in MiB) in the top-level
metadata.badger block. Setting one dimension still auto-sizes the other:
metadata:
badger:
block_cache_mb: 2048# 0 (default) = auto-size from available RAM
index_cache_mb: 1024# 0 (default) = auto-size from available RAM
Recommended sizing vs. object/metadata count. As a rule of thumb the block
cache should hold the hot directory/inode working set. Each cached file/inode is
on the order of a few hundred bytes of decompressed LSM data, so:
Hot metadata objects
Suggested block_cache_mb
Suggested index_cache_mb
up to ~1 M
auto (≥512)
auto (≥256)
~1–10 M
1024–2048
512–1024
~10–50 M
2048–4096
1024–2048
> 50 M
4096+ (raise host RAM)
2048+
If you still see hit-ratio below ~0.8 or sets-rejected in the Badger logs,
the cache is undersized for the working set — raise block_cache_mb first.
These global sizes apply to every BadgerDB metadata store on the node. A single
store can be overridden via its config-map keys when it is created (see below):
--config '{"path":"...","block_cache_mb":2048,"index_cache_mb":1024}'.
Applies to the badger and postgres metadata stores. Defaults to true.
Namespace operations (create, unlink, rename, mkdir, rmdir,
attribute-only setattr) commit without an inline flush. On badger a
background syncer makes them durable within ~100 ms; on postgres the
transaction runs with synchronous_commit = off, so the window is whatever the
server’s wal_writer_delay allows (PostgreSQL default 200 ms). Writes paired
with file data commit synchronously either way, so this is bounded loss, never
corruption.
Terminal window
# Strict: fsync every namespace commit (roughly a third of the create throughput)
Only an event that takes the kernel down — power loss, kernel panic, hypervisor
reset — can lose that window. Killing the dfs process (SIGKILL, OOM-kill,
panic) loses nothing at either setting, because an acknowledged write is already
in the kernel page cache. See
Durability → Namespace durability.
SQLite: Persistent embedded database - single-node deployments, pure-Go with no cgo. Shares its implementation with PostgreSQL, so the two behave alike apart from concurrency.
PostgreSQL: Persistent distributed database - multi-node deployments with horizontal scaling. Survives restarts and supports multiple DittoFS instances sharing the same metadata.
DittoFS supports two complementary quota layers, both enforced by NFS and SMB:
Per-share quota (dfsctl share create/edit --quota-bytes) — a single byte
ceiling for the whole share. Exceeding it returns NFS3ERR_NOSPC /
STATUS_DISK_FULL.
Per-identity quotas (dfsctl quota …) — per-user (uid) and per-group
(gid) limits, plus an optional default-user fallback applied to any user
without an explicit quota. Each quota bounds both bytes and inodes
(file count) and supports a soft threshold with a grace period before
the soft threshold is enforced as a hard limit. Usage is charged to the file
owner (standard quota semantics). Exceeding a hard limit (or an expired
soft+grace) returns NFS3ERR_DQUOT / NFS4ERR_DQUOT /
STATUS_QUOTA_EXCEEDED. df / FSSTAT and SMB FS_FULL_SIZE_INFORMATION
report the smallest applicable quota for the calling identity.
Terminal window
# Per-user quota: uid 1000 limited to 10 GiB / 100k files, soft at 8 GiB,
# 7-day grace (604800s) before the soft byte threshold becomes hard.
Per-identity quota usage is tracked incrementally by every metadata backend
(memory / badger / postgres), keyed by owner uid and gid, and is reconstructed
from the file rows on startup. A chown that changes a file’s owner moves its
bytes and inode count between identities. Limits live in the control-plane DB
and are also manageable via the REST API
(/api/v1/shares/{name}/quotas[/{scope}/{id}]).
The soft → grace → hard state machine records when an identity first crosses
its soft threshold and enforces the soft limit as hard once the grace window
elapses. For an explicit user/group quota the grace timer lives on the quota
row. For the default-user fallback the timer is inherently per-user (each
user trips soft at a different time, and the single shared template row cannot
hold per-user state), so it is recorded in a small side table keyed by
(share, uid), written the first time a default-user breaches soft and reaped
when usage drops back under soft. This makes default-user grace durable across
a server restart — a restart no longer hands every over-soft default user a
fresh grace window.
Note: enforcement is best-effort (matching the per-share soft quota):
under high write concurrency a few operations may briefly exceed a limit until
usage catches up. This is standard for userspace NFS/SMB servers.
Kubernetes operator: the operator manages infrastructure only — there is
no Share/Quota CRD. Quotas are managed via the REST API / dfsctl as above.
DittoFS supports a unified user management system for both NFS and SMB protocols. Users, groups, and their permissions are stored in the control plane database (see Database Configuration) and can be managed via:
REST API - For programmatic management and integrations
Permission resolution follows a priority order: user explicit permissions > group permissions (highest wins) > share default.
The server config bootstraps only the administrator through admin settings;
it does not import top-level user or group lists. Create other accounts and
memberships through the CLI or REST API.
NFS Authentication: NFS clients authenticate via AUTH_UNIX. The client’s UID is matched against DittoFS user UIDs. If a match is found, the user’s permissions are applied.
SMB Authentication: SMB clients authenticate via NTLM. The username is matched against DittoFS users, and permissions are applied from the user’s configuration.
There is no top-level guest bootstrap section. Configure anonymous NFS identity
mapping per export, for example with
dfsctl share nfs-config set /public --squash root_to_guest. SMB guest sessions
are governed by adapter authentication policy and share permissions. See
NFS access controls and SMB authentication.
Users and groups live in the control-plane database, not the config file. Manage them with
dfsctl against a running server (run dfsctl login first). See CLI.md for the
complete, generated reference.
User Commands:
Terminal window
# Create a user (password prompted interactively)
dfsctlusercreate--usernamealice
dfsctlusercreate--usernamealice--host-uid# map to your current host UID
Protocol adapters and their settings live in the control-plane database. Log in
with dfsctl login, then use adapter enable, adapter disable, and
adapter edit for lifecycle and ports; use adapter settings for protocol tuning.
These settings do not belong in the server YAML/TOML file and are not read from
DITTOFS_ADAPTERS_* environment variables.
Terminal window
# Set listen ports (the default NFS/SMB ports are 12049/12445)
The old names below were nested under the removed SMB adapter config section.
Rows without a supported flag are explicitly marked. For keys without an
exposed override, adapter edit smb --config is not a workaround.
dfsctl adapter settings smb update --signing required; modes are disabled, enabled, and required. SMB 3.1.1 still requires signing.
signing.preferred_algorithms
No exposed override; algorithms are negotiated from built-in support.
encryption.encryption_mode
No equivalent three-mode flag. dfsctl adapter settings smb update --enable-encryption can enable encryption support, but does not select a global required policy; see encryption below.
encryption.allowed_ciphers
No exposed override; ciphers are negotiated from built-in support.
leases.enabled
No exposed global lease toggle.
leases.directory_leases
REST-only directory_leasing_enabled in PATCH /api/v1/adapters/smb/settings; no CLI flag.
leases.lease_break_timeout
No SMB CLI override: lease breaks use a fixed 5-second bound. --oplock-break-timeout controls traditional oplocks only.
The shipped adapter uses preferred encryption mode. The
--enable-encryption flag enables support; setting it to false does not switch
the running adapter to disabled. There is no CLI flag for selecting a global
required mode.
Require encrypted SMB traffic for a particular share with --encrypt-data:
Terminal window
dfsctlsharecreate--name/secure--metadatadefault\
--block-stores3-remote--encrypt-data
See the SMB security model for protocol-level
enforcement and the distinction between session and per-share encryption.
Use dfsctl adapter settings smb update --signing required to require signing.
enabled offers signing and disabled relaxes it for older dialects; SMB 3.1.1
requires signing regardless of this setting. Signing algorithms are negotiated;
there is no preferred_algorithms CLI flag.
Use --min-dialect and --max-dialect on dfsctl adapter settings smb update,
with names such as SMB2.1, SMB3.0, SMB3.0.2, and SMB3.1.1. The example
above excludes SMB 2.x clients.
SMB lease breaks have a fixed 5-second bound. For traditional oplocks,
dfsctl adapter settings smb update --oplock-break-timeout 35 sets the wait in
seconds (supported range: 5–120). This does not change lease-break timing.
Directory leasing is controlled by the REST setting listed in the migration table.
Durable handles require a metadata store that implements durable-handle storage.
The adapter does not expose the removed timeout, capacity, enable, or scavenger
fields through the CLI or server config. The migration table records the current
runtime timeout and scan interval.
Cross-protocol coordination uses programmatic defaults rather than SMB YAML or
environment overrides. The delegation recall timeout is 90 seconds and the
anti-storm TTL is 5 seconds; neither has an SMB adapter settings flag.
DittoFS can advertise itself on the LAN so it appears in macOS Finder →
Network, Linux file managers (via Avahi), and the Windows Explorer →
Network view — the same job the external avahi-daemon and wsdd/wsdd2
daemons do for Samba, but built in-process. Discovery is off by default and
managed through the live adapter settings (dfsctl / REST), so it applies
immediately without an adapter restart:
This opens additional listeners: mDNS on UDP 5353, and — for WS-Discovery —
UDP 3702 plus an HTTP metadata endpoint on TCP 5357. NFS advertises the
first export’s path in a path= TXT record so Finder mounts the right share.
Advertised name. All advertisers share one instance-wide name, so a server
shows up consistently across Finder and Explorer. It defaults to
DittoFS-<hostname> (e.g. DittoFS-VM2) — distinct per host so several DittoFS
servers on one LAN stay distinguishable — and is overridable:
Terminal window
dfsctlsettingssetdiscovery.name"Marketing Files"# custom instance name
dfsctlsettingssetdiscovery.name""# revert to DittoFS-<hostname>
Each adapter formats the name for its own protocol: mDNS uses it verbatim, while
WS-Discovery folds it to a NetBIOS-legal computer name (upper-cased, illegal
characters replaced with -, capped at 15 characters) since Explorer renders it
as a Windows computer name. A name change takes effect the next time an
advertiser (re)starts — toggle discovery off/on, or restart the adapter.
Note: discovery is multicast-based and LAN-local. It works on a host
network (bare metal, VM, or a hostNetwork pod) but does not traverse
standard Kubernetes / overlay networks — Explorer and Finder will only see the
server on the same L2 segment. Mounting by name/IP always works regardless.
WS-Discovery makes the machine appear in Explorer; Windows still connects to
SMB on port 445, so the SMB adapter must be reachable there.
Windows host firewall: when DittoFS runs on a Windows host, the built-in
“Network Discovery” firewall rules only cover Windows’ own services (they are
scoped to System / svchost), so inbound traffic to dfs.exe is dropped by
default. Explorer then discovers the server over multicast but silently fails
the follow-up metadata fetch (a TCP 5357 connection that never completes),
and the machine never renders. Add inbound allow rules for the dfs.exe
program on TCP 5357 and UDP 3702/5353:
When netbios_domain is set, the SMB server advertises the AD domain in the
NTLM challenge (MsvAvNbDomainName / MsvAvDnsDomainName) and stamps it on
authenticated sessions, so domain users authenticate against the correct domain.
Unset → the server advertises WORKGROUP / local (pre-AD-4 standalone
behavior).
Offline keytab import (one keytab, both protocols)
A keytab can hold multiple service principals, so a single file serves SMB and
NFS. Pre-create the computer/service account in AD and export a keytab
containing bothcifs/<host>@REALM (SMB) and nfs/<host>@REALM (NFS):
Terminal window
# On a domain-joined admin host (samba-tool / adcli / Windows ktpass):
Point kerberos.keytab_path at the combined keytab. The SMB handler selects
the cifs/ principal (deriving it from the NFS service_principal, or via an
explicit override); NFS RPCSEC_GSS uses the nfs/ principal.
NTLM pass-through for AD domain users (NETLOGON machine account)
The keytab above authenticates AD users over Kerberos (mounting by the SPN
FQDN, e.g. \\server.example.com\share). But when a client connects by a name
that has no Kerberos SPN — an IP address, or the LAN-discovery name a user
gets by double-clicking the server in Explorer → Network (§10) — Windows
falls back to NTLM, which the KDC never sees. To let AD domain users
authenticate on that path, DittoFS validates their NTLM response against a
Domain Controller via NETLOGON pass-through (MS-NRPC NetrLogonSamLogon).
This is opt-in and requires a dedicated machine (computer) account —
distinct from the cifs/ service account in the keytab, because NETLOGON needs
a workstation-trust secure channel that only a machine account can establish.
The secure channel rides a Kerberos-authenticated SMB session to the DC’s
\PIPE\netlogon (reusing the same krb5_conf / realm as above). On success the
DC returns the user’s SID and group SIDs, which resolve to a UID/GID through
the directory idmap (§13 / the LDAP identity provider, or idmap_rid) and are
matched against share grants — the same SID-based authorization as the
Kerberos/PAC path, so a domain user or group (e.g. Domain Admins) granted
on a share is authorized identically over NTLM.
kerberos:
# ... realm / netbios_domain / keytab_path as above ...
machine_account:
enabled: true# opt-in; false (default) => no NTLM pass-through
account_name: "DITTOFS$"# the machine account sAMAccountName (trailing '$')
dc_address: # optional; empty => discover the DC via DNS SRV
- "10.0.0.10"
# Provisioning — pick ONE:
# (A) OFFLINE: you pre-create the computer account and give DittoFS its
# password. Nothing is written to AD at runtime.
secret: "the-machine-account-password"
# (B) ONLINE JOIN: DittoFS creates the computer object itself over LDAPS on
# first domain logon, owns the password, and rotates it. Requires a
# privileged bind account that can create computer objects. Omit
# `secret` when using this. LDAPS (or ldap:// + start_tls) is mandatory —
# AD refuses a machine-password write over an unencrypted connection.
realm and netbios_domain are required for pass-through (the secure
channel and NTLM TargetInfo both need them). The secret / bind_password
are redacted in dfs config show. For a full walkthrough — pre-creating the
DITTOFS$ account and testing the Explorer double-click — see
docs/guide/windows-ad-setup.md.
Controls the background scheduler that drives per-share snapshot policies
(schedule + retention). Policies themselves are configured per share via
dfsctl share snapshot-policy or the REST API — see
SNAPSHOTS.md §12. These
knobs only tune the daemon-wide scheduler loop.
snapshot:
# How often the daemon scans for due policies. The per-share policy
# interval (not this knob) controls how often a share is snapshotted.
scheduler_poll_interval: 1m# default 1m
# Turn the scheduler off entirely. Policies are still stored and can be
# run manually with `dfsctl share snapshot-policy run`.
scheduler_disabled: false# default false
# Per-request budget for the synchronous restore endpoint.
Resolves directory principals (a user@REALM form or an AD SID) to a Unix
identity by querying LDAP/AD. It reads the RFC2307uidNumber/gidNumber
POSIX attributes (the idmap_ad model) or falls back to RID-based
derivation (idmap_rid), and resolves the user’s group memberships — including
nested AD groups via the LDAP_MATCHING_RULE_IN_CHAIN matching rule. The
provider is registered in the identity-resolution chain after Kerberos, so an AD
principal/SID with no local user mapping is resolved against the directory.
Security: the connection is encrypted by default. A plaintext ldap://
connection is refused unless it is upgraded with start_tls: true or the
operator explicitly opts in with allow_plaintext: true. Prefer ldaps://.
ldap:
enabled: true
url: ldaps://dc.example.com:636# ldaps:// (preferred) or ldap:// + start_tls
start_tls: false# upgrade an ldap:// connection to TLS
allow_plaintext: false# explicit opt-in for an unencrypted bind (off)
(required when enabled; empty triggers an anonymous bind and is rejected)
ldap.user_attr
DITTOFS_LDAP_USER_ATTR
sAMAccountName
ldap.realm
DITTOFS_LDAP_REALM
(empty)
ldap.idmap
DITTOFS_LDAP_IDMAP
rfc2307
ldap.nested_groups
DITTOFS_LDAP_NESTED_GROUPS
false
ldap.max_group_results
DITTOFS_LDAP_MAX_GROUP_RESULTS
200
ldap.timeout
DITTOFS_LDAP_TIMEOUT
10s
ldap.tls.ca_cert_file
DITTOFS_LDAP_TLS_CA_CERT_FILE
(system roots)
ldap.tls.client_cert_file
DITTOFS_LDAP_TLS_CLIENT_CERT_FILE
(empty)
ldap.tls.client_key_file
DITTOFS_LDAP_TLS_CLIENT_KEY_FILE
(empty)
ldap.tls.insecure_skip_verify
DITTOFS_LDAP_TLS_INSECURE_SKIP_VERIFY
false
ldap.tls.min_version
DITTOFS_LDAP_TLS_MIN_VERSION
1.2
Admin-configured identity links (dfsctl idmap add) take precedence over the
directory query: a link for an ldap external ID resolves to the mapped local
user before any LDAP search is issued.
Samba AD-DC self-signed certificates. A default Samba AD-DC serves an
auto-generated TLS certificate that commonly has a negative serial number.
Go’s crypto/x509 rejects such certificates at parse time — before TLS
verification runs — so ldap.tls.insecure_skip_verify does not bypass it.
The dfs binary is built with x509negativeserial=1 so ldaps:// against a
default Samba AD-DC works out of the box. For production directories, prefer a
properly-issued DC certificate (the Go toggle is slated for removal in a future
release).
Managing identity providers over the API (no restart). The ldap.* and
kerberos.* keys above seed the configuration on first boot. After that, both
providers can be read, updated, and tested over the control-plane API without
editing files:
Method & path
Purpose
GET /api/v1/identity-providers
List providers + enabled state (no secrets).
GET /api/v1/identity-providers/{type}/config
Read config (bind password redacted to ********).
PUT /api/v1/identity-providers/{type}/config
Create/replace config (validated).
POST /api/v1/identity-providers/{type}/test
Dry-run dial+bind (LDAP) / keytab check (Kerberos); never persists.
{type} is ldap or kerberos; all routes are admin-only. A persisted config
takes precedence over the file/env config on subsequent boots. LDAP
changes hot-reload the live resolver; Kerberos changes take effect on the
next server restart (the NFS/SMB adapters bind it at startup). The bind
password is write-only — submit ******** (or omit it) on PUT to keep the
stored secret. Equivalent CLI: dfsctl identity-provider {list,get,set,test}
(see CLI.md).
Current servers store remote data as packed blocks/<id> containers. The
automatic startup conversion for shares still holding standalone-CAS state
(per-chunk cas/ objects and locators from v0.16-v0.21 servers) has been
removed: such a store cannot be upgraded in place by this build, and a
locator still pointing at a standalone object fails closed on read. See
the migration guide.
Pre-v0.16 .blk layouts: migrate with dittofs ≤ v0.21 first
The offline migrate-to-cas command was removed after v0.21. On startup,
dfs start still probes each share for the legacy .blk layout (a
.cas-migrated-v1 sentinel from an old migration short-circuits the probe)
and refuses to start un-migrated shares:
Exits with code 78 (EX_CONFIG per sysexits(3)).
Prints a directive to stderr naming the offending share and pointing at
dittofs ≤ v0.21 for the .blk migration.
Halts on the FIRST share that surfaces the legacy layout.
Run migrate-to-cas with a v0.21 binary, verify the per-share .cas-migrated-v1
sentinel exists, then upgrade — the automatic conversion above finishes the
job.
DittoFS exposes a Prometheus /metrics endpoint on a dedicated listener,
separate from the control-plane API and the protocol adapters. It is opt-in
and disabled by default. When enabled it serves the standard Go/process
collectors plus the DittoFS instruments (request RED metrics per adapter,
connection counts, sync/remote/local-store/quota/GC gauges, and snapshot
timestamps).
# Top-level metrics block (NOT nested under `server:`).
metrics:
enabled: true# opt-in; default false
host: 127.0.0.1# bind interface; default loopback only
port: 9090# default 9090
path: /metrics# default /metrics
auth: none# "none" (default) or "token"
token_file: ""# path to a file holding the Bearer token (auth: token)
tls: # optional; reuses the control-plane TLS shape
cert_file: ""
key_file: ""
client_ca: ""# set for mTLS
min_version: "1.2"# minimum TLS version: "1.2" (default) or "1.3"
Option
Default
Description
enabled
false
Turn the metrics listener on.
host
127.0.0.1
Bind interface. Binds loopback by default; set 0.0.0.0 to expose it deliberately (pair with a firewall/NetworkPolicy).
port
9090
TCP port for the endpoint.
path
/metrics
HTTP path. Must start with /.
auth
none
none (rely on bind host + network policy) or token (require a Bearer token).
token_file
File whose trimmed contents are the expected Bearer token. Required when auth: token.
tls.cert_file / tls.key_file
Serve the endpoint over HTTPS.
tls.client_ca
Require + verify client certificates (mTLS).
tls.min_version
1.2
Minimum negotiated TLS version: 1.2 or 1.3.
Security: the endpoint binds 127.0.0.1 by default so it is not reachable
off-host without an explicit host change. In production prefer one of:
bind loopback and scrape via a sidecar; restrict with a firewall/NetworkPolicy;
or enable auth: token (and/or TLS). Never expose 0.0.0.0 unauthenticated on
an untrusted network.
DittoFS is only a scrape target. It never runs, bundles, manages, or depends
on a Prometheus/Thanos/Mimir instance. Point your cluster’s existing Prometheus
at the endpoint. For production, the standard path is the
kube-prometheus-stack
(prometheus-operator) with long-term storage via Thanos or Mimir. DittoFS ships
no dashboards or compose artifacts — the guidance below is enough to wire it
into any standard stack.
DittoFS Pro bundles a turnkey monitoring stack on top of this endpoint: a
docker compose --profile monitoring profile that stands up Prometheus +
Grafana with a pre-provisioned DittoFS dashboard, plus an in-dashboard Metrics
section. The endpoint and config documented here are the contract it builds on.
The DittoFS Kubernetes operator wires this up for you when you opt in on the
DittoServer spec. It renders the metrics container port, a dedicated metrics
Service carrying prometheus.io/scrape annotations (for annotation-based
discovery), and — only if the monitoring.coreos.com CRDs are installed and
you enable it — a ServiceMonitor for the prometheus-operator. If those CRDs
are absent the operator skips the ServiceMonitor (logged) and never fails the
reconcile.
apiVersion: dittofs.dittofs.com/v1alpha1
kind: DittoServer
metadata:
name: example
spec:
storage:
metadataSize: 10Gi
metrics:
enabled: true# renders the metrics Service + scrape annotations
port: 9090
path: /metrics
# bearerTokenSecret: # optional authed scrape
# name: dittofs-metrics-token
# key: token
serviceMonitor:
enabled: true# requires the prometheus-operator CRDs
interval: 30s
labels: # match your Prometheus serviceMonitorSelector
Unix time of the last successful snapshot create (backup-freshness signal).
⭐ dittofs_integrity_findings{share,kind}
Findings from the last structural manifest scan, by kind: payloads_with_findings, damaged_payloads, claimed_uncovered_ranges, unplaceable_rows, unknown_hash_rows.
Unix time the structural manifest scan last completed. 0 means it has never run since process start or its last attempt failed.
⭐ dittofs_integrity_last_scan_failed{share}
1 when the most recent structural manifest scan failed, 0 when it completed or has never run. Pairs with the timestamp above to tell “never scanned” from “scans are erroring”.
Cost and reach of the last structural manifest scan.
⭐ dittofs_offline_safe{share}
1 when every byte the share holds can be served with the remote unreachable, else 0. A share whose residency cannot be determined reports 0 and publishes no byte counts.
Build (or import) a Grafana dashboard around the ⭐ core series: an adapter RED
row (rate of dittofs_adapter_requests_total, error ratio, latency percentiles
from dittofs_adapter_request_duration_seconds), a durability row
(dittofs_remote_up, dittofs_sync_pending_bytes, and snapshot success rate
from dittofs_snapshot_operations_total), and a capacity row
(dittofs_localstore_disk_used_bytes, dittofs_quota_used_bytes). DittoFS does
not ship a dashboard JSON; these series are stable and named for direct use.
Note: DITTOFS_ADMIN_INITIAL_PASSWORD is only used during the very first server start when the admin user is created. It has no effect on subsequent starts. When set, the admin account’s MustChangePassword flag is not enabled.
On the first start (when no admin user exists yet) the initial admin password is chosen in this precedence:
DITTOFS_ADMIN_INITIAL_PASSWORD (env, plaintext) — sets a known password and also derives the NT hash, so the admin can authenticate over SMB as well as the REST/control-plane API.
admin.password_hash (config, bcrypt $2a$/$2b$/$2y$) — sets a known credential without writing a plaintext secret to disk. No NT hash is derivable from a bcrypt hash, so an admin bootstrapped this way works for the control-plane/REST API only, not SMB (use option 1 for SMB). A value that is not a valid bcrypt hash is rejected at startup.
Auto-generated — a random password is generated. It is printed once only when dfs start runs with stdout attached to an interactive terminal (i.e. dfs start --foreground in a real TTY). In background/daemon mode, and under Docker/systemd/CI where stdout is a pipe, the password is not written to the log and cannot be recovered — the log only records that an admin user was created. Recovery via dfsctl user password admin is not possible in this state (it requires an authenticated admin session). Options 1 and 2 are consulted only while bootstrapping a new admin, so once this first start has created the admin user, setting them and restarting will not change the password — you must delete the admin row from the users table of the control-plane database (not the database itself, which also holds your shares, stores, mounts and other users) and re-bootstrap with option 1 or 2 set. For any non-interactive deployment, set option 1 or 2 before the very first start.
Options 1 and 2 also skip the forced first-login password change (the operator already chose the password). All three apply only on first start; later starts never change an existing admin’s password.
admin:
username: admin
# bcrypt hash, e.g. from `htpasswd -bnBC 10 "" 'my-secure-password' | tr -d ':\n'`
Store instances and protocol adapters are managed through the control-plane
API and dfsctl; their settings are persisted in the database.
After logging in, use dfsctl store metadata add to create a metadata store,
dfsctl adapter edit nfs --port 12049 to set the NFS port, and
dfsctl adapter settings nfs show or dfsctl adapter settings smb show to
inspect protocol settings. See the CLI reference for supported
settings and update flags.
Server settings are applied in the following order (highest to lowest priority).
This precedence does not apply to stores or adapters managed through dfsctl: