Skip to content

Advanced Configuration

This guide covers advanced Kingfisher features for power users.

Table of Contents

Baseline Management

There are situations where a repository already contains checked‑in secrets, but you want to ensure no new secrets are introduced. A baseline file lets you document the known findings so future scans only report anything that is not already in that list.

The easiest way to create a baseline is to run a normal scan with the --manage-baseline flag (typically at a low confidence level to capture all potential matches):

kingfisher scan /path/to/code \
  --confidence low \
  --manage-baseline \
  --baseline-file ./baseline-file.yml

--manage-baseline automatically enables --no-dedup so the baseline captures every individual occurrence.

Use the same YAML file with the --baseline-file option on future scans to hide all recorded findings:

kingfisher scan /path/to/code \
  --baseline-file /path/to/baseline-file.yaml

Running the scan again with --manage-baseline refreshes the baseline by adding new findings and pruning entries for secrets that no longer appear. See BASELINE.md for full detail.

Understanding Confidence Levels

The --confidence flag sets a minimum confidence threshold, not an exact match.

  • If you pass --confidence medium, findings with medium and higher confidence (medium + high) will be included.
  • If you pass --confidence low, you'll see all levels (low, medium, high).
# Only show high-confidence findings
kingfisher scan /path/to/code --confidence high

# Show medium and high confidence findings
kingfisher scan /path/to/code --confidence medium

# Show all findings (low, medium, high)
kingfisher scan /path/to/code --confidence low

Filtering and Suppression

Skip Known False Positives

Use --skip-regex and --skip-word to suppress findings you know are benign. Both flags may be provided multiple times and are tested against the secret value and the full match context.

With --skip-regex, these should be Rust compatible regular expressions, which you can test out at regex101

# Skip any finding where the finding mentions TEST_KEY
kingfisher scan --skip-regex '(?i)TEST_KEY' path/

# Skip findings that contain the word "dummy" anywhere in the match
kingfisher scan --skip-word dummy path/

# Combine multiple patterns
kingfisher scan \
  --skip-regex 'AKIA[0-9A-Z]{16}' \
  --skip-word placeholder \
  --skip-word dummy \
  path/

If a --skip-regex regular expression fails to compile, the scan aborts with an error so that typos are caught early.

Skip Canary Tokens (AWS)

Canary/honey tokens are intentionally leaked credentials used to catch misuse. Kingfisher can recognize and skip known AWS canary accounts so hygiene scans don't set off alerts.

How to skip
Pass the 12-digit AWS account IDs for your canaries via --skip-aws-account (comma-separated) or --skip-aws-account-file (one ID per line; blank lines and # comments allowed). Kingfisher also ships with a pre-seeded (but not exhaustive) list of Thinkst Canary account IDs used by canarytokens.org, so many are skipped automatically.

kingfisher scan /path/to/code \
  --skip-aws-account "171436882533,534261010715"

# or combine preloaded canary IDs with a just-created decoy account
printf '999900001111 \n534261010715' > /tmp/canary_accounts.txt

kingfisher scan /path/to/repo \
  --skip-aws-account-file /tmp/canary_accounts.txt

What you'll see
Findings tied to a skip-listed account report Validation: Canary Token (Skipped) and note in the Response: that the entry came from the skip list:

AWS ACCESS TOKEN => [BETTERLEAKS.AWS-ACCESS-TOKEN]
 |Finding.........: <REDACTED>
 |Fingerprint.....: 2141074333616819500
 |Confidence......: medium
 |Entropy.........: 5.00
 |Validation......: Canary Token (Skipped)
 |__Response......: (skip list entry) AWS validation not attempted for account 171436882533.
 |Language........: Unknown
 |Line Num........: 21
 |Path............: /tmp/test_canary_accounts.log

Why this matters
Skipping prevents noisy tripwires in prod telemetry while keeping the status explicit—"Canary Token (Skipped)" signals that the credential likely belongs to an active honeypot but was intentionally not validated. If needed, verify these credentials out-of-band or with a safe, non-triggering method.

Common CLI flows

# Skip a few in-house canaries during a filesystem scan
kingfisher scan repo/ \
  --skip-aws-account "111122223333,444455556666"

# Read a longer list from disk
kingfisher scan repo/ \
  --skip-aws-account-file /tmp/scripts/canary_accounts.txt

# Combine preloaded canary IDs with a just-created decoy account
printf '999900001111\n534261010715\n' > /tmp/new_canary.txt

kingfisher scan /path/to/repo \
  --skip-aws-account-file /tmp/new_canary.txt

Tip: if you manage multiple canary fleets (Thinkst, self-hosted alternatives, or bespoke decoys), checkpoint the account IDs alongside your infrastructure-as-code so security teams can rotate or expand the skip list without editing pipelines.

Inline Ignore Directives

Add kingfisher:ignore anywhere on the same line as a finding to silence it. Multi-line strings and PEM-style blocks may also be ignored by placing the directive on the closing delimiter line (for example, """ # kingfisher:ignore), on the next logical line after the string, or on a comment immediately before the value:

# kingfisher:ignore
API_KEY = """
line 1
line 2
"""
# kingfisher:ignore

Kingfisher searches the surrounding lines for these tokens without requiring language-specific comment markers. To reuse existing inline directives from other scanners, add them with repeatable --ignore-comment flags (for example --ignore-comment "gitleaks:allow" --ignore-comment "NOSONAR"). Use --no-ignore when you want to disable inline suppressions entirely.

Validation Tuning

Use these options with kingfisher scan to customize live validation behavior:

# Set per-request timeout (default: 10 seconds; 0 disables timeouts)
kingfisher scan /path/to/code --validation-timeout 15

# Set number of retry attempts (default: 1, range: 0-5)
kingfisher scan /path/to/code --validation-retries 2

# Increase validation response storage limit (default: 2048 bytes)
kingfisher scan /path/to/code --max-validation-response-length 8192

# Disable validation response storage truncation entirely (0 = unlimited)
kingfisher scan /path/to/code --max-validation-response-length 0

# Include full validation response bodies end-to-end (no validation or reporter truncation)
kingfisher scan /path/to/code --full-validation-response

# Combine options
kingfisher scan /path/to/code \
  --validation-timeout 20 \
  --validation-retries 3 \
  --max-validation-response-length 8192
  • --validation-timeout SECONDS: per-request and per-match timeout for validation (default: 10; 0 disables timeouts).
  • --validation-retries N: number of retry attempts for validation requests (default: 1, range: 0-5).
  • --max-validation-response-length BYTES: maximum bytes stored from validation response bodies (default: 2048; 0 disables truncation at storage time).
  • --full-validation-response: include complete validation response bodies end-to-end. This bypasses both storage-time truncation and reporter display truncation, and takes precedence over --max-validation-response-length.

Repository discovery and scanning overlap. Streaming output and audit records can arrive before discovery completes. Treat a nonzero exit caused by discovery or artifact-fetching errors as an incomplete scan, even if findings were printed. The audit log emits run_failed on these failures; only a successful run emits run_completed.

Scanning in CI Pipelines

Changes in the last N hours

Use --since-hours <HOURS> to scan recent Git changes without choosing a baseline commit. HOURS must be a positive whole number. With no --branch, Kingfisher walks all locally available refs (local branches, remote-tracking branches, and tags) plus HEAD:

# Clone once with all branches and complete history; no checkout is needed.
git clone --no-checkout --no-single-branch "$REPO_URL" ./repo

# Scan commits from the last 24 hours across all available refs.
kingfisher scan ./repo --since-hours 24 --git-history full --format toon

# Restrict the same window to one branch.
kingfisher scan ./repo --since-hours 24 --branch origin/main --format toon

--git-history full is the default, so it may be omitted. --since-commit is not needed.

For example, assuming the following committer timestamps:

                  scan start minus 24h        scan start
                            |                     |
      A---B-----------------|---C---D              |  main (HEAD)
           \                |                     |
            E---------------|---F---G              |  origin/feature
                            |                     |
         older commits      |   selected window   |

The first scan includes changes in C, D, F, and G. It excludes A, B, and E as scan candidates. Adding --branch main includes only C and D. A secret added in C and removed in D is still detectable because each selected commit is inspected, not just the final branch tree.

Time-window semantics:

  • Kingfisher fixes one inclusive window at scan start, from that instant minus HOURS through scan start, at Unix-second precision. Every repository in the same invocation uses the same window, even if cloning or scanning takes time. Commits dated after scan start are excluded.
  • Selection uses the committer timestamp, not the author timestamp, push/fetch time, or file modification time. Rebasing and cherry-picking can change committer dates. This option cannot reconstruct when a server received a push.
  • Commit dates need not follow ancestry order. Kingfisher continues traversing older commits to find qualifying ancestors; a time window reduces scanned file content but does not necessarily reduce graph traversal. The repository timeout still applies.
  • Each selected commit is compared with its first parent, even if that parent predates the window. Kingfisher scans the complete added or modified file versions, not only added lines. An older secret in a file modified during the window can therefore be reported. Deleted files and removed content are not scanned from the parent tree; a removed secret is detectable only if an added or modified file version containing it belongs to a selected commit. Unchanged files and uncommitted working-tree changes are excluded.
  • Only locally available history can be scanned. A shallow clone may omit qualifying commits or parent trees. When a selected shallow boundary has no usable parent history, Kingfisher scans its full tree, which may include older content. Prefer complete history for accurate parent comparisons; git clone --shallow-since is not equivalent to this scan filter.

--since-hours requires full Git history mode and cannot be combined with --since-commit, --staged, --branch-root, or --branch-root-commit. It filters Git commit content, not repository-host artifacts or other non-Git inputs. GitHub public-event scanning uses its separate --event-lookback-hours option and cannot be combined with --since-hours.

Audit output identifies this scope as commit_time_range and records the inclusive window in since_timestamp and until_timestamp (Unix seconds). All-ref scans use (all refs and HEAD) as tip_ref; an explicit branch records its resolved tip.

Commit ranges

With the default --git-history full, --since-commit <baseline> scans changes in every commit reachable from all locally available refs (including local branches, remote-tracking branches, and tags) plus HEAD, excluding the baseline ref and all its ancestors. Add --branch <ref> to restrict the walk to one tip (baseline..tip); explicit --branch HEAD retains HEAD-only scanning. This includes merged history and secrets added and removed before a tip. With --git-history none, Kingfisher instead scans the net tree diff from the baseline to --branch (default HEAD).

For example, after cloning all branches, scan their available history with:

kingfisher scan ./repo --since-commit "$BASE_COMMIT" --git-history full --format toon

Only refs and history present in the local clone can be scanned. Shallow boundaries limit traversal; a boundary commit whose parent is unavailable is scanned as a full tree. The baseline must resolve locally. --since-commit is an ancestry boundary, not a time filter: divergent branches may include older commits outside the baseline’s ancestry.

Use --branch-root-commit alongside --branch when you need to include a specific commit (and everything after it) in a diff-focused scan without re-examining earlier history. Provide the branch tip (or other comparison ref) via --branch, and pass the commit or merge-base you want to include with --branch-root-commit. If you omit --branch-root-commit, you can still enable --branch-root to fall back to treating the --branch ref itself as the inclusive root for backwards compatibility. This is especially useful in long-lived branches where you want to resume scanning from a previous review point or from the commit where a hotfix forked.

How is this different from --since-commit?
--since-commit excludes the baseline commit and its ancestors, and by default inspects each subsequent commit’s changes. --branch-root-commit instead computes a net tree diff from the parent of the supplied commit to the tip, including the supplied commit’s changes but potentially missing secrets removed before the tip. Use --since-commit <baseline> to audit a commit range, or add --git-history none when only the final delta is wanted.

kingfisher scan . \
  --since-commit origin/main \
  --branch "$CI_BRANCH"

Another example:

cd /tmp
git clone https://github.com/micksmix/SecretsTest.git

cd /tmp/SecretsTest
git checkout feature-1
#
# scan diff between main and feature-1 branch
kingfisher scan /tmp/SecretsTest --branch feature-1 \
  --since-commit=$(git -C /tmp/SecretsTest merge-base main feature-1)
#
# scan only the snapshot of a specific commit
kingfisher scan /tmp/SecretsTest \
  --branch baba6ccb453963d3f6136d1ace843e48d7007c3f --git-history none
#
# scan feature-1 starting at a specific commit (inclusive)
kingfisher scan /tmp/SecretsTest --branch feature-1 \
  --branch-root-commit baba6ccb453963d3f6136d1ace843e48d7007c3f
#
# scan feature-1 starting from the commit where the branch diverged from main
kingfisher scan /tmp/SecretsTest --branch feature-1 \
  --branch-root-commit $(git -C /tmp/SecretsTest merge-base main feature-1)
#
# scan from a hotfix commit that should be re-checked before merging
HOTFIX_COMMIT=$(git -C /tmp/SecretsTest rev-parse hotfix~1)
kingfisher scan /tmp/SecretsTest --branch hotfix \
  --branch-root-commit "$HOTFIX_COMMIT"

When the branch under test is already checked out, --branch HEAD or omitting --branch entirely is sufficient. Kingfisher exits with 200 when any findings are discovered and 205 when validated secrets are present, allowing CI jobs to fail automatically if new credentials slip in.

Tip: You can point Kingfisher at a local working tree and scan another branch or commit without changing checkouts. The CLI now resolves repositories from their worktree roots, so commands like the following work without needing to pass the .git directory explicitly:

kingfisher scan /path/to/local/repo --branch <ref>
kingfisher scan C:\\src\\repo --branch <commit-hash>

The same diff-focused workflow works when cloning repositories on the fly by passing a Git URL directly to scan. Kingfisher automatically tries remote-tracking names like origin/main and origin/feature-1, so you can target the branches involved in a pull request without performing a local checkout first.

kingfisher scan https://github.com/org/repo.git \
  --since-commit main \
  --branch development

When no explicit diff options (--branch-root, --branch-root-commit, or --staged) are supplied, --branch scans all history reachable from the requested ref, including merged branches. This is the default --git-history full behavior and finds secrets deleted in later commits without scanning unrelated branches or checking out the selected ref. Use --since-commit <ref> to exclude that ref and its ancestors from the history scan. Use --git-history none to scan only the selected ref’s snapshot, or the net diff when paired with --since-commit. Full-history enumeration takes more time and buffers blob metadata before scanning; the revision walk and commit diffs share one --git-repo-timeout budget. Increase that timeout for large histories when needed.

# Scan a branch from an existing checkout
kingfisher scan ~/tmp/repo --branch feature-123

# Or scan a branch when cloning on the fly
kingfisher scan https://github.com/org/repo.git \
  --branch origin/feature-123

In CI systems that expose the base and head commits explicitly, you can pass those SHAs directly while scanning a Git URL:

kingfisher scan https://github.com/org/repo.git \
  --since-commit "$BASE_COMMIT" \
  --branch "$PR_HEAD_COMMIT"

If you want to know which files are being skipped, enable verbose debugging (-v) when scanning, which will report any files being skipped by the baseline file (or via --exclude):

# Skip all Python files and any directory named tests, and report to stderr any skipped files
kingfisher scan ./my-project \
  --exclude '*.py' \
  --exclude tests \
  -v

Compiled Rule Cache

Kingfisher persists the compiled Vectorscan rule database by default so repeated short runs, such as pre-commit hooks and CI jobs, do not pay the full database compilation cost every time.

If no cache directory is specified, Kingfisher uses a platform default:

  • Windows: %LOCALAPPDATA%\Kingfisher\rule-cache
  • macOS: ~/Library/Caches/kingfisher/rule-cache
  • Linux/Unix: $XDG_CACHE_HOME/kingfisher/rule-cache, then ~/.cache/kingfisher/rule-cache
kingfisher scan . --staged

KF_RULE_CACHE_DIR=.kingfisher-cache kingfisher scan . --staged

kingfisher scan . --staged --rule-cache-dir .kingfisher-cache

For Docker runs, the default cache directory lives inside the container and is lost when the container is removed. On a Unix host, create a private directory, mount it and run as its owner so repeated docker run --rm scans reuse the cache:

install -d -m 0700 "$HOME/.cache/kingfisher-rule-cache"
docker run --rm \
  --user "$(id -u):$(id -g)" \
  -v "$PWD":/src \
  -v "$HOME/.cache/kingfisher-rule-cache":/kf-cache \
  -e KF_RULE_CACHE_DIR=/kf-cache \
  ghcr.io/mongodb/kingfisher:latest scan /src --staged

Docker-managed volumes or user-namespace mappings must likewise give the cache directory and entries to the UID visible inside the container. A root container cannot reuse an unprivileged user's cache merely because the directory is mounted.

Use --no-rule-cache to disable the cache for a scan:

kingfisher scan . --staged --no-rule-cache

To pre-warm the cache before the first scan, run:

kingfisher rules compile-cache

Cache pruning is opt-in. To remove old compiled databases during a scan, pass --prune-rule-cache. By default, pruning keeps at least 10 cache entries and removes only entries older than 30 days. Recognized cache temporary files left by interrupted writes are also pruned after the configured age, with a minimum one-day grace period; the entry retention floor does not apply to those temporary files. Tune those thresholds with --rule-cache-max-entries and --rule-cache-max-age:

kingfisher scan . --staged --prune-rule-cache

kingfisher scan . --staged \
  --prune-rule-cache \
  --rule-cache-max-entries 20 \
  --rule-cache-max-age 14d

To inspect or prune the cache without scanning, use kingfisher rules prune-cache:

kingfisher rules prune-cache --dry-run

kingfisher rules prune-cache \
  --rule-cache-max-entries 20 \
  --rule-cache-max-age 14d

By default, Kingfisher logs the cache directory in use. Pass --debug or -v to see cache hit/miss details and when new entries are written.

The cache key includes resolved rule order and patterns, CPU architecture, pointer width, endianness, cache format, exact binding/native crate versions and full Vectorscan build identity. Use the same wheel or engine build for deployment prewarming. A different operating system or native build may require recompilation; Vectorscan also checks CPU-feature compatibility during deserialization. This applies to built-in and custom rules. Missing, corrupt, stale or incompatible entries fall back to compilation and best-effort replacement; older cache formats recompile once.

Compiled databases are trusted native engine input. Cache directories and files must belong to the current user. Unix cache directories/files exclude group/other writes, and parent directories must prevent replacement. Windows DACLs also permit Administrators/SYSTEM. New Windows cache objects receive the process user SID as owner at creation, including for SYSTEM services. Existing paths still require that ownership. Protected Windows ancestors may additionally be owned or maintained by TrustedInstaller; other accounts must not be able to replace the cache. New Unix directories/files use 0700/0600; unsafe ownership, permissions, symlinks or nontrusted writable ACLs disable caching. There is no shared temporary-directory fallback: without a per-user location, scans compile in memory. SHA-256 payload verification detects corruption before native deserialization; it cannot authenticate writers with the same privileges. rules compile-cache requires successful persistence or reuse and fails if the cache is disabled, unsafe or unwritable.

These ownership checks apply to the deployed files; cache bytes are not bound to the builder's account or host. For images built as root and run as a service UID, transfer the directory and every entry with COPY --chown or recursive chown before deployment. Valid entries can load from a read-only filesystem; rejected entries still recompile, but storing a replacement requires a writable cache. See the SDK container prewarming example.

The Python SDK uses the same cache by default, including KF_RULE_CACHE_DIR. Pass Rules(cache_dir=...) to override the location or Rules(cache=False) to opt out. Prewarm with matching rules and confidence to share entries between the CLI and SDK. Only the main Vectorscan database is persisted; rule loading and confirmation/path/filter initialization still run.

Default Betterleaks Rules

Betterleaks supplies Kingfisher's main candidate detector catalog, with selected Veles detectors filling gaps. They load automatically and use the betterleaks. and veles. namespaces:

kingfisher scan --rule betterleaks.openai-api-key /path/to/repo
kingfisher validate --rule betterleaks.github-pat TOKEN

The betterleaks. prefix is optional for short selectors such as --rule github-pat. Custom rules loaded through --rules-path are additive by default. Use --load-builtins=false for a custom-only scan, or --no-builtins with validate and revoke.

Kingfisher omits Betterleaks' generic-api-key, generic-password, and generic-username rules from the built-in catalog because their broad patterns provide low signal at disproportionate scan cost. Use a targeted custom TOML or YAML rule when your organization needs generic credential detection for a known naming convention.

Kingfisher embeds 488 built-in rules from a prepared compressed bundle. Normal compilation does not download rule sources. Exact upstream inputs, licenses, hashes, and source revisions are archived under crates/kingfisher-rules/generated/. Maintainers regenerate the bundle with cargo run --locked -p kingfisher-rule-bundle; --refresh fetches pinned upstream sources, and --check verifies the prepared artifacts offline. License text, source headers, and provenance are preserved under crates/kingfisher-rules/generated/. Veles import is not available through --rules-path.

Detection regexes, path constraints, confidence changes, component dependencies, and Betterleaks validation expressions are translated into Kingfisher's runtime model. The top-level Betterleaks prefilter is stored once on the compiled rule database, compiled into its own Vectorscan database, and evaluated once per source path before content matching. The top-level finding filter and each rule's filter are combined and evaluated only after a candidate match; their regex helpers are likewise compiled once into shared Vectorscan databases. Betterleaks keywords are intentionally ignored because Vectorscan already performs candidate selection. Production validation runs in Rust; Go and Betterleaks itself are not runtime dependencies. Provider endpoint overrides are exposed to Betterleaks validation environment variables.

The pinned Betterleaks v2.0.0-rc.1 catalog uses both older namespaced helpers and newer expression helpers such as type, string, int, toJSON, len, filter, all, and strings.splitTrim. Kingfisher evaluates the imported validation expressions and preserves their provider-result checks. Upstream analyze expressions are not executed; blast-radius analysis continues to use Kingfisher's reviewed access-map handlers. Capture selection accepts both the v2 valueGroup field and the older secretGroup field in Betterleaks TOML.

The source prefilter applies only to Betterleaks rules. Custom and Veles rules still run on paths that Betterleaks excludes.

--blast-radius (alias --access-map) works for validated Betterleaks findings when Kingfisher has an access-map handler for the credential shape. Kingfisher maintains a checked-in capability overlay for access-map bindings and selected safe revocation actions. It contains no candidate detector regexes, but may add narrow operational filters and capability metadata; it is validated against the downloaded imported-detector IDs and components during bundle generation. Betterleaks 2.x supports revoke expressions; Kingfisher's importer currently uses the reviewed overlay instead of executing those expressions. Checksum templates remain available to Kingfisher custom rules; Betterleaks detectors rely on their upstream regex/filter behavior until the upstream schema exposes checksum metadata.

The betterleaks.gcp-api-key binding uses this path to run the bounded, read-only Google API-key mapper. It records exact accepted, restricted, invalid, and inconclusive probe outcomes without inferring access to untested methods; see Blast Radius for its allowlist and safety boundaries.

New generally useful rules and validation improvements must be contributed to the Betterleaks repository first. Do not add a new Kingfisher-owned built-in YAML catalog.

Custom Rules

Both the Kingfisher rule format (.yml/.yaml) and Betterleaks TOML (.toml) are fully supported for custom rules loaded with --rules-path.

First, review RULES.md to learn how to create custom Kingfisher rules.

Scan with only custom rules

To scan using only your own my_rules.yaml:

kingfisher scan \
  --load-builtins=false \
  --rules-path path/to/my_rules.yaml \
  ./src/

Add custom rules alongside built-ins

To add your rules alongside the built‑ins:

kingfisher scan \
  --rules-path ./custom-rules/ \
  --rules-path my_rules.yml \
  --rules-path team_rules.toml \
  ~/path/to/project-dir/

Scan a custom rules directory

--rules-path accepts a directory containing Kingfisher .yml/.yaml and Betterleaks .toml rule files, not just a single file. Point it at any rules directory to load every rule file it contains.

Custom-only (no built‑ins loaded):

kingfisher scan --load-builtins=false --rules-path path/to/rules/ <target>

Additive (your rules run alongside the Betterleaks built‑ins):

kingfisher scan --rules-path path/to/rules/ <target>

Note that a custom-only scan (--load-builtins=false) has no Betterleaks source prefilter, so paths that the built-in catalog would otherwise exclude are scanned by your rules. See Default Betterleaks Rules above.

Check custom rules

# Check custom rules - ensures all regexes compile and match rule examples
kingfisher rules check --rules-path ./my_rules.yml

# List all built-in rules
kingfisher rules list

Scan using a rule family

(prefix matching: --rule betterleaks.aws loads the Betterleaks AWS family)

# Only apply Betterleaks AWS-related rules
kingfisher scan /path/to/repo --rule betterleaks.aws

Rule Performance Profiling

Use --rule-stats to collect timing information for every rule. After scanning, the summary prints a Rule Performance Stats section showing how many matches each rule produced along with its slowest and average match times. Useful when creating rules or debugging rules.

kingfisher scan /path/to/repo --rule-stats

Control Scan Concurrency

--jobs sets the scanner worker pool used for matching and Git scans. By default, Kingfisher uses the available logical CPUs, capped at approximately one worker per GiB of RAM. Reduce the value on memory-constrained CI runners or when scanning very large Git histories; a lower value trades some throughput for lower peak memory usage.

# Use four scanner workers for a large repository
kingfisher scan /path/to/large-repo --jobs 4

For project config, set scan.jobs. Pass --jobs on the command line when the Tokio runtime and scanner pool must use the same explicit value; see Project Configuration caveats.

Disk Offload

--disk-offload trades temporary storage I/O for lower memory use while findings accumulate across repositories or input roots. It is disabled by default. It preserves captured values, commit metadata, dependency helpers, validation results, and global deduplication, without changing worker count or Git history coverage.

kingfisher scan /path/to/repos --scan-nested-repos --disk-offload --format toon

When to use it

Enable it for scans of many repositories or input roots when findings from completed inputs consume memory needed by the inputs still being scanned. It is most useful when those completed inputs produce many findings or large associated metadata or validation responses, and a fast local temporary-storage volume has sufficient free space. It lets you reduce that accumulation without reducing --jobs or narrowing history coverage.

It is less useful for a single repository or scans with few findings. It does not limit the memory needed to scan an individual repository, hold concurrently active repositories, or perform final processing. All findings are restored to RAM before the final consumers run, so it cannot prevent an out-of-memory failure if that final working set is too large. Digest and blob-ID indexes also remain in memory throughout the scan.

For example, when kingfisher scan ~/example-repo --disk-offload --format toon scans ~/example-repo as a single repository, all of that repository's findings accumulate in RAM until its scan finishes. They are then written to disk and loaded back into RAM for final processing. Similar memory usage with and without the flag is expected for this workload, including similar peak memory. Offloading happens between completed repositories or input roots, not between files, commits, or finding batches within an active repository. The flag does not bound total scan memory or reduce Git scanning and worker memory.

Disk-backed storage adds serialization, writes, and reads; it is not a speed optimization. Measure elapsed time and peak memory on your workload before enabling it routinely. A memory-backed temporary filesystem such as tmpfs still consumes system memory, so it may not relieve overall memory pressure even if the scanner's own resident memory falls.

How it works

  1. Kingfisher opens a temporary file when the option is enabled.
  2. After a completed repository or input root's results are merged into the accumulated store, Kingfisher appends those findings as internal JSON Lines records and flushes the write buffer. Only after the write succeeds does it release that accumulated in-memory batch and its metadata maps. With parallel scans, active repositories still keep their own working sets.
  3. Before final processing, Kingfisher reads the records back in batches, reconstructing the findings and shared metadata. Final deduplication, validation, and reporting still operate on an in-memory store. After successful restoration, the temporary file is truncated to zero bytes; its handle is closed when the store is dropped or the process exits.

This internal representation is independent of report serialization. It contains unredacted credentials, including when --redact is used, so later processing receives the original values. It is not an exported report, saved scan, or resumable checkpoint. If temporary storage fills up, Kingfisher attempts the in-memory fallback described below. Other creation, write, or restore failures fail the scan explicitly rather than silently discarding findings.

If temporary storage fills up

If the OS reports that temporary storage is full while creating the file, Kingfisher prints a warning and continues with findings in memory. If it fills up during a write, Kingfisher:

  1. Rolls back the incomplete append to the last successfully written batch. The current batch is still in memory and has not been discarded.
  2. Prints a warning, restores all previously stored findings to memory, and keeps the pending findings alongside them, preserving deduplication and metadata.
  3. Closes the temporary file and disables disk offload for the rest of that scan. It does not repeatedly retry a full disk.

This allows the scan to continue without losing findings, provided enough RAM is available. Memory use may increase substantially; the fallback cannot guarantee completion if RAM is also exhausted. The disk must still be readable. If rollback or restoration fails, Kingfisher stops with an error instead of continuing with an incomplete result. Other I/O errors, such as denied access or corrupt records, remain fatal. This fallback applies only to temporary findings storage; it does not recover failures writing reports, fetching repositories, or extracting archives.

Privacy and location

The file is created with tempfile::tempfile() in the operating system's temporary directory. Kingfisher retains an open file handle, not a user-facing filename. It does not place the temporary file in the scanned repository, report output directory, or --git-clone-dir.

“Private” and “anonymous” describe file access and lifetime, not encryption:

  • Linux: where supported, the file is created without a directory entry using O_TMPFILE. Otherwise, it uses the Unix fallback below.
  • macOS and the Unix fallback: a randomly named file is created with owner-only read/write permissions (0600, further restricted by the process umask), then immediately unlinked. Unlinking removes its directory entry while the open handle keeps its contents available to Kingfisher. After unlinking, another process cannot simply open it by its former pathname.
  • Windows: a randomly named temporary file is opened without file sharing and with delete-on-close enabled. Its name can remain visible until the handle closes, so “anonymous” does not mean an invisible directory entry on Windows. Ordinary attempts to open the file while Kingfisher holds it are denied by the sharing mode.

These protections do not isolate the data from an administrator or a process with sufficient rights to inspect Kingfisher or its open handles. The contents are not encrypted by Kingfisher, and deletion is not secure erasure. Use suitably protected temporary storage if captured credentials must be encrypted at rest.

On Unix, set TMPDIR before launching Kingfisher to choose a different existing temporary directory, for example on a protected local disk with enough free space:

TMPDIR=/path/to/private-temp kingfisher scan /path/to/repos \
  --scan-nested-repos --disk-offload --format toon

Cleanup and crashes

Normal completion truncates the temporary file after restoration and closes its handle when the store is released. On an error, closing the handle also releases the temporary file, even if its contents were never restored. No separate Kingfisher cleanup command or scheduled sweep is needed for the normal case.

If Kingfisher panics, aborts, or is forcibly terminated (including SIGKILL on Unix), the OS closes its handles when the process exits. An already-unlinked Unix file is reclaimed after its last handle closes; Windows deletes the file through its delete-on-close setting. This cleanup does not depend on Rust destructors or a signal handler running. The temporarily stored findings are lost and cannot be used to resume the interrupted scan.

There are limits to that guarantee. In the Unix named-file fallback, termination in the brief interval between creation and unlinking can leave a randomly named file behind; an unlink failure can also leave a name behind. A whole-machine crash or power loss additionally depends on filesystem recovery and is not a secure-erasure guarantee. Kingfisher does not run a startup sweep for such remnants. If a leftover is found in the chosen temporary directory, treat it as sensitive data and remove it after confirming it is no longer in use. Filesystem snapshots, backups, or recoverable storage blocks can also retain data after deletion.

Temporary input content

Cloned repositories, staged stdin, Docker layers and decoded archive members can contain credentials independently of findings. Kingfisher's owned staging directories use Unix mode 0700 (the umask may restrict it further). Windows inherits the temporary parent's DACL. Choose a protected temporary parent on every platform; on Unix, TMPDIR selects it. If you supply --git-clone-dir, its permissions remain your responsibility. Normal cleanup removes owned temporary directories unless clones are retained, but crashes can leave remnants. Deletion is not secure erasure; use encrypted or memory-backed storage when needed.

Unlimited scans

Use --no-limits when scan completeness takes priority over bounded time, memory, and temporary disk usage:

kingfisher scan /path/to/code --no-limits --format toon

The flag overrides individual resource limits, including values supplied in a configuration file. Defaults are unchanged when it is absent. It removes:

  • File-size limits, archive entry/count/output budgets, and nested archive depth limits, including archives in Git history and Docker layers.
  • SQLite row/output budgets and lock waits, Python bytecode extraction budgets, Base64 size/depth limits, and credential candidate-count budgets.
  • Repository scan and Git clone/update deadlines, source-provider request timeouts, validation deadlines, and validation response body/storage caps.
  • Repository clone counts and provider search result counts. API page sizes remain bounded, and pagination continues until the provider reports completion.

Worker counts, bounded queues, request rate limits, finite retry counts, scan selectors/exclusions, --no-base64, --no-extract-archives, detection rules, and archive path checks and ancestor-cycle detection remain in effect. ZIP extraction still switches to disk for large inputs. Scan-time blast-radius network deadlines are also disabled, while its enumeration and evidence caps remain in effect. Report presentation limits, operating system limits, remote service limits, and format/protocol validity checks are outside this scan resource policy. A stalled scan can wait indefinitely; memory or disk exhaustion can still prevent completion.

For selective control, --max-file-size 0, --extraction-depth 0, --git-repo-timeout 0, and --validation-timeout 0 disable their respective limits. Validation timeouts and archive depth also accept positive values above the former 60-second and 25-layer maxima. KF_GIT_CLONE_TIMEOUT_SECS=0 and KF_GIT_UPDATE_TIMEOUT_SECS=0 disable the corresponding Git subprocess deadline. Negative, NaN, and infinite file sizes are rejected.

To enable the preset in an explicitly loaded configuration file:

scan:
  no_limits: true

Notable Scan Options

Git clone/update commands ignore global and system Git configuration and enable core.longpaths=true on Windows. Set KF_GIT_BINARY to select a Git executable; the same executable is used for local staged scans. Provider credentials must be supplied through Kingfisher's authentication options rather than inherited Git credential helpers.

For HTTPS repositories using certificates in the Windows Certificate Store, set KF_GIT_SSL_BACKEND=schannel when using Git for Windows. Other installations can select openssl if their Git supports it. With OpenSSL, GIT_SSL_CAINFO and GIT_SSL_CAPATH can specify a CA file or directory. Selecting a TLS backend does not disable certificate verification.

  • --jobs <N>: Set the number of parallel scanner workers; see Control Scan Concurrency.
  • --disk-offload: Store accumulated repository findings in a private temporary file; see Disk Offload.
  • --no-dedup: Report every occurrence of a finding instead of grouping repeated credential content
  • --include-hidden-findings: Include hidden helper-rule matches in reports and scan summary counts (diagnostic use)
  • --no-base64: By default, Kingfisher finds and decodes base64 blobs and scans them for secrets. This adds a slight performance overhead; use this flag to disable
  • --confidence <LEVEL>: (low|medium|high)
  • --min-entropy <VAL>: Override default threshold
  • --include-contributors: When scanning GitHub or GitLab URLs, include contributor-owned repos in the scan
  • --git-clone-dir <DIR>: Choose the parent directory for cloned repos and scan artifacts (use with Git URL scans)
  • --keep-clones: Preserve cloned repositories on disk after a scan completes
  • --repo-clone-limit <N>: Cap GitHub and GitLab clone targets when enumerating users, orgs/groups, or contributor repos; this includes opted-in GitHub gists and GitLab snippets. The first unique repositories discovered are selected, so provider ordering changes can change the selected subset across runs.
  • --no-binary: Skip binary files
  • --no-extract-archives: Do not scan inside archives
  • --extraction-depth <N>: Specifies how deep nested archives should be extracted and scanned (default: 2; 0 = unlimited)
  • --redact: Replaces discovered secrets with a one-way hash for secure output
  • --exclude <PATTERN>: Skip any file or directory whose path matches this glob pattern (repeatable, uses gitignore-style syntax, case sensitive)
  • --baseline-file <FILE>: Ignore matches listed in a baseline YAML file
  • --manage-baseline: Create or update the baseline file with current findings (automatically enables --no-dedup)
  • --skip-regex <PATTERN>: Ignore findings whose text matches this regex (repeatable)
  • --skip-word <WORD>: Ignore findings containing this case-insensitive word (repeatable)
  • --skip-aws-account <ACCOUNT_ID>: Skip live AWS validation for findings tied to the specified AWS account number (repeatable, accepts comma-separated lists)
  • --skip-aws-account-file <FILE>: Load AWS account numbers to skip from a file (one account per line; # comments allowed)
  • --ignore-comment <DIRECTIVE>: Honor additional inline directives from other scanners (repeatable; e.g. --ignore-comment "gitleaks:allow")
  • --no-ignore: Disable inline directives entirely so every match is reported
  • --no-ignore-if-contains: Ignore the ignore_if_contains filter in rules so placeholder words still produce findings
  • --validation-timeout SECONDS: per-request and per-match timeout for validation (default: 10; 0 disables timeouts).
  • --validation-retries N: number of retry attempts for validation requests (default: 1, range: 0-5).
  • --max-validation-response-length BYTES: maximum bytes stored from validation response bodies (default: 2048; 0 disables truncation at storage time).
  • --full-validation-response: include complete validation response bodies end-to-end (bypasses storage and reporter truncation).

Exclude specific paths

# Skip all Python files and any directory named tests
kingfisher scan ./my-project \
  --exclude '*.py' \
  --exclude '[Tt]ests'

Scan while ignoring likely test files

--exclude skips any file or directory whose path matches this glob pattern (repeatable, uses gitignore-style syntax, case sensitive)

# Scan source but skip likely unit / integration tests
kingfisher scan ./my-project \
  --exclude='[Tt]est' \
  --exclude='spec' \
  --exclude='[Ff]ixture' \
  --exclude='example' \
  --exclude='sample'

Limit maximum file size scanned

By default, Kingfisher skips files larger than 256 MB. You can raise or lower this cap per run with --max-file-size, which takes a value in megabytes. Use --max-file-size 0 for unlimited file size, or --no-limits to disable scan resource budgets and timeouts.

# Scan files up to 500 mb in size
kingfisher scan /some/file --max-file-size 500

Customize the HTTP User-Agent

Kingfisher identifies its HTTP requests with a user-agent that includes the binary name and version followed by a browser-style string. Some environments require extra context, such as a contact address, a change-ticket number, or a temporary test label. Use the global --user-agent-suffix flag to append this information between the Kingfisher identifier and the browser portion:

# Attach a contact email to all outbound validation requests
kingfisher --user-agent-suffix "contact=security@example.com" scan path/

# Label a one-off experiment
kingfisher --user-agent-suffix "Sept 2025 testing" scan github --user my-user --list-only

When omitted, Kingfisher defaults to kingfisher/<version> Mozilla/5.0 .... The suffix is trimmed; passing an empty string has no effect.

Finding Fingerprints

Kingfisher separates its location-sensitive reported fingerprint from its default, credential-focused scan deduplication. The document below explains both identities, why repeated locations are normally grouped into one actionable credential, and how --no-dedup reports every individual occurrence.

See FINGERPRINT.md for complete details.

Update Checks

Kingfisher automatically queries GitHub for a newer release when it starts and tells you whether an update is available. The check is informational only — the binary is not modified unless you explicitly opt in.

  • Update and exit – Run kingfisher self-update (alias kingfisher update) to download the latest release, replace the running binary in place, and exit. No scanning occurs.

  • Update then run with the new version – Pass the global --self-update flag (alias --update) on any scan or other command. If a newer release exists, Kingfisher downloads it, replaces the on-disk binary, and re-execs into the freshly installed binary so the current invocation completes with the new code (including the latest detection rules). On Unix this is a true exec() (same PID); on Windows the new binary is spawned and the parent exits with its status code. If no update is available, the command runs normally with no extra steps.

  • Disable version checks – Pass --no-update-check to skip both the startup and shutdown checks entirely. Recommended for CI runs to keep behavior reproducible.

Self-update writes to wherever the running binary lives, so it requires the calling user to have write access to that location. If you installed Kingfisher via a package manager (Homebrew, the .deb/.rpm packages, the PyPI wrapper, etc.), update through that package manager instead, even if you have root access: self-update bypasses the package manager's version and file tracking. For downloaded Linux packages, download and verify the newer release, then install the local package with DNF/Yum or APT; a repository upgrade command alone does not fetch GitHub release assets. See Linux package installation and RPM migration.

Self-update supports all six release platforms: Linux x64/arm64, macOS x64/arm64, and Windows x64/arm64.

Exit Codes

Code Meaning
0 No findings
1 Scan or runtime error
2 Invalid command-line arguments
3 No inputs discovered to scan
200 Findings discovered
205 Validated findings discovered

Input discovery can return exit code 3 when, for example, a GitHub user has no matching repositories or all discovered repositories are excluded. Orchestrators can use this code without matching the No inputs to scan error message. This condition stops before a scan report is produced.