Skip to content

Kingfisher Library Crates

For in-process Python detection, validation and revocation, see the Python SDK.

The prepared library releases are kingfisher-core 1.0.3, kingfisher-rules 1.2.0, and kingfisher-scanner 1.3.0. They require Rust 1.99 or newer and are versioned independently of the kingfisher-bin CLI, currently 2.11.1. See publishing.

Crate Overview

Crate Use it for
kingfisher-core Content buffers (Blob), identifiers, locations, provenance, entropy, ValidationOutcome
kingfisher-rules Custom rule loading, embedded catalog, RuleSyntax, compiled RulesDatabase
kingfisher-scanner Synchronous Scanner, ScannerConfig, owned Finding results, optional validators

The scanner re-exports Blob, Rule, RuleSyntax, RulesDatabase, and get_builtin_rules for common embedding tasks. Depend on the other crates directly when you need their additional APIs. None of the three depends on kingfisher-bin. Use Validator from kingfisher-scanner for live checks and Revoker for explicit, rule-driven revocation. Both are available with validation.

The kingfisher-bin package also exports the full application library as kingfisher. Its existing module paths remain available for consumers that need CLI orchestration or provider integrations. Prefer the focused crates above for new embedding work. CLI configuration merging, configuration generation, and rule commands are private binary modules, rather than additional public APIs. Public fallible scanner, rule-compilation, validator-builder, and revocation methods document their error conditions in rustdoc.

For lower-level matching, kingfisher_scanner::primitives::CandidateMatchIndex::new_in_range indexes a byte range while retaining offsets relative to the complete input. Confirmation windows outside that range use the original regex search, allowing segment-sized indexes without limiting match length. CandidateMatchCache::get_or_insert_with defers index construction until repeated endpoints justify the work; consider each distinct endpoint once, before widening its confirmation window. These existing low-level embedding APIs retain their signatures and original confirmation enum. The CLI's additional indexes and opaque optimized iterator are exposed only through __cli-internals, an unsupported feature whose API may change without notice. The rule crate's __scanner-internals feature is likewise an unsupported implementation detail enabled by the scanner crate.

Quick Start

Add the scanner crate to your application's Cargo.toml:

[dependencies]
kingfisher-scanner = "1.3.0"
anyhow = "1"
use std::sync::Arc;
use kingfisher_scanner::{get_builtin_rules, RulesDatabase, Scanner};

fn main() -> anyhow::Result<()> {
    let database = RulesDatabase::from_rule_collection(get_builtin_rules(None)?)?;
    let scanner = Scanner::new(Arc::new(database));
    let findings = scanner.scan_bytes(b"ordinary application configuration")?;
    for finding in findings.into_iter().filter(|finding| finding.rule().visible()) {
        // Avoid logging the secret or capture values.
        println!("{} at line {}", finding.rule_id, finding.line());
    }
    Ok(())
}

The quick starts in each crate README are also its rustdoc documentation, and Cargo executes their Rust examples as doctests. This keeps the advertised entry points checked against the implementation.

Runnable Examples

Each publishable package includes runnable source examples. Run these from the repository:

Package Example Command
kingfisher-core Borrow a blob and resolve locations cargo run -p kingfisher-core --example blob_locations
kingfisher-rules Load and compile rules cargo run -p kingfisher-rules --example load_rules
kingfisher-scanner Inspect rules, regexes and configured actions cargo run -p kingfisher-scanner --example inspect_rules -- --with-revocation
kingfisher-scanner Scan, redact, and share across threads cargo run -p kingfisher-scanner --example scan_content
kingfisher-scanner CLI-compatible detection with explicit limits cargo run -p kingfisher-scanner --features context --example detection_policy
kingfisher-scanner Read-only Git scopes and provenance cargo run -p kingfisher-scanner --features git --example git_inputs -- path/to/repo
kingfisher-scanner Explicit local validation cargo run -p kingfisher-scanner --example local_validation --features validation
kingfisher-scanner Scan file batches to JSON Lines cargo run -p kingfisher-scanner --example scan_files -- Cargo.toml README.md
kingfisher-scanner Scan with private YAML/TOML rules cargo run -p kingfisher-scanner --example scan_custom_rules -- crates/kingfisher-scanner/examples/fixtures/acme-http.yml README.md
kingfisher-scanner Bound scanning work in Tokio cargo run -p kingfisher-scanner --example scan_async
kingfisher-scanner Validate a file with built-in rules cargo run -p kingfisher-scanner --example validate_file --features validation -- path/to/config.env
kingfisher-scanner Scan and validate via YAML HTTP cargo run -p kingfisher-scanner --example http_validation --features validation
kingfisher-bin Embed the full application's library cargo run -p kingfisher-bin --example embedded_application

The load_rules example accepts a custom TOML/YAML path after --. The scan_content example accepts a file path after -- to use the built-in catalog; with no arguments it scans synthetic tokens, redacts output, and makes no network requests. scan_files and scan_custom_rules require the paths shown above. CI exercises these examples, including both custom rule formats and HTTP outcome classification. They are included in the published crate archives. The HTTP example starts its own loopback mock; it makes no external provider requests. validate_file explicitly contacts providers for detected credentials; enable the additional validator features your project needs.

To install the CLI from crates.io:

cargo install --locked kingfisher-bin --version 2.11.1
kingfisher scan path/to/project --no-validate --format toon --no-update-check

For the kingfisher-bin library example in your own application, use kingfisher = { package = "kingfisher-bin", version = "2.11.1" } and anyhow = "1". The focused library examples use the crate versions listed above; kingfisher-rules and kingfisher-scanner examples also use anyhow = "1".

Integration Recipes for Rust Projects and LLM Agents

Start with a complete example above and copy it into src/main.rs in a small Rust application. These are public-API consumers; no CLI process or private Kingfisher module is required. Choose dependencies from this table in addition to kingfisher-scanner = "1.3.0" and anyhow = "1":

Recipe Additional dependencies / features
scan_content None
detection_policy Scanner feature context
inspect_rules serde_json = "1"; kingfisher-rules = "1.2.0" for loading custom YAML/TOML
scan_files serde_json = "1"
scan_custom_rules kingfisher-rules = "1.2.0"
scan_async tokio = { version = "1.53", features = ["macros", "rt", "sync"] }
local_validation Scanner feature validation; kingfisher-rules = "1.2.0"
http_validation Complete manifest below; also copy fixtures/acme-http.yml and support/mod.rs into src/fixtures/ and src/support/
validate_file Scanner feature validation; serde_json = "1"; tokio = { version = "1.53", features = ["macros", "rt"] }
git_inputs Scanner feature git

The examples use the published crate versions. For local development, replace each Kingfisher dependency version with a path to the corresponding crate in a local checkout, preserving its features. Use Rust 1.99+ and the native build prerequisites in Build and Deployment.

Inspect catalog rules and validation/revocation definitions

Use the inspect_rules example for offline catalog discovery or an exact rule's complete loaded definition. It requires no validation feature:

cargo run --locked -p kingfisher-scanner --example inspect_rules -- --with-validation --with-revocation
cargo run --locked -p kingfisher-scanner --example inspect_rules -- betterleaks.aws-access-token
cargo run --locked -p kingfisher-scanner --example inspect_rules -- betterleaks.aws-access-token --field pattern --field validation --field revocation
cargo run --locked -p kingfisher-scanner --example inspect_rules -- --rules-path company.yml --no-builtins

The example loads all confidence levels and lists ID, name, visibility and configured action support. Use --id-prefix or capability flags for catalog filters, an exact ID for details, and repeated --field to select detail sections. The code comments show where to add predicates for confidence, entropy, visibility, dependencies or pattern requirements.

For your own application, use these public APIs:

API Inspection data
database.rules() Loaded Arc<Rule> entries; iterate or find by exact rule.id()
rule.syntax() Full serializable definition, including pattern, path, confidence, entropy, filters, capture selection, dependencies, examples and references
database.anchored_regexes()[index].as_str() Already-compiled Rust confirmation pattern with comments removed; index matches database.rules()
rule.syntax().validation Configured validation type and content, or None
rule.syntax().revocation Configured revocation type and content, or None

Serialize rule.syntax() or individual fields with serde_json::to_string_pretty. This avoids maintaining a second schema and preserves the loaded definition. HTTP/gRPC configurations expose requests and response matchers; Betterleaks validation exposes a portable expression tree, components and operational capabilities. Original expression text may be absent in release builds. Typed/raw validators identify Rust dispatch handlers; their implementation source is not serialized. Absent configurations become JSON null.

The historical anchored_regexes() accessor returns original search patterns, without the internal endpoint wrapper. A regex match alone is a candidate: reported findings still pass capture selection, entropy, filters and dependency requirements. Inspection makes no validation/revocation requests, and configured actions alone do not establish successful validation or revocation. Custom rule literals are returned as configured; this is not finding redaction.

Scan strings, uploads, and application configuration

Use scan_content for the smallest integration. Compile the rules once at startup and retain an Arc<Scanner> in application state. Pass bytes from configuration, uploads, or generated source to scan_bytes; use scan_blob_at_path when a logical filename matters to rule filters. Decide whether your product blocks on any visible finding or only flags it for review. An empty successful scan means no matches under the selected rules, not proof that the input contains no secrets.

Bound a scan or cancel it from another thread

Use scan_bytes_with_control, scan_file_with_control, or scan_blob_at_path_with_control to apply a per-call ScanControl:

use std::time::Duration;
use kingfisher_scanner::{CancellationToken, ScanControl};

let cancellation = CancellationToken::default();
let control = ScanControl::default()
    .with_timeout(Duration::from_secs(2))?
    .with_cancellation(cancellation.clone());
let findings = scanner.scan_bytes_with_control(content, &control)?;

Another thread can call cancellation.cancel(). Cancellation is permanent for that token. An interrupted call returns an error containing ScanAborted::TimedOut or ScanAborted::Cancelled; it does not return partial results or update the dedup cache. Existing methods remain unlimited. Deadlines are cooperative: checks run between matching/filtering operations, in Vectorscan callbacks, and during Base64 enumeration. File reads, decoding, rule compilation and individual native operations cannot be preempted. Use a separate process if your service needs a hard execution limit. A deadline starts when the control is constructed and covers the entire call.

Scan files in a build tool or CI gate

scan_files accepts paths as OS strings, preserves each file's source path, and emits one JSON object per visible finding. Identical credentials in separate files remain separately reportable because cross-call deduplication is off. Select files using your application's existing traversal or Git integration; the scanner itself does not traverse a repository. The example returns errors for unreadable files and reports findings without failing the process. For an enforcing gate, track whether any findings were returned and choose a separate nonzero exit code for detections.

The report deliberately includes only path, rule ID, location, and validation state. Avoid serializing entire Finding objects into logs or API responses.

Add private detectors without rebuilding Kingfisher

scan_custom_rules loads a YAML file, a Betterleaks TOML file, or a directory of both formats, then compiles and scans with that collection. This example uses only the custom collection; it does not implicitly add built-ins. TOML IDs gain a custom. prefix. The included Acme YAML fixture contains a synthetic token pattern, positive/negative samples, and a validation request. Loading it still performs detection only. See rule authoring for adapting patterns and adding supporting credential components.

Add scanning to an async request handler or ingestion worker

scan_async compiles on a blocking worker, shares the resulting scanner, moves owned input into spawn_blocking, and holds a semaphore permit until scanning finishes. Both task failures and scan failures propagate with await??. Retain the scanner and semaphore in your service state instead of rebuilding them per request. The example processes two small inputs; for an unbounded stream, also drain completed results incrementally and enforce input-size limits. Cancelling the async caller does not stop an already-running blocking scan.

Scan and validate HTTP credentials end to end

Copy http_validation.rs and its YAML fixture with this manifest:

[package]
name = "secret-check"
version = "0.1.0"
edition = "2024"
rust-version = "1.99"

[dependencies]
anyhow = "1"
kingfisher-scanner = { version = "1.3.0", features = ["validation"] }
kingfisher-rules = "1.2.0"
serde_json = "1"
tokio = { version = "1.53", features = ["macros", "rt", "net", "io-util", "time"] }

Run cargo run. The four output rows have outcomes verified_active, verified_inactive, unavailable, and unavailable. The mock checks that the synthetic token reached its Authorization header. A 200 response with unrelated content does not count as active; a rate limit does not count as inactive.

The example loads a YAML rule, scans, and passes the full result set to Validator::validate_findings. The validator binds captures and supporting credentials, renders templates with Kingfisher's filters, dispatches the rule's validator family, and returns ValidatedFinding values. Each contains the original finding, a ValidationOutcome, an optional credential-free ValidationReason, and an optional HTTP status. Call into_redacted() after validation, then emit the metadata your application needs. Redacting before validation yields Skipped with RedactedInput rather than sending [REDACTED] to a provider.

For your own service, replace the mock and synthetic detector with the built-in catalog or your private rules. The validate_file example shows a complete application using the built-ins. Provider requests happen only when you call the validator; scanning alone remains network-free.

Use the validation builder

use std::time::Duration;
use kingfisher_scanner::Validator;

let validator = Validator::builder()
    .timeout(Duration::from_secs(10))
    .concurrency(4)
    .max_response_bytes(1024 * 1024)
    .build()?;
let results = validator.validate_findings(findings).await;
for result in results {
    if result.finding.rule().visible() {
        println!("{}: {:?}", result.finding.rule_id, result.outcome);
    }
}

Reuse the scanner and validator in application state. validate_finding(&finding) checks a standalone finding. Use validate_finding_with_context(&finding, &findings) for one finding with supporting credentials, or validate_findings(findings) for a whole scan. Batch results preserve input order and include invisible helpers; filter those only after validation. Clones share the client's connection pool and the concurrency limit. Dropping the validation future cancels its pending orchestration; no detached validation tasks are created by the batch runner.

Builder option Default / behavior
concurrency(n) 8 checks across the validator and its clones; zero is rejected
timeout(duration) 10 seconds per started check, including waiting for a shared permit, DNS, and multi-step requests; zero disables Kingfisher timeouts
retries(n) Zero YAML HTTP retries by default; retries share the total deadline and rebuild multipart bodies
max_response_bytes(n) 1 MiB for YAML HTTP responses; oversized bodies yield Unavailable; zero disables YAML HTTP, Betterleaks, and gRPC body caps
client(reqwest_client) Default client verifies TLS and disables redirects; injected clients must supply their own TLS, proxy, and no-redirect policy
variable(name, value) Trusted template/Betterleaks environment variable, for example GITHUB_API_BASE_URL; capture/component values take precedence
allow_internal_ips(true) Opt in for trusted local/private services; false by default

The HTTP client is used by YAML HTTP, Betterleaks HTTP requests, Raw HTTP flows, and Coinbase. SDK, database, and gRPC helpers own their transports. All dispatches share the outer deadline and concurrency bound. timeout(Duration::ZERO) also disables Kingfisher protocol-helper timeouts. Injected HTTP clients retain their own settings. Positive body-size settings apply to YAML HTTP; Betterleaks and gRPC retain their 1 MiB defaults unless max_response_bytes(0) disables body caps. These policies are scoped to each validation future and do not change concurrent validators. The resolver checks are not a DNS-pinning guarantee or a substitute for application network policy.

Shared execution with the CLI

Validator, CLI scans, and kingfisher validate all call the same validation::ValidationEngine. It dispatches rules, executes protocols, and returns explicit outcomes. CLI candidate selection, rate limiting, and scan caches remain outside the engine. Provider responses cannot silently turn an inconclusive outcome into an inactive credential through HTTP-status inference.

Most applications should use Validator. Advanced integrations that already resolve rule variables can use ValidationEngine::new(&client, &parser).validate(&rule, &globals). Supply uppercase scalar variables including TOKEN, a parser configured with kingfisher_rules::register_liquid_filters, and a client with redirects disabled. Configure its deadline, retries, private-network policy, and typed-validator TLS policy explicitly. This lower-level API does not associate findings or bound concurrency. ValidationResult exposes outcome, reason, HTTP status, and a potentially sensitive response_body; Debug omits that body and the type has no Serialize implementation. The high-level ValidatedFinding omits provider responses entirely.

YAML HTTP checks require a nonempty status, word, or header matcher. Empty matchers and JSON-validity checks alone cannot prove authentication. Multipart file parts send rendered content as bytes; they do not read paths from the host filesystem.

Supporting credentials and validation outcomes

Pass the complete raw results from one scan input into each batch. Components must have the same blob ID and encoding as the primary finding, and satisfy the rule's within constraint when present. Distinct competing values yield Skipped with AmbiguousDependency; required absent values yield MissingDependency. Repeated occurrences of the same component value are accepted. The API deliberately does not try multiple credential combinations, even for verify_candidates rules. For ambiguous inputs, scan a narrower credential block with its needed context. Do not combine unrelated file results into one batch.

The dispatcher supports YAML HTTP (including inline multipart), Betterleaks expressions and multi-step flows, and the enabled typed/Raw/gRPC validator families. validation exposes the high-level API and all supported families, including offline key-material checks. Rules without a validator produce NotAttempted; Assumed stays distinct from live proof; non-authoritative rules remain NotAttempted with a NonAuthoritative reason. The API does not revoke credentials or perform access mapping.

HTTP matchers must encode reliable authentication evidence; negative matchers retain the YAML rule semantics. A matching response reports activity; HTTP 401 is rejection, while unexpected bodies, 403, rate limits, redirects, server errors, and network failures remain inconclusive unless a provider-specific validator has stronger evidence. Legacy protocol helpers that return an ambiguous false are conservatively reported as Unavailable. Local Ethereum derivation remains LocallyDerived, not VerifiedActive. Inspect outcomes directly instead of inferring them from the optional HTTP status.

ValidatedFinding retains raw credentials until explicitly redacted. Its debug output omits credentials and it has no automatic serialization implementation. Serialize selected metadata or its redacted finding. Error reasons omit raw provider bodies, URLs, and capture values.

Verify an integration before shipping it

Run copied examples with synthetic inputs before connecting real providers. Cover a clean input, a detected token, unreadable input, and provider rejection, throttling, and unexpected-success bodies. Preserve detection results when validation is inconclusive. In this repository, run the portable example checks with:

python3 -m unittest discover -s scripts/tests -p test_library_examples.py -v

Loading and Compiling Rules

Compile once and share the database. Preserve catalog metadata with RulesDatabase::from_rule_collection; converting the loaded collection into a Vec<Rule> loses its database-level source prefilter.

For custom files, add kingfisher-rules = "1.2.0" and use:

use kingfisher_rules::{Confidence, Rules, RulesDatabase};

let rules = Rules::from_paths(["rules/company.toml"], Confidence::Low)?;
let database = RulesDatabase::from_rule_collection(rules)?;

Both the Kingfisher rule format (.yml/.yaml) and Betterleaks TOML (.toml) are fully supported by Rules::from_paths, including directories containing both formats. Betterleaks TOML rules receive the custom. namespace. See rule authoring. Missing files, malformed rules, and compilation failures return errors. The built-in catalog is embedded; loading it performs no rule downloads.

Programmatic private rules can use a constructor instead of a complete schema literal:

use kingfisher_scanner::{Rule, RuleSyntax, RulesDatabase};

let mut syntax = RuleSyntax::new(
    "acme.service-token", "Acme service token", r"(acme_[a-z0-9]{16})",
);
syntax.min_entropy = 2.0;
let database = RulesDatabase::from_rules(vec![Rule::new(syntax)])?;

Construction supplies medium confidence, visibility, and no validators, filters, or entropy threshold. Compilation validates the pattern. Use one capture for the reported secret; Vectorscan patterns cannot use lookaround.

Scanner Configuration

use kingfisher_scanner::ScannerConfig;

let config = ScannerConfig {
    redact_secrets: true,
    enable_dedup: false,
    ..Default::default()
};

Pass it to Scanner::with_config(Arc::clone(&database), config).

Setting Default Contract
enable_base64_decoding true Scan one Base64 decoding layer in addition to ordinary content
enable_dedup false Opt in to suppressing previously reported content at the same source path
min_entropy_override None Use each rule's threshold; an override applies to every rule
redact_secrets false Replace returned secret and capture values with [REDACTED]

Deduplication is scoped to one scanner and cleared by reset_dedup() or dropping that scanner. Concurrent first scans may both report findings, and in-flight scans may commit entries after a reset. Finish them before resetting a batch. Source paths are compared as supplied, without canonicalization. The cache retains entries for scans that returned findings; it grows until reset. Leave it disabled for request-oriented services or when every file occurrence must be reported.

Scanning Methods

For reuse across processes, add kingfisher-rules = "1.2.0" to your dependencies and load the collection with kingfisher_rules::RulesDatabase::from_rule_collection_with_cache and kingfisher_rules::RuleCacheConfig::from_dir_or_env(None) (or RuleCacheConfig::new(path)). This is the CLI/Python compiled content-database cache and honors KF_RULE_CACHE_DIR. Entries require the same native engine build and exact binding versions, plus compatible architecture, pointer width and endianness. Cache read failures and native CPU/version rejection fall back to compilation; writes are best effort. Confirmation regexes and path/finding filters still initialize per database. Existing uncached Rust constructors remain uncached. RulesDatabase::cache_status() reports Loaded, Stored or Bypassed, so applications can require persistence when prewarming. Cache directories must be trusted; unsafe ownership, permissions or symlinks bypass caching. Without a per-user directory, automatic cache configuration disables disk use. SHA-256 checks payload integrity before native deserialization.

All scanning methods return anyhow::Result<Vec<Finding>>; propagate or explicitly handle errors. A failed scan is not an empty successful scan.

An error evaluating a Betterleaks filter, including a malformed custom betterleaks_filter, fails the entire scan call for that input. No partial findings are returned, including findings from other rules. This applies to ordinary and Base64-decoded content: propagating the error prevents a failed filter from silently keeping a finding that should have been filtered. Fix the rule and retry the input.

Method Input and behavior
scan_bytes(&[u8]) Borrows input except when encoding normalization requires allocation; no source path
scan_file(path) Opens a file, preserving its path for path-aware rules; returns I/O errors
scan_blob(&Blob) Reuses an existing blob; no source path
scan_blob_at_path(&Blob, &str) Reuses a blob with an explicit path for filters and path predicates

A scan performs detection and filtering, including entropy, catalog filters, and component requirements. It does not run validators, revoke credentials, traverse remote repositories or reproduce every CLI pipeline stage. Optional git, archives, extraction and context features provide explicit local enumeration, transforms and detection policies; see the sections below. Enabling a validation feature does not change this boundary.

Working with Findings

Findings own their strings and remain usable after the scan input is dropped. They hold an Arc<Rule> for rule metadata. Results may include invisible component helpers; filter with finding.rule().visible() when displaying user-facing detections. Use rule_id, rule_name, confidence, entropy, location, and is_base64_encoded for reporting. secret and captures contain credentials unless redaction is enabled.

Redaction happens after component matching and fingerprint computation. It covers all returned capture values as well as the primary secret. It is output redaction, not secure erasure of input, rules, or temporary memory. Findings can still contain sensitive metadata such as locations and fingerprints.

Offsets are zero-based, end-exclusive byte offsets; lines are one-based and columns are zero-based byte columns. UTF-16/32 input is normalized to UTF-8, and reported positions refer to that normalized content. For Base64 matches, the location covers the encoded region in the outer content. It is not a decoded-secret offset. Finding order is unspecified. Numeric fingerprints and native database caches are implementation details, not stable persisted identifiers or storage formats.

Parallel Scanning

Scanner, RulesDatabase, and ScannerPool are Send + Sync. Compile once and share with Arc; native scratch space is allocated per worker thread. Default scans are independent, including repeated scans of the same content.

The API is synchronous and CPU-bound. Async applications should use blocking workers or their own bounded thread pool. The library does not initialize a Tokio runtime, logging subscriber, process allocator, or global executor for scanning.

ScannerPool is a lower-level API for native scanners. Prefer try_with, which returns allocation and reentrant-borrow errors. A callback must not recursively borrow the same pool on the same thread. with panics for those errors; it does not permit overlapping mutable borrows.

The application matcher, embeddable scanner, source-path prefilter, and finding-filter helpers share one pool implementation in the rules crate. The existing kingfisher::scanner_pool::ScannerPool and kingfisher_scanner::ScannerPool paths refer to that same type. The pool owns its database and drops thread-local scanners first. Its callback cannot return a scanner borrowing that database. A fallible callback produces a nested Result; propagate both layers with pool.try_with(|scanner| scanner.scan(bytes, callback))??.

Credential Validation (Optional)

No validator features are enabled by default:

[dependencies]
kingfisher-scanner = { version = "1.3.0", features = ["validation"] }

validation enables all supported validators and revocation: HTTP, gRPC, Betterleaks expressions, Raw, Ethereum, AWS, Azure, Coinbase, GCP, JWT, and databases. Without it, the crate provides scanning only. The old validation-* feature names remain compatibility aliases; each enables the complete validation feature.

Validate findings

Enable validation for the full built-in validator catalog. Validator makes provider requests only when called. Use it on findings from one input and redact each result before logging or serializing it:

[dependencies]
anyhow = "1"
kingfisher-scanner = { version = "1.3.0", features = ["validation"] }
tokio = { version = "1", features = ["macros", "rt-multi-thread"] }
use kingfisher_scanner::{Finding, Validator};

async fn validate(findings: Vec<Finding>) -> anyhow::Result<()> {
    let validator = Validator::builder().concurrency(4).build()?;
    let report: Vec<_> = validator
        .validate_findings(findings)
        .await
        .into_iter()
        .filter(|result| result.finding.rule().visible())
        .map(|result| {
            let result = result.into_redacted();
            (result.finding.rule_id, result.outcome)
        })
        .collect();
    println!("{report:?}");
    Ok(())
}

Validation is a separate, explicit step. It may send candidate credentials to provider APIs according to the selected rules. Use it only for credentials and accounts you are authorized to check.

Prefer Validator for rule dispatch and result handling; low-level protocol helpers remain available. Call validation explicitly. Embedding applications own clients, runtime, timeouts, and concurrency. Some validator helpers expose process-wide configuration; isolate or coordinate such configuration instead of changing it per concurrent request. ValidationOutcome distinguishes verified activity from assumptions and local cryptographic derivation. Do not equate an actionable finding with a live credential. Provider behavior and network availability are outside the Rust API contract.

Revoke a credential

kingfisher-scanner::Revoker dispatches built-in or custom rules to HTTP, multi-step HTTP, AWS, and GCP revocation. Enable validation for all of them. No dependency on the CLI crate is needed.

[dependencies]
anyhow = "1"
kingfisher-scanner = { version = "1.3.0", features = ["validation"] }
tokio = { version = "1", features = ["macros", "rt-multi-thread"] }
use std::collections::BTreeMap;
use kingfisher_scanner::{Revoker, Rule, get_builtin_rules};

#[tokio::main]
async fn main() -> anyhow::Result<()> {
    let rules = get_builtin_rules(None)?;
    let syntax = rules.rules.get("betterleaks.github-pat")
        .ok_or_else(|| anyhow::anyhow!("Rule not found"))?;
    let rule = Rule::new(syntax.clone());
    let secret = std::env::var("TOKEN_TO_REVOKE")?;
    let variables = BTreeMap::from([
        ("GITHUB_API_BASE_URL".into(), "https://api.github.com".into()),
    ]);
    let result = Revoker::new()?.revoke(&rule, &secret, &variables).await?;
    println!("{}: revoked={}", result.rule_id, result.revoked);
    Ok(())
}

This example makes a real revocation request. Scanning and validation never invoke revocation automatically. Select and authorize the credential before calling it. Supply companion variables and endpoint overrides as case-sensitive map entries (e.g. AKID, KEY_ID, or the endpoint variable used by your rule). TOKEN is reserved for the secret argument. The runner does not read environment variables or apply CLI endpoint defaults.

The default client uses strict TLS and disables redirects. Revoker::with_client accepts an application-owned client for HTTP rules; AWS and GCP use their own transports. HTTP rules can reach internal addresses, so callers own endpoint policy. HTTP rule requests are not retried; AWS retains its provider-specific retry policy. The default total deadline is 10 seconds; change it with Revoker::timeout. Revoker::error_category classifies failures without returning endpoint URLs or provider bodies, for applications that need safe operational logs. A timed-out request may already have revoked the credential. Results include the rule identity, success flag, optional HTTP status, and response message. Provider responses and errors may contain sensitive data. Missing revocation configuration, execution failures (including transient HTTP statuses such as 429 or 503) return errors; a response evaluated by the rule that fails its matcher returns revoked: false.

See the runnable example.

Build and Deployment

The native Vectorscan backend supports the repository's macOS, Linux, and Windows builds. Windows uses GNU/LLVM MinGW targets; MSVC is not supported by the standard native backend. Cargo builds may download its platform archive. For offline builds, cache Cargo dependencies and native prerequisites, configure VECTORSCAN_PREBUILT_DIR or HYPERSCAN_ROOT, and set VECTORSCAN_OFFLINE=1. See publishing for extracted-package verification. --offline alone does not sandbox build scripts.

The rules and scanner crates configure docs.rs to compile vendored Vectorscan source for Linux x64, avoiding archive downloads in its network-blocked build environment.

Large files can be memory-mapped by Blob::from_file; the application must ensure mapped files are not modified or truncated while in use. Set application-level input limits and bound worker counts for untrusted or large workloads.

The full application's library is available separately:

kingfisher = { package = "kingfisher-bin", version = "2.11.1" }

It carries CLI dependencies and all validator features. The stable embedding contract described here applies to the three 1.x library crates; prefer those for detection, validation, and revocation integration.

API Stability

Scanner and Python SDK 1.3.0 add opt-in CLI detection policies, shared SQLite/bytecode extraction and richer Python Git scopes/provenance. Existing scanner configuration, method signatures and defaults remain unchanged.

Rules 1.1.0 adds endpoint-confirmation helpers and reduces filter input copying. Scanner 1.2.0 adds optional cooperative scan deadlines/cancellation and improves candidate confirmation and span tracking while retaining existing scan methods and public helper signatures. Python SDK 1.1.0 adds per-call scan controls and Rules.detail() inspection. Core 1.0.3 avoids duplicate hashing of borrowed content while preserving its public API; the CLI is versioned separately.

Scanner 1.1.0 introduced Revoker and made validation all-or-nothing. Existing provider feature names remain aliases, so existing manifests keep working while compiling all validators. Python SDK 1.0.2 delegated revocation to the shared Rust runner.

The first published 1.0.0 establishes the stable contract. Within 1.x:

  • Existing reachable public Rust items, signatures, fields, trait implementations, and feature names remain compatible. Hidden-but-public items are not exempt.
  • Removing public items, changing field types, adding fields to exhaustive structs, adding variants to exhaustive enums, or incompatible public dependency type changes requires a major release. Deprecation can precede removal but does not permit it in 1.x.
  • Documented defaults, error propagation, location semantics, and redaction behavior are behavioral contracts covered by regression tests.
  • Existing documented serialized field names and enum spellings are retained. New optional data may be added where compatible; consumers should accept unknown fields.
  • Rust 1.99 is the minimum supported compiler for this release. Dependency updates must preserve that minimum and pass feature and consumer checks.

Catalog content, match counts, provider responses, diagnostics, finding order, fingerprints, and native cache bytes can change without a major API release. Fixing incorrect detections is not a promise to preserve previous findings. Pin exact crate versions and retain Cargo.lock when reproducible detector results matter, and review catalog provenance when upgrading.

CI runs public consumer regression tests, README doctests, default/all-validator builds, and cargo-semver-checks against the pull request's base revision. After publication, release checks must also compare against the last published compatible library version (see publishing). Static API checks do not prove all behavior; review and behavioral tests remain required.

Migrating from the 0.1 API

  • scan_bytes now returns a Result; use ? or handle the error.
  • Repeated calls return findings by default; explicitly enable cross-call deduplication when wanted. Deduplication now includes the source path.
  • Remove unused language_hint and max_base64_depth fields from ScannerConfig. Default Base64 detection scans one decoding layer; the opt-in detection policy supports bounded nested decoding.
  • Redaction now uses [REDACTED] for secrets and captures, and preserves the unredacted fingerprint. It no longer reveals a secret prefix.
  • SerializableCapture owns its name and value strings. The leaking intern helper is removed; use owned strings. raw_value() borrows from the capture.

See Also

Shared archive extraction

The optional kingfisher-scanner feature archives exposes the CLI/Python archive helpers in archive::decompress and their budgets in archive::limits. Scanner methods continue to scan their supplied content without enumeration or extraction. Extract explicitly and pass member bytes and logical paths to scan_blob_at_path. Shared extraction is best effort and can skip entries or truncate content at format limits. Python users can compose native filesystem and reachable Git history iterators with archive expansion through the Python input APIs.

For strict byte budgets, use archive::decompress::decompress_file_with_strict_single_stream_cap_and_limits with ResourceLimits::default(). It fails on stream-cap exhaustion before parsing a partially decoded TAR. decompress_file_with_budget also accepts a strict entry budget and ScanControl, returning an ArchiveExpansion with inspected entry usage so nested callers can debit one aggregate root budget. extract_zip_archive_in_memory_with_budget performs bounded ZIP extraction without temporary storage and reports inspected entry usage. Existing best-effort entry points keep their truncation behavior. TAR members use independent physical staging paths while preserving their logical names, including repeated names.

With no output directory, TAR/ZIP extraction returns CompressedContent::Archive with member bytes in memory. To retain disk-backed members, use decompress_file_to_temp_with_limits and keep its returned TempDir alive. ArchiveExpansion::inspected_entries counts TAR/ZIP directories and skipped members, ASAR indexed files, and HWP streams.

Owned directories from decompress_file_to_temp use Unix mode 0700; standalone decoded files created without an output directory use mode 0600 and require caller cleanup. The umask may restrict these modes further. Windows inherits the temporary parent's DACL. Choose a protected temporary parent on every platform; caller-supplied output directories retain their permissions and must also be protected. Cleanup is not secure erasure and can leave remnants after a crash; use encrypted or memory-backed storage when needed.

Optional CLI detection policies and content extraction

Enable context for opt-in CLI matching, bounded Base64 decoding, inline-ignore and HTML/CSS parser policies. Existing ScannerConfig and scanner methods retain their signatures and behavior. Use context::DetectionOptions with scan_blob_at_path_with_options for unlimited controls, or scan_blob_at_path_with_options_and_control for a deadline/cancellation:

# use std::sync::Arc;
# use kingfisher_scanner::{Blob, RulesDatabase, Scanner, ScanControl, get_builtin_rules};
# #[cfg(feature = "context")]
# fn main() -> anyhow::Result<()> {
# let scanner = Scanner::new(Arc::new(RulesDatabase::from_rule_collection(get_builtin_rules(None)?)?));
use kingfisher_scanner::context::DetectionOptions;
let blob = Blob::from_bytes(b"ordinary configuration".to_vec());
let findings = scanner.scan_blob_at_path_with_options_and_control(
    &blob, "config.html", &DetectionOptions::default(), &ScanControl::default(),
)?;
# Ok(())
# }
# #[cfg(not(feature = "context"))]
# fn main() {}

The options entry point enables CLI matching by default: a 4 KiB initial confirmation window (widened when needed), full-match component windows, per-rule secret containment suppression, and overlapping Betterleaks credential-URI fallback suppression. cli_match_semantics = false retains the SDK's 64 KiB confirmation alignment and secret-based component windows. Full-match metadata stays private; reported locations and the public Finding shape are unchanged.

Base64 decoding defaults to two layers and skips the Base64 pass when the original input exceeds 64 MiB. Raw matching still runs above that limit. Set base64_max_depth independently (zero disables decoding), and use base64_max_input_bytes = None to remove the input cap. ScannerConfig's enable_base64_decoding = false disables decoding regardless of these options. Nested findings keep the outer encoded region's offsets. Existing scanner methods retain their one-layer, uncapped Base64 behavior. Containment checks are separate for raw input and each decoded buffer, preserving separately encoded sibling secrets even when they share the outer region's offsets.

Inline-ignore and containment filtering run before markup verification, followed by component requirements, URI fallback suppression, catalog deduplication and redaction. Interrupted calls return no partial findings and do not commit dedup state. The dedup key is the blob ID and path; it does not include detection options. With deduplication enabled, use a separate scanner per detection policy. The markup gate uses shared language inference, bypasses self-identifying/Base64 candidates and retains candidates above 2 MiB or with invalid UTF-8 secrets. Set inline_ignores = false or markup_context = false independently. The parser, inline-ignore, index and confirmation helpers are isolated behind the explicitly unstable __cli-internals feature for CLI orchestration. Embedders should use the supported scanner entry points.

The optional extraction feature enables archives and exposes extraction::sqlite::extract_sqlite_contents[_with_limits] and extraction::pyc::extract_pyc_strings[_with_limits]. SQLite extraction opens the supplied file read-only and emits SQL per user table; bytecode parsing extracts marshal strings without executing Python. These low-level helpers use the CLI's best-effort limits and may skip/truncate content. Strict counterparts extract_sqlite_contents_with_budget and extract_pyc_strings_from_bytes_with_budget enforce output budgets while extracting and propagate ExtractionLimitExceeded without partial text. SQLite uses read-only/defensive connections, disables trusted schema expressions, bounds native SQLite allocations and checks controls through a VM progress hook. The bytecode helper borrows input bytes and checks controls during parsing. Extraction is explicit and separate from scanning; findings refer to extracted content. The Python expand_content() adapter stages a separate copy and adds per-input budgets and cooperative controls. Choose a protected temp_dir parent on every platform. SDK staging directories use owner-only Unix mode 0700 (the umask may restrict it further); Windows inherits the parent DACL. See the SDK guide.

Optional local Git enumeration

Enable git to use git::GitInputs, GitScope, GitOptions and GitEvent. Prepare descriptors once, then scan each GitEvent::Input payload with its repository-relative path. GitEvent::Skipped reports explicit oversized or missing-blob skips; decide whether incomplete coverage is acceptable in the embedding. The Git example shows scoped selection, budgets, controls and redacted reporting.

History traverses every selected merge parent and compares each commit to its first parent. Identical subtrees are skipped; commit metadata is shared with Arc. Payloads load lazily, while descriptors and ancestry prepare eagerly. max_commits bounds distinct visited commits (including exclusion ancestry), max_inputs bounds distinct raw-path/blob descriptors, and max_blob_size checks object headers before reading payloads. Preparation budget exhaustion fails rather than truncating coverage. The ScanControl deadline covers iterator lifetime, including time spent by the consumer. Limits are optional and unlimited by default; configure them for service workloads.

Snapshots, net diffs, staged index content, time ranges, branch roots and stored unreachable objects are explicit GitScope choices. Native acquisition is read-only and offline. Missing blobs fail unless skip_missing_blobs explicitly requests skipped events; missing trees/commits still fail. Partial clones are never fetched. Non-UTF-8 paths retain raw_path bytes with a lossy display name. Default ancestor discovery is preserved; set discover = false for explicit repository roots. Revisions use local Git grammar, including reflog expressions. See Python Git scope semantics for provenance and selection details shared by both SDKs.