Oblive Docs

Analytical storage and execution

Stable sources, exact amounts, immutable definitions, and the bounded DuckDB runtime.

PostgreSQL coordinates source collections and indexes visible segments. Immutable Parquet objects retain analytical history. DuckDB CLI 1.5.5 filters and deduplicates pinned files inside the backend; it is not another database service. Deterministic adapters run in the observation loop. The dashboard, chat and task CLI use one shared query service.

Install the engine

Run bun run insights:install before backend analytics tests or local host execution. Backend images install it during their build. The installer supports macOS and Linux on ARM64 and AMD64 and checks the official release asset SHA-256 values before decompressing or executing the binary. The absolute runtime path is node_modules/.cache/oblive/duckdb/1.5.5/duckdb under the repository root.

The query process uses an argument vector, no shell, no startup file, and no inherited credentials. Only server-owned SQL templates run. Exact input files are staged in a private temporary directory and checked against their indexed lengths and hashes. External access and extension installation/loading are disabled, allowed paths list only these files, and configuration is locked before the query.

Each backend admits at most two processes. Each process has two threads, 256 MiB of memory, at most 128 MiB of scratch, 256 MiB of staged input, a 30-second deadline, and a 5 MiB public output limit. Internal corrected facts stream through a bounded 64 MiB/250,000-row reader before exact aggregation. The reader reuses its validation schema and yields every 128 records so cancellation timers and unrelated HTTP requests remain responsive during reduction. Query-local indexes avoid searching or sorting the full receipt history for every fact; calendar and coverage results are reused within the query. Cancellation kills the process and removes its temporary directory. There is no persistent file cache. Oversized results return insight_query_narrow_required, never a silently truncated result.

Durable meaning

Source identity consists of provider, upstream account and resource, dataset, and environment. Credential rotation changes the collection generation, not source identity. Generation also includes enforced scope, authorization epoch, adapter version, and schema version. Collection checkpoints have separate refresh, reconciliation and backfill cursors, typed leases, coverage and batch receipts. Discovered resourceName is display metadata outside source identity; repository renames cannot create new history. The metrics catalog marks selected working defaults as preferred. Explorer defaults and collection priority prefer these resources while other authorized sources remain available.

Facts contain stable analytical keys, typed minimal content, effective/collection times and optional upstream sequence or timestamp revisions. A deletion requires explicit evidence or a completed exhaustive snapshot. Message bodies and contact profiles are not accepted analytical payloads. Segments are addressed by source, dataset, logical event/report month and content hash; upload month does not determine history.

Active metric definitions are code-owned and have immutable versions. Legacy custom definitions remain in storage for historical receipt parsing but are inactive and cannot execute. Reusing a version with different content fails; a PostgreSQL trigger also prevents direct semantic updates. Availability may change independently.

Amounts cross public and Parquet boundaries as validated decimal strings, with twenty integral and eighteen fractional digits. Shared arithmetic uses integer coefficients and half-even rounding at the 18-digit boundary. Ratios use numerator/denominator components. Median final arithmetic uses the middle samples; it must not use DuckDB’s floating-point division or decimal median implementation.

Currency and source timezone belong to series identity. Real report boundaries, including 23/25-hour DST days, stay intact. Complete zero, missing, partial, suppressed and unavailable coverage remain different outcomes. Duration samples carry stable sample identities. Twelve months is a backfill bound, not a history retention policy.

Collection publication

The analytical collector uses typed analytical_connector, analytical_source, and analytical_receipt records in existing system durability. They are backend-owned; generic durability mutations cannot change them and ordinary agent Context does not include their internal state.

The loop resolves one published catalog revision per cycle and executes at most four work units concurrently. It drains continuations for up to 45 seconds or 64 claims per cycle instead of waiting a full polling interval after every page batch. Claims serialize within a process to avoid competing for the same organization lock; provider requests remain concurrent and database leases coordinate replicas. It queries indexed due sources, prioritizing current refresh over reconciliation and backfill. Claims lock the integration before its organization. PostgreSQL permits two active connector leases per organization and one per connector. Collection units retain at most twenty pages, ten thousand normalized records, or 32 MiB, and stop after thirty seconds. Separate phase cursors retain hourly refresh, daily reconciliation, and historical backfill progress. Asynchronous report job IDs and poll times survive restarts. Provider Retry-After wins over the capped exponential backoff with jitter. Interrupted discovery becomes eligible again after its lease expires, even if the next daily scan was scheduled for a later time. The collector reclaims it through the normal authorization checks; active leases, provider retry delays and permission pauses remain effective. Reaching a work-unit deadline after retained progress is a continuation, not a provider outage. Metric availability uses the relevant dataset; an issue inventory cannot establish PR coverage and a workflow permission failure cannot make unrelated issue metrics unavailable. Past-due refresh times display as queued work. Missing connector pagination is unsupported_capability, not an instruction to change company Context or reconnect valid credentials.

Files and their collection manifest upload before the publication transaction. Publication rechecks credentials, enforced scope, authorization epoch, source identity, checkpoint version and lease token; expiry is also checked against the current PostgreSQL clock. It allocates the next organization revision while holding the short organization lock. Another source’s publication does not invalidate otherwise valid uploaded files. A committed batch receipt makes replay idempotent.

Incomplete multi-page reports retain immutable segment references in the corresponding phase checkpoint. Their segments and current projections from every retained page become visible together only when the report completes. Refresh can interrupt backfill without losing either phase’s retained pages. Current projections order by reporting period before collection time, and previous means a preceding distinct period. Scope failures pause the affected source until its connection changes; disconnect preserves historical objects.

The agent-authored insight setup protocol and generic JSON-pointer collector are removed. Resend and Instantly use deterministic adapters; lifecycle-only integrations retain connection health. Additional provider rollout is recorded in the milestone ledger. The UI distinguishes collection, backfill, reporting permission, required configuration, source retention and connection status.

All nineteen catalog providers now have a tested collection capability. Upstash requires explicit numeric/collection keys and never scans. MongoDB counts scoped collections using verified indexes; optional creation-date reports preserve UTC boundaries and retained-document limitations. Reporting selection changes fence collection separately from task authorization. The same reviewed fixed read providers serve these collectors and the agent integration gateway.

MongoDB new-document definitions retain the selected creation-date fields. Collection allocates a new immutable version when that mapping changes; credential rotation preserves the version. Publication deactivates older definitions without deleting their history. Agent queries and retained receipts recheck the original timestamp field against current grants, including when a newer metric uses a different permitted field. Owners can still inspect the earlier definition and results.

Shared queries and retained evidence

GET /organizations/:id/insights/metrics lists immutable definitions, supported sources, dimensions, grains, and actual coverage. POST /organizations/:id/insights/query selects exact metric versions and source IDs, a positive range of at most twelve calendar months, day/week/month grain, reviewed dimensions/filters, and an optional target currency. Limits are twenty selections, one hundred series, ten thousand points, five MiB, and thirty seconds. Oversized queries return a typed request to narrow. The current/history views and oblivectl insight query use this same service. Each source remains a separate series; primary integration preference never hides secondary-source data.

GET .../insights/current?view=catalog returns definitions, source metadata and collection status without reading Parquet or retaining query receipts. The Insights route loads this projection first. Collapsed connector groups never request values. An explicitly opened connector or metric uses a deferred browser query with a local loading/retry boundary, so the page and other groups stay usable. Direct links and reloads also render metadata first; their selected query starts after hydration. Router caching reuses results for thirty seconds; URL search keys separate connector, range, metric, source and version. Polling refreshes catalog status and only the open metric view; collapsed groups remain unloaded. The schedule list also uses the catalog projection for its integration-refresh status.

GET .../insights/current?view=overview projects all current metric/source combinations, optionally filtered by connectorKey. It batches up to twenty selections with matching calendar windows, runs at most two queries concurrently and shares a 25-second deadline. Invalid or oversized batches are split to isolate failing selections; storage and access failures are not retried. Unfinished selections remain explicit errors. Each successful batch keeps its own immutable receipt and publication revision; an overview is not an atomic snapshot across batches. The ordinary single metric view remains the default API behavior and provides the detailed history drill-down. The production frontend keeps idle sockets open for 120 seconds, beyond the bounded backend requests. CI exercises a response longer than Bun’s default ten-second socket timeout.

A repeatable-read transaction pins the analytical publication revision and eligible sources. The query reads only hash-verified visible files, deduplicates corrections, and respects completed report replacement windows. Matching-period distinct counts are never reconstructed by summing daily users. Flow totals use disjoint periods, ratios use aggregate components, gauges choose latest reporting periods, and duration medians use individual samples. Completed empty inventories can publish zero; partial inventories cannot. MRR history begins with observed complete subscription snapshots.

Every successful query uploads an immutable bounded result and commits its receipt after rechecking access. GET .../insights/receipts/:receiptId preserves cited results through later corrections and compaction. Every enabled profile in the active organization can read shared aggregates independently of live integration tool grants. List/query/receipt operations recheck connection readiness, verified credential account and enforced resource scope. Disabled sources are unavailable to new reads. Concurrent collection does not invalidate a query’s pinned revision.

Compaction and cleanup

The observation loop runs one bounded maintenance unit after collection. PostgreSQL leases one maintenance operation per organization, including disconnected sources. Compaction reads one source/dataset/content/month partition, at most 128 files, 64 MiB and 100,000 input rows. It uses the same hash-verified DuckDB reader and thirty-second deadline as queries. Trusted SQL checks the complete output’s row and serialized-byte bounds before the CLI writes a temporary result file. The backend reads that file incrementally, checks the bounds again, and removes it on completion or cancellation. This avoids native-process pipe short writes under load. It preserves deletion tombstones, original publication precedence, separate inventory snapshots, definition versions and completed report replacements. Uploaded replacements, input supersession, the next publication revision and a replay receipt commit together. Concurrent appends stay visible.

Cleanup examines 200 objects per page, resumes a durable cursor, and preserves all current files, unfinished paginated reports and receipt objects. Superseded files remain for at least twenty-four hours after compaction; unreferenced uploads remain for at least twenty-four hours after upload. Missing storage timestamps or uncertain database references prevent deletion. Query receipt commits recheck their deadline after acquiring the database lock, so an abandoned query cannot publish after its orphaned upload becomes eligible for cleanup. Cited query results survive input-file cleanup. Collected history is retained until organization deletion; twelve months bounds initial backfill, not retention. Bootstrap orphan cleanup is described in workspace creation.

GA4 and PostHog restart interrupted reports if the source timezone changes. The collection lease fences removal of the affected cursor and unfinished pages, with ordinary retry backoff; this does not revoke connection access. Missing X Ads async jobs restart the retained reporting window.

Optional currency conversion is a separate period_end_reference_v1 projection: a historical ECB reference rate for the source-local period end (or current day for a provisional period). The actual rate date is retained; weekends/holidays use the most recent available rate within seven days. Frankfurter’s official time-series API is restricted to the ECB provider and CSV decimal lexemes so rates never pass through binary floating point. Missing FX produces an explicit conversion gap and preserves the original amount. This is not accounting recognition or an assertion that every event transacted at the period-end rate.

Fixed definitions and explicit refresh

Adapters and the catalog own metric meaning. Definition comparison canonicalizes JSON keys, so a PostgreSQL JSONB round trip cannot create a new MongoDB metric version. When an earlier override differs from the code default, publication allocates one new immutable version and retires the old active definition. Collection generation schema version 2 resets unfinished reports and recollects fixed meanings while preserving prior segments and receipts. No custom formula execution, definition write endpoint, or recalculation worker remains.

POST /organizations/:id/insights/refresh and its agent counterpart accept an optional connectorKey. They validate the current organization/profile, connection readiness and source scope under the organization lock. An eligible source’s existing refresh deadline is moved to now; active leases, provider retry deadlines, paused sources, work already due, and collection completed in the last minute remain unchanged. A connector without discovered sources can bring discovery forward using its existing checkpoint. The ordinary collector claims this work; the request creates no action or new scheduler. oblivectl insight refresh [--connector <key>] exposes the same operation to agents.

The dashboard uses shadcn’s Recharts components. Coordinates are normalized with integer decimal arithmetic, preserving small differences between large values; tooltips retain exact original decimal strings. Missing observations remain line gaps. Neither chart scaling nor display rounding changes the stored metric value.