XpectraFlow docs

System specs

What one on-prem server absorbs — 10 kHz across 100 channels, how much of it you keep, and the arithmetic behind both numbers.

The reference workload on this page is 10 kHz across 100 channels — one million values a second, into a single dataset, on one on-prem server.

The short answer:

One server at the recommended spec ingests 10 kHz across 100 channels continuously, and stores about 310 hours of it.

That is twelve hours of recording a week for six months, or eighteen hours a week for four — without adding hardware, changing the schema, or migrating anything.

The rest of this page is the arithmetic, so you can check it against your own duty cycle rather than take the claim on faith.

The server

One machine. Sized for on-prem, where you buy the box once and want it to still be the right box in two years.

MinimumRecommended
CPU8 cores16 cores
RAM16 GB32 GB
System disk100 GB SSD250 GB SSD
Data disk1 TB NVMe SSD, dedicated2 TB NVMe SSD, dedicated
Data disk write throughput≥ 250 MB/s sustained≥ 500 MB/s sustained
Network1 GbE1 GbE
OS64-bit Linux64-bit Linux

Two things about that table are worth saying out loud:

  • The data disk is separate, and it is the component that matters. Telemetry is written continuously and read in bulk; keep it on its own NVMe device rather than sharing the system disk. Its size is what decides how much history you keep, and nothing else on this page changes that.
  • Everything else runs on that one machine. The software layers are installed and tuned during deployment; you do not size them separately, and they are not your problem to operate.

Against that hardware, the reference workload is a small fraction of the box:

At 10 kHz × 100 channelsRecommended specUsed
Network64 Mbit/s1 Gbit/s≈ 6%
Disk write≈ 20 MB/s500 MB/s≈ 4%
CoresIngest is a single writer per dataset16Leaves the rest for queries

The headroom is deliberate. A server that runs at 90% of its disk throughput on the day it is installed is a server that fails the first time somebody adds a second test stand.

What it absorbs

Figures are for 10 kHz × 100 channels, one dataset. Every line is linear in rate × channels — divide or multiply freely.

Per second
Values1,000,000
Rows10,000
Over the wire, streaming≈ 8.0 MB/s (64 Mbit/s)
Over the wire, JSON batches≈ 13 MB/s
Written to disk, before compression≈ 9 MB/s

Each row costs about 900 bytes on disk — 800 for the hundred float64 readings, the rest for the timestamp, the dataset reference and the index entry that makes it findable. At 10,000 rows a second that is ≈ 32 GB per hour of recording, before compression.

That number, not throughput, is what decides how long you go without touching anything.

How much it holds

Storage is a function of hours actually recorded, not calendar time:

Per recorded hourSize
Rows36,000,000
Raw≈ 32 GB
Stored, conservative (5× compression)≈ 6.4 GB
Stored, typical for smooth analog telemetry (10×)≈ 3.2 GB

Telemetry is compressed automatically once an hour of data has closed, so the steady-state figure is the compressed one. At the conservative 5×:

Data diskRecorded hours heldUnbroken single session
1 TB≈ 156 hours≈ 31 hours
2 TB≈ 310 hours≈ 62 hours

The second column is the peak case: compression only helps after data has settled, so one continuous session writes at the raw 32 GB/hour until it does. Even so, a 2 TB disk swallows a 62-hour unbroken run of 10 kHz × 100 channels. No realistic test campaign gets near it.

Read the first column against your duty cycle:

Recorded per weekHeld on 1 TBHeld on 2 TB
2 hours18 months3 years
6 hours6 months12 months
12 hours3 months6 months
20 hours8 weeks15 weeks

Twelve hours of 10 kHz × 100-channel recording every week, for six months, on the recommended 2 TB disk, with nothing touched. At 10× compression — common for smooth analog telemetry — double every figure in the table.

When you do eventually reach the end of the disk, none of the answers is a migration:

Archive. Completed datasets move to long-term storage on a schedule, and stay restorable. This keeps the working disk approximately flat however long the campaign runs. Already built — it is a setting, not a project.

Grow the disk. Adding capacity does not move or rewrite existing data.

Set a retention window. Older data is dropped automatically past a cutoff. Off by default, because archiving is meant to own cold storage — turn it on only once archiving is confirmed working.

100 channels is not a special case

Nothing widens when you go from 8 channels to 100. Each dataset gets its own storage table at registration, with one column per declared channel:

time         timestamp
dataset_id   uuid
ch_0         float64
…
ch_99        float64
ValueHeadroom at 100 channels
Columns in the table102The platform's ceiling is 1600. You are at 6%
Channels per registration4096Above what the column ceiling permits, so it is not the binding constraint
Adding channels laterAdditiveNo table rewrite, no downtime, no re-upload

So "does 100 channels need a migration" has a flat answer: no, because the storage shape is created per dataset from the channel list you send. There is no shared telemetry table whose schema anyone has to change.

Measured against projected

Be clear about which parts are proven:

Validated (2026-07-16)Reference workload
Rows/s20,00010,000
Channels8100
Wire1.3 MB/s8.0 MB/s
ResultZero loss, zero duplicates, publish p99 23 msProjected from the same pipeline

The pipeline has been run at twice the reference row rate with no loss and no duplicates. Rows per second is what the writer's per-batch overhead scales with — the batching, the flush timer, the duplicate check — and 10 kHz sits comfortably inside proven territory.

The untested axis is width: 12.5× the bytes per row through the same bulk write path. Bytes per second is the axis to watch, and at ≈ 20 MB/s it is using about 4% of the recommended disk. There is room, but it is a projection, not a measurement.

The validated runs were on a developer workstation. Re-run the load test at 10 kHz × 100 channels on the actual delivered hardware before go-live. Correctness is proven; your box's ceiling is not.

Burst tolerance

Producers do not write to the disk directly. A durable buffer sits in between, so a slow disk, a restart or a maintenance window costs latency rather than data:

At the reference rate
Buffered backlog before anything is refused≈ 8 minutes of full-rate data
Unacknowledged backlog before back-pressure≈ 3 minutes
Write batchingCommitted every 200 ms, or every 50,000 rows

Frames are acknowledged only after the data is committed, so an interrupted consumer replays rather than loses. Replayed rows collapse onto the rows they already wrote — a duplicate is never a duplicate row.

Keep frames at or below 1000 rows at 100 channels. A frame carries channels × rows × 8 bytes — 500 rows is 400 KB, 1000 rows is 800 KB — and a single frame is capped at 1 MB. The default of 500 rows/frame leaves plenty of margin; 2000 rows at 100 channels is refused, visibly, but only once you are at rate.

Where the JSON path caps out

The reference workload assumes streaming. Sustaining 10 kHz over HTTP append is a different proposition, and the honest arithmetic is:

At 100 channels
Rows per request (hard cap)5000
Requests/s needed for 10 kHz2
JSON per request≈ 6.5 MB
Internal write statements per request9

That last row is the one that bites. A request that wide is split into nine internal statements of 588 rows each, every value is written individually rather than as a bulk stream, and every row carries its own timestamp instead of sharing one base timestamp and a sample period across a whole frame.

HTTP append remains right for backfills, slow sensors and anything a person triggers, up to a few hundred rows a second. Above that, streaming is not an optimisation, it is the supported path — and switching does not change the data model: same dataset, same ch_N columns, same channel list.

What would actually force a change

Three triggers, in the order you are likely to meet them:

Query load is the wildcard none of the above covers: reads on recent data share the same machine as the writer. Bulk reads already downsample server-side for exactly this reason, but a dashboard polling a live 10 kHz dataset is spending from the same budget as the ingest — which is the other reason the recommended core count is what it is.

Checking it yourself

A support bundle reports every dataset's storage table with its estimated row count and total size on disk. Two bundles a week apart give you your real growth rate — the one number this page cannot predict for you.

On this page