Building 1‑Minute K‑Lines From Tick Data: Handling Cross‑Minute Trades and Duplicate Packets

avatar
· Views 188


Building 1‑Minute K‑Lines From Tick Data: Handling Cross‑Minute Trades and Duplicate Packets



Introduction

In algorithmic and quantitative research, 1‑minute K‑lines (OHLCV candles) form one of the most widely‑used datasets for factor mining, strategy backtesting and simulation validation.

There are two common approaches to obtain 1‑minute candles. The first is to call ready‑made candle endpoints provided by market‑data platforms. The second approach is to pull raw tick‑by‑tick snapshots through an A‑Share API, then aggregate ticks into OHLCV on your local machine or backend service.

Many quant researchers choose the second method for greater data flexibility. Self‑aggregation lets you implement custom aggregation logic, apply specialized sampling rules, and perform self‑hosted data validation and secondary processing.

On paper, aggregation seems straightforward: collect all trades occurring within the same calendar minute, compute open, high, low and close prices, and sum total trading volume. This produces a complete 1‑minute candle.

However, once you work with real‑time streaming market data, network conditions, WebSocket reconnection events, and upstream feed‑behaviour introduce many edge‑cases rarely covered in basic tutorials. These issues do not trigger program crashes. They create silent data bias, which propagates directly into backtest outputs, distorting strategy performance metrics and causing misjudgement regarding factor effectiveness.

While building research‑oriented market‑data pipelines, I have repeatedly encountered mismatches between self‑aggregated candles and authoritative benchmark datasets. Post‑mortem debugging revealed that discrepancies rarely stem from flawed OHLCV arithmetic. Instead, root causes lie in three frequently‑overlooked areas: correct time attribution for each tick, handling out‑of‑order packets crossing minute boundaries, idempotent filtering for duplicate tick records, and deciding when a minute candle should be considered final.

This article walks through real‑world anomaly cases, explains underlying causes, and shares practical implementation patterns. It also discusses how these pre‑processing steps improve reliability for backtesting and model‑driven research. All code snippets are for demonstration purposes only and should be adapted to match your own research environment.



1. Real‑world Data‑anomaly Scenarios

In the early iteration of our pipeline, we implemented a naive but widespread pattern: we used the WebSocket packet arrival timestamp on our server to decide which minute bucket each incoming tick belonged to.

Consider realistic intraday A‑Share trade timestamps: 09:30:59.800, 09:30:59.950, 09:31:00.020

Based on the actual exchange matching time, the first two ticks belong to the 09:30 candle, while the final tick belongs to 09:31.

Networks, however, do not guarantee in‑order packet delivery. Your backend may receive packets in this jumbled sequence: 09:30:59.950 → 09:31:00.020 → 09:30:59.800

Under the arrival‑time‑based grouping logic, the delayed tick 09:30:59.800 would be incorrectly assigned to the 09:31 candle. Price ranges and volume values for both adjacent candles become distorted.

If these corrupted candles feed directly into backtesting, entry‑and‑exit trigger conditions, return figures, maximum drawdown and other core performance metrics will shift. In severe scenarios, you may observe excellent backtest results that completely fail to replicate in simulation or live trading.

A second frequent anomaly originates from feed replay after connection recovery. When a WebSocket session drops and reconnects, some A‑Share API providers replay a short window of recent tick history. Without deduplication safeguards, identical trades get counted multiple times.

Example: a single real trade with volume 100 arrives twice in duplicate packets. Aggregated volume becomes 200. Volume is a critical input for volume‑price factors and liquidity indicators. Inflated volume values distort factor distributions and introduce bias into statistical analysis and machine‑learning feature engineering.

📝 Research Note: Silent data bias cannot be caught by program‑level exception handling. You must periodically sample and compare aggregated outputs against authoritative benchmark datasets. For quantitative workflows, data validation should be treated as a mandatory pre‑backtesting step.

2. Root‑cause Analysis

2.1 Cross‑minute out‑of‑order packets: distinguish trade time versus packet‑arrival time

When processing tick‑based market data, two separate time dimensions must be strictly differentiated. This principle forms the foundation of reliable candle aggregation.



  1. Trade timestamp embedded inside each tick: the exact moment when the exchange matched the trade. This is the single source‑of‑truth for time‑bucket assignment.
  2. Packet‑arrival timestamp: system time when the tick payload reaches your local machine or server. This value is affected by network jitter, upstream feed scheduling and server‑side load. Never use arrival timestamps for candle grouping.

In streaming‑data environments, it is normal for a trade that occurred earlier to arrive later than trades belonging to the subsequent calendar minute.

Another subtle pitfall: receiving the first tick of a new minute does not guarantee that all ticks belonging to the prior minute have already arrived. Even after your aggregation logic switches to a new minute window, delayed late‑arriving packets from the previous minute can still appear. Blindly dropping these late ticks creates incomplete candles and introduces another form of data distortion.



2.2 Duplicate tick packets: greater research‑side harm than out‑of‑order delivery

Duplicate packets are commonly triggered by WebSocket disconnect‑reconnect cycles, resubscription to market‑data feeds, or duplicate consumption inside message queues.

Out‑of‑order delivery merely misplaces trades into the wrong candle bucket. Duplicate ticks directly inflate aggregated‑volume statistics and corrupt raw volume‑price feature distributions. For momentum, volume‑ratio and liquidity‑related factors, this type of data pollution creates substantial model bias.

Sample duplicate tick records:



09:30:12.123  15.20  100
09:30:12.123  15.20  100

Without deduplication, aggregated volume equals 200, while true market volume is only 100.

⚠️ Critical research caveat: building deduplication purely from a composite key of timestamp + price + volumecannot achieve 100 % accuracy. Within real A‑Share exchange matching, independent trades can coincidentally share identical timestamps, prices and volume sizes. Composite‑key deduplication carries inherent risk of false‑positive filtering that discards valid trades. Document this constraint inside your research notes. Whenever your A‑Share API exposes trade‑unique‑IDs or global sequence numbers, prioritize those identifiers for idempotency checks.

3. Production‑ready Implementation Patterns for Quant Research

The solutions outlined below are designed for two common research workflows: real‑time streaming tick processing, and off‑line historical‑tick replay for backtest‑dataset construction.



3.1 Time‑bucket assignment: use only native tick trade timestamps

Enforce this rule: minute‑candle bucket assignment must rely exclusively on the exchange‑provided trade timestamp embedded within each tick payload. Packet‑arrival timestamps must never participate in grouping logic.

Convert raw Unix timestamps into minute‑granularity bucket identifiers:



minute = tick_timestamp // 60

You may also format timestamps into human‑readable string keys such as 2026‑09‑07 09:30, which map directly to in‑memory candle‑aggregation objects.

This rule applies equally for real‑time streaming pipelines and offline batch replay of historical tick archives.



3.2 Bounded sliding memory buffer for cross‑minute late‑arriving ticks

Many publicly‑shared demo snippets finalize the previous candle immediately upon detecting a minute boundary transition:



if tick_minute != current_minute:
    finalize(current_kline)
    current_kline = create_kline(tick)
    current_minute = tick_minute
else:
    update_kline(current_kline, tick)

This implementation works in perfectly‑ordered demo environments. In real‑world streaming feeds, late‑arriving ticks for the prior minute may arrive after you have already switched aggregation context. Simply discarding these ticks causes permanent loss of genuine trade samples.

✅ Recommended approach: maintain a bounded sliding in‑memory buffer that retains candle‑aggregation instances for only the most recent 3‑5 minutes.



  1. For every incoming tick, parse its native exchange trade timestamp.
  2. Locate the corresponding minute bucket inside the buffer and update the target candle object.
  3. Discard a tick only if its timestamp falls completely outside the buffer’s time window.

This design gracefully absorbs brief out‑of‑order behaviour caused by network jitter. The same buffered logic is also useful during offline batch replay, to handle out‑of‑order records already present inside historical tick datasets.



3.3 Two‑tier idempotent deduplication, adaptable to API field capabilities

We adopt a hierarchical deduplication strategy that accommodates different field‑sets exposed by various A‑Share API providers.



  1. If your upstream API returns trade‑level unique identifiers or global sequence numbers, use ID‑based idempotent filtering as your primary mechanism. This delivers the highest possible reliability:
if tick_id in processed_ticks:
    return
processed_ticks.add(tick_id)
  1. When unique identifiers are unavailable upstream, construct composite deduplication keys assembled from multiple business fields:
dedup_key = (
    symbol,
    timestamp,
    price,
    volume
)
📌 Research‑workflow note: Composite keys reduce duplication probability but cannot eliminate mis‑classification risk. If your research demands high volume‑data fidelity, prioritize data sources that expose per‑trade unique identifiers. Schedule periodic sampling comparisons between your aggregated candles and benchmark market data to validate volume distributions.

3.4 Decouple processing‑pipeline layers for easier debugging and reproducibility

Avoid monolithic functions that mix ingestion, cleansing and aggregation logic together. Decouple your workflow into three clearly‑defined layers to simplify anomaly troubleshooting and offline replay reproducibility.

LayerCore ResponsibilityIngestion LayerMaintain WebSocket sessions and consume raw tick payloads from the A‑Share API. For offline workflows, read local historical‑tick datasets.Cleansing LayerValidate timestamp sanity, perform idempotent deduplication, filter malformed and abnormal market‑data samples.Aggregation LayerConsume cleansed tick records and compute minute‑level OHLCV candle metrics.

Minimal WebSocket client demonstration (structural example only):



import websocket
import json

def on_message(ws, message):
    data = json.loads(message)

    for tick in data.get("data", []):
        process_tick(tick)

ws = websocket.WebSocketApp(
    "wss://api.alltick.co/stock/websock...",
    on_message=on_message
)

ws.run_forever()
Implementation note: Adjust subscription parameters and field‑mapping strictly following your selected market‑data API official documentation. For offline backtest pipelines, remove WebSocket‑related logic and iterate over local tick files directly.

3.5 Separate real‑time preview and back‑test persistence; re‑evaluate candle‑finalization logic

One pervasive engineering misconception: using server‑system clock time to judge whether a minute candle can be marked as complete.

When your server clock strikes 09:31:00, calendar time has advanced into a new minute. This does not guarantee that every single tick belonging to 09:30 has finished arriving at your backend. If you finalize candles immediately at wall‑clock minute boundaries, subsequent late trades get lost from your dataset.

Recommended practice: implement a short grace‑waiting window. Optionally combine this window with upstream‑provided feed sequence numbers to safely decide when a minute bucket can be closed. Split data outputs into two independent pipelines:



  1. Real‑time preview pipeline: Optimized for low‑latency live dashboard rendering. Allow candle values to keep revising during the grace window. Do not use outputs from this pipeline directly for backtesting or model training.
  2. Backtest‑ready persistence pipeline: Accept modest latency in exchange for dataset correctness. Wait until the grace window expires and confirm no further late ticks are incoming. Complete deduplication and correction work before writing finalized OHLCV records to storage. Treat these persisted candles as your authoritative baseline dataset for backtesting, factor calculation and model training.
✍️ Practical research takeaway: Always feed backtesting workflows with finalized records produced by the persistence pipeline. Avoid consuming unfinalized real‑time candles for research.

4. Resource‑consumption Optimizations for Multi‑symbol Workloads

When your research pipeline subscribes to or processes ticks for dozens or hundreds of instruments simultaneously, memory usage and CPU overhead can escalate rapidly. Below are three production‑proven optimization practices.



  1. Constrain candle‑buffer time range: Do not retain full‑trading‑day candle objects resident in‑memory. Align with A‑Share exchange trading hours and limit your sliding buffer to the latest 3‑5 minutes. Release expired candle instances and rely on disk persistence for historical records. Apply time‑shard partitioning during offline batch processing to bound memory footprint.
  2. Periodically prune deduplication collections: Sets storing tick_id values or composite dedup_key entries must not grow infinitely. Schedule cleanup jobs to evict entries falling outside your active time window, reducing garbage‑collection pressure and mitigating potential memory leaks.
  3. Symbol‑aware differentiated processing: Optimize candle‑update routines for high‑liquidity heavily‑traded instruments and minimize unnecessary object copies. Reuse generic aggregation logic for thinly‑traded low‑activity symbols; avoid over‑engineering which increases maintenance and debugging burden.

5. Research‑oriented Takeaways & Conclusion

Generating reliable 1‑minute candles from raw tick data is far more sophisticated than a simple GROUP‑BY operation. Data quality hinges on seemingly minor architectural decisions: time‑bucket assignment rules, tolerance for out‑of‑order network packets, idempotent handling for duplicate payloads, and carefully‑designed candle‑finalization semantics.

Within quantitative research workflows, practitioners commonly devote most attention to strategy logic, hyper‑parameter tuning and model‑building. It is easy to overlook upstream market‑data pre‑processing. Real‑world experience demonstrates that many puzzling mismatches between backtest results and simulated/live performance originate from silent bias inside raw datasets, not flaws inside strategy algorithms.

We recommend adding formal data‑validation steps to your research workflow. Periodically sample self‑aggregated OHLCV outputs and compare them against authoritative benchmark datasets. Run statistical sanity checks over price and volume distributions to catch aggregation‑logic defects early.

All code examples inside this article serve educational demonstration purposes only. Production‑grade research systems need additional supporting components: comprehensive exception handling, automatic WebSocket reconnection, structured logging and observability tooling for reproducible debugging. When evaluating market‑data suppliers, you may source raw tick feeds from AllTick API and apply the aggregation patterns covered in this article to build your own custom research datasets.



Discussion

Have you encountered backtest bias introduced by tick‑data pre‑processing while constructing your own quantitative datasets? Feel free to share your observations, pitfalls and mitigation approaches in the comment thread.


08 Sep 2026, 11:43 を編集しました

免責事項:本記事で述べられている見解は著者の見解のみであり、Followmeの公式見解を反映するものではありません。Followmeは、提供された情報の正確性、完全性、信頼性について一切責任を負いません。また、書面で明示的に記載されている場合を除き、本記事の内容に基づいて行われたいかなる行動についても責任を負いません。

この記事が気に入ったら、著者にチップを送って感謝の気持ちを表しましょう。
応答 0

  • tradingContest