Intelligent period and current fusion market analysis system and method

The intelligent futures-spot fusion market analysis system utilizes the Transformer model to automatically fuse and standardize futures and spot data, solving the problems of data silos and timeline asynchrony. This enables high-precision trend prediction and trading signal generation, improving trading efficiency and accuracy.

CN122023007APending Publication Date: 2026-05-12PUSHAN TECHNOLOGY DEVELOPMENT (SICHUAN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610316411.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-16
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing technologies in futures-spot combined trading suffer from problems such as data silos, asynchronous data timelines, outdated analysis methods, low prediction accuracy, and cumbersome operations, making it difficult for traders to make accurate decisions when faced with high-frequency market fluctuations.

Method used

The intelligent futures and spot market analysis system adopts a layered architecture consisting of a data source layer, a data access and processing layer, a core computing layer, and an application presentation layer. It uses the Transformer model for data processing and prediction to achieve automatic fusion and standardization of futures and spot data, and provides high-precision trend warnings through AI models.

Benefits of technology

It has achieved automated integration and standardization of futures and spot data, significantly improving the accuracy of market forecasts, reducing manual operations, increasing trading efficiency, and reducing the rate of human error.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122023007A_ABST
    Figure CN122023007A_ABST
Patent Text Reader

Abstract

The invention relates to an intelligent period-present fusion market analysis system and method. The system comprises a data source layer, a data access and processing layer, a core calculation layer and an application presentation layer. The data source layer is configured to acquire data information of a futures exchange, a spot trading platform and a macroscopic / industry database in real time; the data access and processing layer is configured to receive, buffer and preprocess the data acquired by the data source layer; the core calculation layer is configured to perform reasoning calculation by loading a pre-trained Transform model and perform data storage; and the application presentation layer is configured to convert the calculation result into a transaction strategy and visual content of user interaction. According to the invention, the problem of data islands is solved, and automatic fusion and standardization of periodic data are realized; an AI deep learning model is introduced, so that the market prediction accuracy is remarkably improved; the full-automatic analysis process is realized, the transaction efficiency is greatly improved, and human errors are avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of big data processing, and in particular to an intelligent futures-spot fusion market analysis system and method. Background Technology

[0002] With the deepening of global economic integration, the price fluctuations of commodities (such as ferrous metals, chemicals, and agricultural products) are becoming increasingly volatile. As the world's largest consumer of commodities, my country's commodity futures market boasts a massive trading volume. According to statistics from relevant industry associations, my country's commodity futures trading volume has ranked first globally for many consecutive years; in 2023, the trading volume exceeded 5 billion lots, and the average daily open interest remained above 40 million lots. To mitigate the risks brought about by drastic price fluctuations, more and more physical enterprises and investment institutions are adopting a "futures-spot combination" trading model, that is, operating simultaneously in the futures and spot markets, using hedging or futures-spot arbitrage to lock in profits. Against this backdrop, the application of financial technology in the trading field is constantly deepening, placing extremely high demands on the speed and depth of data processing.

[0003] Currently, futures and spot trading data mainly comes from two channels: 1. Futures market data: This comes from the official market data interfaces of major futures exchanges (such as the Shanghai Futures Exchange, Dalian Commodity Exchange, and Zhengzhou Commodity Exchange). This data is highly standardized, with mainstream commodities experiencing a refresh rate of 500 milliseconds during peak trading hours, generating tens of millions of ticks daily. 2. Spot market data: This comes from spot e-commerce platforms (such as Zhaogang.com and SteelHome), information portals (such as Mysteel and Zhuochuang Information), or offline manual price inquiries. This type of data has highly inconsistent formats and vastly different update frequencies, often exhibiting discrete characteristics. For some non-mainstream commodities, the price update frequency is even as low as 10 minutes or less.

[0004] Current market analysis techniques mainly remain at the following stages: 1. Manual statistics and Excel calculations: Traders obtain spot quotes via telephone or instant messaging tools, manually enter them into Excel spreadsheets, and calculate the price difference using simple linear formulas (e.g., basis = spot price - futures price). 2. Independent market analysis software: Traditional futures trading software (such as WenHua Finance and Boyi Master) is used to view futures trends, while web browsers are used to view spot information. Although some high-end terminals provide cross-market arbitrage monitoring, they are mostly based on fixed rules and lack the ability to deeply integrate heterogeneous data. 3. Simple threshold alerts: Static thresholds are set based on linear logic, such as "an alarm is triggered when the price difference is greater than 500 points." However, such static rules often fail when facing high-frequency fluctuations and complex market sentiment.

[0005] For existing futures-spot combined scenarios, current technologies have the following significant shortcomings and deficiencies when dealing with massive, high-frequency, and unstructured data: 1. Severe data silos and difficulty in real-time automatic correlation of heterogeneous data: Futures market data is high-density time-series data, while spot market data is mostly unstructured text or discrete data. Current technologies lack effective data fusion middleware, making it difficult to achieve high-precision alignment of these two types of data in the time dimension. In practice, due to the lack of automated data cleaning pipelines, data engineers and traders typically spend 30% to 40% of their work time on data cleaning and reconciliation. In addition, the average acquisition latency of spot data is usually 1 to 5 minutes, while futures data is in the millisecond range. This severe "time lag" on the timeline causes the calculated basis to lag behind the true market value, easily leading to trading decisions based on erroneous information. 2. Lagging analytical methods fail to capture microsecond-level market fluctuations in real time: Existing analytical methods mainly rely on static historical statistics or manual experience review, which is essentially ex-post analysis. This lag is fatal when facing algorithmic high-frequency trading (HFT). In extreme market conditions (such as during major data releases), the price volatility of active commodities (such as rebar and crude oil) can surge within seconds. The calculation and response delays of traditional analytical systems typically exceed 10-30 seconds, while the optimal arbitrage window for high-frequency trading is often only 1-3 seconds. This technological lag means that by the time traders see an opportunity, it has already been absorbed by the market. 3. Low prediction accuracy and lack of in-depth analysis of massive nonlinear data: The futures-spot price spread is influenced by hundreds or even thousands of variables, including macroeconomic factors, supply and demand fundamentals, position structure, and investor sentiment, exhibiting highly nonlinear characteristics. Existing technologies mostly use linear regression models or simple moving averages for analysis, which assume a simple linear relationship between variables. According to industry backtesting data, in volatile markets or trend reversal phases, the prediction accuracy of traditional linear models is usually below 55% (only slightly higher than random guessing), and the prediction error rate for the basis regression direction even exceeds 15%. This low-precision prediction cannot meet the needs of enterprises for refined risk control, resulting in enterprises facing huge basis risk exposure in hedging. 4. Cumbersome operation, high labor costs, and low error tolerance: Due to the lack of an integrated intelligent analysis system, futures and spot traders often need to operate multiple independent software programs simultaneously. This multi-screen, multi-window operation mode not only distracts attention but also brings extremely high human error risks. Summary of the Invention

[0006] The purpose of this invention is to overcome the shortcomings of the prior art and provide an intelligent futures and spot market analysis system and method, which solves the deficiencies of the prior art.

[0007] The objective of this invention is achieved through the following technical solution: an intelligent futures and spot market analysis system, the system comprising a data source layer, a data access and processing layer, a core computing layer, and an application presentation layer;

[0008] The data source layer is configured to acquire data from futures exchanges, spot trading platforms, and macro / industry databases in real time.

[0009] The data access and processing layer is configured to receive, buffer, and preprocess the data acquired from the data source layer.

[0010] The core computing layer is configured to perform inference computations by loading a pre-trained Transformer model and to store data.

[0011] The application presentation layer is configured to transform the calculation results into interactive trading strategies and visual content for users.

[0012] The data access and processing layer includes a data access module, a message queue cluster, a data alignment and cleaning module, and a feature engineering module;

[0013] The data access module is deployed on a dual-path server, runs a multi-threaded program, and parses the binary stream data of the futures / spot interface through a protocol.

[0014] The message queue cluster adopts a distributed message queue, configured with 10 partitions and 3 replication factors to buffer traffic fluctuations caused by the data source structure.

[0015] The data alignment and cleaning module: uses a linear interpolation algorithm to fill the discrete spot prices forward to the futures time granularity, and performs outlier detection to remove noise caused by network jitter or input errors;

[0016] The feature engineering module calculates multiple features in real time based on the cleaned data and stores the calculation results in the feature engineering database.

[0017] The inference computation by loading a pre-trained Transformer model includes:

[0018] A1. The input layer of the Transformer model receives time series data from futures, spot, and macro / volume-price auxiliary channels;

[0019] A2. The feature extraction layer extracts local morphological features of each channel through parallel 1D-CNN;

[0020] A3. The attention fusion layer calculates the dynamic weights of the futures and spot channels to achieve feature-weighted fusion.

[0021] A4. The temporal inference layer uses LSTM to capture long-term dependencies, and the output layer outputs the basis trend prediction probability.

[0022] A5. The output layer determines the trading signal based on the probability value and threshold, and then pushes the trading signal to the front-end visualization terminal.

[0023] The output layer determines the trading signal based on probability values ​​and thresholds, including:

[0024] Set a dynamic threshold θ. If the probability value P>θ, it is determined to be a high-confidence bullish signal. If P<(1-θ), it is determined to be a high-confidence bearish signal. If θ<P<(1-θ), it is determined to be an oscillation signal.

[0025] If a high-confidence bullish signal is triggered, a JSON-formatted trading signal package is automatically generated, including the instrument code, operation type, expected take-profit level, stop-loss level, and signal generation timestamp.

[0026] The output layer determines the trading signal based on probability values ​​and thresholds, including:

[0027] Set a dynamic threshold θ. If the probability value P>θ, it is determined to be a high-confidence bullish signal. If P<(1-θ), it is determined to be a high-confidence bearish signal. If θ<P<(1-θ), it is determined to be an oscillation signal.

[0028] If a high-confidence bullish signal is triggered, a JSON-formatted trading signal package is automatically generated, including the instrument code, operation type, expected take-profit level, stop-loss level, and signal generation timestamp.

[0029] A smart futures-spot fusion market analysis method, the analysis method comprising:

[0030] S1. Monitor the ThostFtdcMdApi protocol data stream of futures and the RESTful API requests of spot in real time. When new tick data arrives, trigger an event interruption, store the input in the memory buffer, and clean and time-align the monitored data.

[0031] S2. Calculate multiple features in real time for the aligned data and perform feature normalization;

[0032] S3. Input the normalized feature vector into the pre-trained Transformer model to obtain the basis increase probability, and generate a trading signal by determining the threshold and push it to the front-end visualization terminal.

[0033] S3 specifically includes the following:

[0034] A1. The input layer of the Transformer model receives time series data from futures, spot, and macro / volume-price auxiliary channels;

[0035] A2. The feature extraction layer extracts the local morphological features of each channel through parallel one-dimensional convolutional layers;

[0036] A3. The attention fusion layer calculates the dynamic weights of the futures and spot channels to achieve feature-weighted fusion.

[0037] A4. The temporal inference layer uses LSTM to capture long-term dependencies, and the output layer outputs the basis trend prediction probability.

[0038] A5. The output layer determines the trading signal based on the probability value and threshold, and then pushes the trading signal to the front-end visualization terminal.

[0039] The output layer determines the trading signal based on probability values ​​and thresholds, including:

[0040] Set a dynamic threshold θ. If the probability value P>θ, it is determined to be a high-confidence bullish signal. If P<(1-θ), it is determined to be a high-confidence bearish signal. If θ<P<(1-θ), it is determined to be an oscillation signal.

[0041] If a high-confidence bullish signal is triggered, a JSON-formatted trading signal package is automatically generated, including the instrument code, operation type, expected take-profit level, stop-loss level, and signal generation timestamp.

[0042] The A1 includes: setting a time window T=60, constructing a futures price channel matrix and a spot price channel matrix, and inputting them into the feature extraction layer;

[0043] The A2 includes: using two parallel one-dimensional convolutional layers to process the futures and spot channel matrices respectively, and extracting local morphological features;

[0044] The A3 includes: concatenating the feature matrices of the futures and spot channels, and calculating the attention weight vector α through a fully connected layer and a Sigmoid activation function;

[0045] The A4 includes: weighting and fusing the feature matrices of the futures and spot channels according to the attention weight vector α to obtain the fused feature vector Xfused, and inputting the feature vector Xfused into a two-layer LSTM network to extract the long-term dependencies in the time series. Finally, the fully connected layer maps the output of the LSTM to the prediction result.

[0046] This invention has the following advantages: an intelligent futures-spot fusion market analysis system and method that solves the data silo problem and realizes automatic fusion and standardization of futures-spot data; introduces an AI deep learning model to significantly improve the accuracy of market prediction; and realizes a fully automated analysis process, greatly improving trading efficiency and avoiding human error. Attached Figure Description

[0047] Figure 1 This is a schematic diagram of the system of the present invention;

[0048] Figure 2 This is a schematic flowchart of the method of the present invention;

[0049] Figure 3 This is a schematic diagram of the process for dual-channel spatiotemporal attention fusion prediction.

[0050] Figure 4 A comparison chart of basis distribution before and after alignment of current data;

[0051] Figure 5 This is a bar chart comparing the prediction accuracy of the AI ​​model of this invention with that of traditional methods;

[0052] Figure 6 This is a comparison chart showing the time consumption of the automated process of this invention with that of the traditional process. Detailed Implementation

[0053] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the detailed description of the embodiments of this application provided below with reference to the accompanying drawings is not intended to limit the scope of protection of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application. The present invention will be further described below with reference to the accompanying drawings.

[0054] One embodiment of this invention relates to an intelligent (AI) futures-spot fusion market analysis system, which solves the problems of heterogeneous data sources between futures and spot markets, asynchronous timelines, and low accuracy and slow response speed of traditional rule-based market analysis and prediction in existing technologies. This system, through deep learning algorithms, achieves fully automated real-time monitoring of the commodity futures and spot markets, multi-dimensional feature fusion, and high-precision trend early warning.

[0055] like Figure 1 As shown, this system adopts a layered distributed microservice architecture, connecting various hardware facilities through a 10 Gigabit switch to achieve efficient data flow and low-latency processing. The architecture is divided into four layers from bottom to top, and the modules and specific parameters of each layer are as follows:

[0056] Data source layer: As the system's data input end, it contains three core data sources, whose data characteristics and access methods are as follows:

[0057] CTP (Commodity Transfer Platform): Connects to mainstream futures markets such as the Shanghai Futures Exchange and the Dalian Commodity Exchange. It obtains Level-2 market data through the CTP integrated market data interface, with a data frequency of 500 milliseconds / tick (daily processing volume of over 50 million ticks). The data format is a binary stream (such as the ThostFtdcMdApi protocol), which needs to be converted into a structured object through a custom decoder.

[0058] Spot trading platform (API): Connects to more than 10 mainstream spot e-commerce platforms such as Zhaogang.com and Zhuochuang Information. It obtains discrete spot quotes through RESTful API, with a data frequency of about 30 seconds / time (average daily processing volume of about 2 million records). The data format is JSON and includes fields such as commodity code, price, and update time.

[0059] Macro / Industry Database: Integrates external data sources such as the National Bureau of Statistics and Wind Information to obtain macroeconomic indicators such as GDP growth rate, PMI, and industry inventory. The data is updated daily (historical data goes back to 2010), and the data format is CSV / Excel. It needs to be cleaned by an ETL tool before importing.

[0060] Data Access and Processing Layer: This layer is responsible for data reception, buffering, and preprocessing. It contains four core modules, and its hardware configuration and processing logic are as follows:

[0061] Data access module: Deployed on a server with dual Intel Xeon Gold 6248R CPUs (2.5GHz, 28 cores), 32GB DDR4 memory, and dual 10 Gigabit Ethernet cards, running a Python 3.9 multi-threaded program (number of threads = number of CPU cores × 2), parsing binary stream data from futures / spot interfaces via TCP / IP protocol, with a throughput of 100,000 data entries per second;

[0062] Message queue cluster: Kafka 2.8.0 distributed message queue is used, configured with 10 partitions and 3 replication factors, with a throughput of 1 million messages / second, to buffer traffic fluctuations caused by heterogeneous data sources (such as sudden bursts of futures tick data).

[0063] Data alignment and cleaning module: Deployed on the Apache Flink 1.14.0 stream computing framework, it uses millisecond-level timestamps of futures contracts as a benchmark and employs a linear interpolation algorithm to "push forward" discrete spot prices to the futures time granularity (time window size = 1 minute). Simultaneously, it performs 3-Sigma outlier detection (standard deviation multiple = 3, sliding window size = 100 data points) to remove noise caused by network jitter (latency > 200ms) or input errors (price fluctuation > 5%). After cleaning, the data accuracy reaches 99.9%.

[0064] Feature Engineering Module: Based on the cleaned data, it calculates 12 types of features in real time (such as basis, basis MA5 / MA20, RSI14, OI_Chg, volume-weighted average price, etc.), with a feature update frequency of 1 second / time, and stores the results in the feature engineering database (Redis / InfluxDB).

[0065] Core Computing Layer: This layer is the core computing engine of the system, responsible for AI model inference and data storage. Its performance metrics are as follows:

[0066] AI analytics engine (GPU inference server): Deployed on an NVIDIA Tesla A100 40GB GPU (or T4 16GB GPU) cluster (4 nodes in total), it loads a pre-trained Transformer model (ONNX 1.12.1 format), utilizes the parallel computing capabilities of CUDA 11.7 to perform matrix multiplication operations on the feature matrix (dimension=256), and controls the inference latency to 3 milliseconds / inference (single node throughput=50,000 times / second).

[0067] Redis caching: Uses Redis 6.2.6 in-memory database, configured with 128GB memory and RDB persistence (backed up every 5 minutes), to store the latest price matrix, basis sequence and other hot data, with read and write latency <1 millisecond;

[0068] InfluxDB Time Series Database: Utilizes InfluxDB 2.0, configured with automatic sharding (shard duration = 1 hour), retention policy (data retention for 30 days), stores historical minute / hourly candlestick data (daily write volume = 10 million records), and has a query latency of <50 milliseconds.

[0069] Application Presentation Layer: This layer is responsible for transforming the calculation results into interactive trading strategies and visualizations for users, specifically including:

[0070] Strategy generation and push service: Based on the Spring Boot 2.7.5 microservice framework, using RabbitMQ 3.9.0 message queue, it receives probability signals output by the AI ​​engine (such as the probability of basis rising >85%), and converts them into specific trading suggestions (entry point, stop loss point, take profit point), with a push latency of <100 milliseconds;

[0071] Visualization terminal (PC / large screen): It adopts the Vue.js 3.2.13 front-end framework, combined with the ECharts 5.4.2 chart library, and uses the Socket.IO 4.5.4 WebSocket protocol to render candlestick charts in real time (update frequency = 1 second), basis trend (with warning signals), and trading signal pop-ups (such as a red warning for "buy arbitrage"). The data push latency is <200 milliseconds.

[0072] The modules at each layer are connected via data flow: data from the data source layer enters the Kafka message queue through the data access module, and is then distributed to the data alignment and cleaning module and the feature engineering module; the feature engineering module interacts with the AI ​​analysis engine through Redis / InfluxDB, and the calculation results of the AI ​​analysis engine are finally transmitted to the strategy generation and visualization terminal of the application presentation layer, forming a complete data processing closed loop, with the overall system latency controlled within 500 milliseconds (from data access to signal push).

[0073] like Figure 2 As shown, another embodiment of the present invention relates to an intelligent futures-spot fusion market analysis method, which specifically includes the following:

[0074] Step ①: Data Stream Monitoring: After system startup, the background daemon maintains communication links with the futures exchange's CTP front-end server (port 10200) and the spot e-commerce platform's API server (port 8080) via long connections. A multi-threaded I / O model (such as Python's asyncio or C++'s libevent) is used to monitor the data streams from the futures' ThostFtdcMdApi protocol and the spot's RESTful API requests, respectively. When new tick data (containing fields such as commodity code, latest price, best bid price, best ask price, trading volume, and timestamp) arrives, an event interrupt is triggered, and the data is stored in a memory buffer (buffer size = 1000 records, using a circular queue structure).

[0075] Step 2: Data Cleaning: After obtaining the raw data, the system first performs a physical validity check.

[0076] Check if the price field is empty or a non-positive number (e.g., the price of rebar must be within the range of 3000-6000 yuan / ton).

[0077] Check if the timestamp has reversed (the current timestamp must be greater than the timestamp of the previous data, and network latency within 5 seconds is allowed).

[0078] Check if the trading volume is 0 (remove invalid market data with no transactions).

[0079] If the data is abnormal, discard the data and record an error log (including error type, timestamp, and original data) through the logging module, then proceed to the "End this processing" process; if the data is normal, proceed to the next step.

[0080] Step 3: Time Alignment: Since futures data is a millisecond-level tick stream (frequency approximately 500ms / tick), while spot data is a discrete price quote (frequency approximately 30s / time), the system uses the "holding pricing method" for time alignment:

[0081] Maintain a 60-second sliding event window (stores the spot price of the most recent 60 seconds). Using the timestamp Tfuture of the futures data as the benchmark, find the nearest spot price Pspot in the spot time window (if there is no price in the window, use the previous valid price). Construct an aligned data pair (Pfuture, Pspot, Tfuture), where Pfuture is the latest futures price and Pspot is the aligned spot price.

[0082] Step 4: Feature Extraction: Based on the aligned data, 12 types of physical features are calculated in real time, specifically including:

[0083] Basis characteristics: St = Pspot − Pfuture (reflects the spot-futures price difference, unit: yuan / ton);

[0084] Momentum characteristics: The rate of change of the basis over the past 10 time windows, ΔSt = St - 10St - St - 10 × 100%;

[0085] Volume and position characteristics: Correlation between the rate of change of futures open interest (OI_Chg) and trading volume (calculated by Pearson correlation coefficient, range [-1,1]);

[0086] Volatility characteristics: Implied volatility calculated using the GARCH(1,1) model (Where ω=0.01, α=0.1, β=0.89 are model parameters);

[0087] Technical indicators: Moving averages (MA5, MA20), Relative Strength Index (RSI14), Bollinger Bands, etc.

[0088] Step 5: Feature Normalization: To eliminate differences in feature dimensions, the system uses the Min-Max normalization method to map feature values ​​to the [0,1] interval.

[0089] ,

[0090] Where Xmin and Xmax are the minimum and maximum values ​​of the features in the training set (e.g., the minimum value of the basis is -500 and the maximum value is 500), and the dimension of the normalized feature vector is 1×12.

[0091] Step 6: AI Model Inference: Input the normalized feature vectors into the pre-trained Transformer model and deploy it on an NVIDIA Tesla A100 GPU server. The model performs the following operations through CUDA parallel computing:

[0092] Input layer: Receives time-series data from futures, spot, and macro / volume-price auxiliary channels;

[0093] Feature extraction layer: Extracts local morphological features of each channel through parallel 1D-CNN (one-dimensional convolutional layer);

[0094] Attention Fusion Layer: Calculates the dynamic weights of the futures and spot channels to achieve feature weighted fusion;

[0095] Temporal inference layer: Utilizes LSTM to capture long-term dependencies and outputs the basis trend prediction probability P;

[0096] Output layer: Determines trading signals based on probability values ​​and thresholds.

[0097] Step 7: Confidence Determination: The system sets a dynamic threshold θ (the threshold θ can be dynamically adjusted according to market volatility, with a default value of 0.85, set based on a historical backtesting accuracy of 85%), and the determination rules are as follows:

[0098] If P > θ, it is considered a high-confidence bullish signal (the basis is likely to rise).

[0099] If P < (1-θ) (i.e., P < 0.15), it is considered a high-confidence bearish signal (the basis is likely to fall).

[0100] Otherwise, it is judged as a oscillation signal (no trading advice is generated).

[0101] Step 8: Generate trading signals: If a high-confidence bullish signal is triggered, a JSON-formatted trading signal package will be automatically generated, including the instrument code, operation type, expected take-profit level, stop-loss level, and signal generation timestamp.

[0102] Step 9: Silent monitoring: If no high-confidence bullish signal is triggered, end this process and continue monitoring.

[0103] Step 10: Push to the front end: The trading signal is pushed to the front-end visualization terminal (Vue.js + ECharts) via the WebSocket protocol (port 9000) at a frequency of once per second. After receiving the signal, the front end marks a red warning pop-up on the candlestick chart (e.g., "Buy Arbitrage: Take Profit 3800, Stop Loss 3650") and alerts the user with an audio-visual alarm.

[0104] Furthermore, if steps ①, ②, and ⑦ are determined to be "no", the system enters a loop and waits (waiting time 100ms) until new data arrives; if data cleaning fails, the system records an exception log (including error type, time, and data content), ends the current processing, and continues to monitor the data stream.

[0105] Furthermore, such as Figure 3 As shown, this invention employs a dual-channel spatiotemporal attention fusion prediction algorithm, designs a dual-channel convolutional structure, and introduces a gated attention mechanism to dynamically capture market focus, thus solving the problem of insufficient prediction accuracy of traditional models in scenarios of futures-spot divergence.

[0106] 1. Input tensor construction:

[0107] Set the time window T=60 (corresponding to approximately 3 minutes of high-frequency data, covering the short-term fluctuation cycle of the futures market), and construct three input matrices:

[0108] Futures price channel matrix: XF∈R1×T, containing the latest futures prices for the past 60 months;

[0109] Spot price channel matrix: XS∈R1×T, containing the past 60 aligned spot quotes.

[0110] 2. Dual-channel local feature extraction: Two parallel 1D-CNNs (one-dimensional convolutional layers) are used to process the futures and spot channel matrices respectively, extracting local morphological features:

[0111] Futures Channel Convolution: Kernel size is 3, stride is 1, number of channels is 32, activation function is ReLU;

[0112] HF = ReLU(WF * XF + bF).

[0113] Spot channel convolution: The parameters are the same as those of the futures channel, and it independently extracts local features of the spot market;

[0114] HS = ReLU(WS*XS + bS),

[0115] Where WF and WS are learnable convolutional kernel weight matrices, and bF and bS are bias terms.

[0116] 3. Dynamic Attention Weight Calculation: To address the issue of varying degrees of market influence from futures and spot markets at different times (e.g., when there is a "futures-spot divergence," futures sentiment dominates pricing; when supply and demand are tight, the spot market plays a dominant role), a gating attention mechanism is introduced to dynamically adjust the weights of the futures and spot channels.

[0117] (1) Feature splicing: The feature matrices of the futures and spot channels are spliced ​​together to obtain ;

[0118] (2) Weight calculation: The attention weight vector is calculated through the fully connected layer and the Sigmoid activation function. ,in, The sigmoid function maps the output to the (0,1) interval, and Wgate is a learnable weight matrix.

[0119] 4. Feature Fusion and LSTM Processing: The dual-channel features are weighted and fused according to weight α to obtain the fused feature vector. , where ⊙ represents element-wise multiplication.

[0120] Xfused is input into a two-layer LSTM network (each layer has 128 hidden units) to extract long-term dependencies from the time series. .

[0121] 5. Output and Decision Threshold: The final fully connected layer maps the LSTM output ht to the predicted result. Where P represents the probability that the basis will increase within a specific future time period.

[0122] like Figure 4 As shown, this invention addresses the core pain points of heterogeneous data sources and asynchronous timelines between futures and spot markets through a futures-spot data alignment and cleaning module. This module employs a stream computing framework (such as Apache Flink), using millisecond-level tick data from futures markets as a benchmark. It employs a linear interpolation algorithm to "push forward" discrete spot market quotes to the futures time granularity, while simultaneously performing 3-Sigma outlier detection (standard deviation multiple = 3, sliding window = 100 data points) to automatically remove price noise caused by network jitter or data entry errors. Compared to traditional manual alignment (which requires manual matching of commodity codes and manual completion of missing data, resulting in low efficiency and a high risk of errors), this module controls data alignment latency to within 1 second, increasing data consistency from 70% to over 99.5% compared to traditional methods.

[0123] like Figure 5As shown, this invention employs a "dual-channel spatiotemporal attention fusion prediction algorithm" (including components such as 1D-CNN, gated attention mechanism, and LSTM). Compared to traditional statistical methods (such as simple arbitrage formulas and manual experience analysis), it can more accurately capture the nonlinear patterns of the futures and spot markets (such as basis regression, futures-spot divergence, and other complex patterns). Traditional methods rely on fixed rules (such as shorting when the basis > 500), which cannot adapt to dynamic market changes, and the prediction accuracy is only about 60%. In contrast, the AI ​​model of this invention, trained on historical data (covering 3 years of tick-level data, approximately 5 million samples), achieves an accuracy of 85% on the validation set, an F1-score of 0.82, and an AUC (area under the curve) of 0.88. For example, in the 2023 rebar futures-spot arbitrage scenario, the model successfully captured the divergence signal of "declining spot inventory + surging futures open interest," predicting a 90% probability of a basis increase, ultimately achieving arbitrage profits 30% higher than traditional methods.

[0124] like Figure 6 As shown, this invention completely replaces traditional manual operations (requiring manual software switching, data export, basis calculation, and trend analysis) through an end-to-end automated process (from data acquisition to signal push). In the traditional process, traders need to spend 2-3 hours a day processing data, and the basis calculation error rate is as high as 15% (such as incorrect spot price entry or timestamp mismatch). This system, however, automatically monitors the data stream through a background daemon process, completing cleaning, alignment, feature extraction, model inference, and signal push in real time, reducing traders' workload by more than 90%, and lowering the human calculation error rate to below 0.1%. For example, when the basis fluctuates abnormally, the system can generate a trading signal and push it to the front end within 500 milliseconds, allowing traders to respond promptly without manual intervention, significantly improving trading efficiency and decision-making accuracy.

[0125] The above description is merely a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein and should not be construed as excluding other embodiments. It can be used in various other combinations, modifications, and improvements, and can be altered within the scope of the concept described herein through the above teachings or related technologies or knowledge. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention should be within the protection scope of the appended claims.

Claims

1. An intelligent futures-spot fusion market analysis system, characterized in that: The system includes a data source layer, a data access and processing layer, a core computing layer, and an application presentation layer; The data source layer is configured to acquire data from futures exchanges, spot trading platforms, and macro / industry databases in real time. The data access and processing layer is configured to receive, buffer, and preprocess the data acquired from the data source layer. The core computing layer is configured to perform inference computations by loading a pre-trained Transformer model and to store data. The application presentation layer is configured to transform the calculation results into interactive trading strategies and visual content for users.

2. The intelligent futures-spot fusion market analysis system according to claim 1, characterized in that: The data access and processing layer includes a data access module, a message queue cluster, a data alignment and cleaning module, and a feature engineering module; The data access module is deployed on a dual-path server, runs a multi-threaded program, and parses the binary stream data of the futures / spot interface through a protocol. The message queue cluster adopts a distributed message queue, configured with 10 partitions and 3 replication factors to buffer traffic fluctuations caused by the data source structure. The data alignment and cleaning module: uses a linear interpolation algorithm to fill the discrete spot prices forward to the futures time granularity, and performs outlier detection to remove noise caused by network jitter or input errors; The feature engineering module calculates multiple features in real time based on the cleaned data and stores the calculation results in the feature engineering database.

3. The intelligent futures-spot fusion market analysis system according to claim 1, characterized in that: The inference computation by loading a pre-trained Transformer model includes: A1. The input layer of the Transformer model receives time series data from futures, spot, and macro / volume-price auxiliary channels; A2. The feature extraction layer extracts local morphological features of each channel through parallel 1D-CNN; A3. The attention fusion layer calculates the dynamic weights of the futures and spot channels to achieve feature-weighted fusion. A4. The temporal inference layer uses LSTM to capture long-term dependencies, and the output layer outputs the basis trend prediction probability. A5. The output layer determines the trading signal based on the probability value and threshold, and then pushes the trading signal to the front-end visualization terminal.

4. The intelligent futures-spot fusion market analysis system according to claim 3, characterized in that: The output layer determines the trading signal based on probability values ​​and thresholds, including: Set a dynamic threshold θ. If the probability value P>θ, it is determined to be a high-confidence bullish signal. If P<(1-θ), it is determined to be a high-confidence bearish signal. If θ<P<(1-θ), it is determined to be an oscillation signal. If a high-confidence bullish signal is triggered, a JSON-formatted trading signal package is automatically generated, including the instrument code, operation type, expected take-profit level, stop-loss level, and signal generation timestamp.

5. The intelligent futures-spot fusion market analysis system according to claim 3, characterized in that: The output layer determines the trading signal based on probability values ​​and thresholds, including: Set a dynamic threshold θ. If the probability value P>θ, it is determined to be a high-confidence bullish signal. If P<(1-θ), it is determined to be a high-confidence bearish signal. If θ<P<(1-θ), it is determined to be an oscillation signal. If a high-confidence bullish signal is triggered, a JSON-formatted trading signal package is automatically generated, including the instrument code, operation type, expected take-profit level, stop-loss level, and signal generation timestamp.

6. A smart futures-spot fusion market analysis method, characterized in that: The analytical method includes: S1. Monitor the ThostFtdcMdApi protocol data stream of futures and the RESTful API requests of spot in real time. When new tick data arrives, trigger an event interruption, store the input in the memory buffer, and clean and time-align the monitored data. S2. Calculate multiple features in real time for the aligned data and perform feature normalization; S3. Input the normalized feature vector into the pre-trained Transformer model to obtain the basis increase probability, and generate a trading signal by determining the threshold and push it to the front-end visualization terminal.

7. The intelligent futures-spot fusion market analysis method according to claim 6, characterized in that: S3 specifically includes the following: A1. The input layer of the Transformer model receives time series data from futures, spot, and macro / volume-price auxiliary channels; A2. The feature extraction layer extracts the local morphological features of each channel through parallel one-dimensional convolutional layers; A3. The attention fusion layer calculates the dynamic weights of the futures and spot channels to achieve feature-weighted fusion. A4. The temporal inference layer uses LSTM to capture long-term dependencies, and the output layer outputs the basis trend prediction probability. A5. The output layer determines the trading signal based on the probability value and threshold, and then pushes the trading signal to the front-end visualization terminal.

8. The intelligent futures-spot fusion market analysis method according to claim 7, characterized in that: The output layer determines the trading signal based on probability values ​​and thresholds, including: Set a dynamic threshold θ. If the probability value P>θ, it is determined to be a high-confidence bullish signal. If P<(1-θ), it is determined to be a high-confidence bearish signal. If θ<P<(1-θ), it is determined to be an oscillation signal. If a high-confidence bullish signal is triggered, a JSON-formatted trading signal package is automatically generated, including the instrument code, operation type, expected take-profit level, stop-loss level, and signal generation timestamp.

9. The intelligent futures-spot fusion market analysis method according to claim 7, characterized in that: The A1 includes: setting a time window T=60, constructing a futures price channel matrix and a spot price channel matrix, and inputting them into the feature extraction layer; The A2 includes: using two parallel one-dimensional convolutional layers to process the futures and spot channel matrices respectively, and extracting local morphological features; The A3 includes: concatenating the feature matrices of the futures and spot channels, and calculating the attention weight vector α through a fully connected layer and a Sigmoid activation function; The A4 includes: weighting and fusing the feature matrices of the futures and spot channels according to the attention weight vector α to obtain the fused feature vector Xfused, and inputting the feature vector Xfused into a two-layer LSTM network to extract the long-term dependencies in the time series. Finally, the fully connected layer maps the output of the LSTM to the prediction result.