A financial knowledge graph driven logic generation and back testing system and method for quantification transaction

CN122840200APending Publication Date: 2026-09-29SHANGHAI QIANYU INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610960080.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-30
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0003]本申请所要解决的技术问题是提供一种面向量化交易的金融知识图谱驱动的逻辑生成与回测系统及方法,以克服现有技术中因子挖掘依赖人工、隐性关联难以提取,以及回测环境缺乏状态突变模拟能力的技术缺陷

Benefits of technology

[0014]本申请的有益效果在于:通过关系图卷积网络将复杂的定性知识图谱拓扑关系转化为定量的隐性因子,拓展了因子库的维度;同时,隐马尔可夫模型与蒙特卡洛马尔可夫链结合的状态转移模拟机制打破了线性历史数据的限制,对逻辑参数在极端市场状态切换下的鲁棒性进行了自动化压力测试,降低了逻辑在实盘应用中的尾部风险与回撤。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122840200A_ABST
    Figure CN122840200A_ABST
Patent Text Reader

Abstract

The application discloses a kind of financial knowledge graph driven logic generation and backtest system and method for quantification transaction.The system and method include: data aggregation server accesses market data and carries out entity alignment extraction, graph database storage cluster persistently stores four-dimensional financial knowledge graph, logic reasoning server inputs the adjacency matrix and node initial feature matrix of the four-dimensional financial knowledge graph into the operation unit based on relationship graph convolution network to extract implicit factor, market environment is divided into implicit market state using hidden Markov model and state transition probability matrix is estimated, market state switching path is generated according to state transition probability matrix using Monte Carlo Markov chain to construct virtual price curve matrix to carry out backtest calculation.The application can convert complex graph topological relationship into quantitative implicit factor to expand factor library dimension, and automatically test extreme market state switching, reduce tail risk and drawdown.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of quantitative trading and financial technology data processing, and specifically refers to a logic generation and backtesting system and method driven by financial knowledge graphs for quantitative trading. Background Technology

[0002] Existing quantitative logic generation systems, due to their combination of time-series data feature extraction and multiple linear regression models, rely heavily on human experience for factor mining, resulting in low efficiency and highly homogenized output. Furthermore, traditional historical backtesting frameworks directly use historical asset price time-series sequences as the benchmark environment, lacking the ability to simulate systemic state changes and implicit topological relationships. When market conditions undergo sudden changes, the logic parameter combinations generated based on static historical data often fail, leading to high-risk drawdowns in actual trading. Summary of the Invention

[0003] The technical problem to be solved by this application is to provide a logic generation and backtesting system and method driven by financial knowledge graph for quantitative trading, so as to overcome the technical defects of existing technologies, such as factor mining relying on manual methods, difficulty in extracting implicit associations, and lack of state change simulation capabilities in the backtesting environment.

[0004] To address the aforementioned issues, this application provides a logic generation and backtesting system driven by a financial knowledge graph for quantitative trading, comprising a data aggregation server, a graph database storage cluster, and a logic inference server. The data aggregation server is used to access market data for entity alignment and extraction. The graph database storage cluster is communicatively connected to the data aggregation server and is used to persistently store a four-dimensional financial knowledge graph covering macroeconomics, industry cycles, underlying assets, and trading rules. The logic inference server is connected to the graph database storage cluster and configured to input the adjacency matrix and initial node feature matrix of the knowledge graph into a computational unit based on a relational graph convolutional network. It calculates node embedding representation vectors using a multi-relational graph convolutional propagation model and extracts them as latent factors. It extracts historical time-series data from the knowledge graph, uses a Hidden Markov Model to divide the market environment into implicit market states, and estimates the state transition probability matrix using the expectation-maximization algorithm. It uses a Monte Carlo Markov chain to generate market state switching paths based on the state transition probability matrix, and concatenates historical price and volume data for corresponding state intervals under the market state switching paths to construct a virtual price curve matrix for backtesting calculations.

[0005] Furthermore, the computation unit is configured to assign an independent weight matrix to each relation type according to the different relation types of the interaction edge set in the four-dimensional financial knowledge graph; based on the adjacency matrix, select the set of neighbor nodes of each type of the target node; use the corresponding weight matrix to perform linear transformation and aggregation calculation on the initial feature vectors of different types of neighbor nodes respectively, and concatenate the aggregation calculation result with the initial feature vector of the target node itself to obtain the updated node embedding representation vector.

[0006] Furthermore, the multi-relation graph convolutional propagation model is configured to construct an updated node embedding representation vector through mathematical operations, and to determine the node embedding representation vector at the end output as a latent factor.

[0007] Furthermore, a Hidden Markov Model (HMM) is used to divide the market environment into implicit market states. Specifically, this involves acquiring multi-dimensional time series data on macroeconomics and industry cycles and standardizing them according to a normal distribution; setting a preset total number of clusters for the implicit market states; initializing the global parameter set of the HMM; the global parameter set includes the initial state probability distribution vector, the state transition probability matrix between different states, and the observation emission probability distribution matrix corresponding to each implicit market state; and inputting the standardized multi-dimensional time series data as the environmental observation sequence into the HMM.

[0008] Furthermore, the state transition probability matrix is ​​estimated using the expectation-maximization algorithm. Specifically, this involves an iterative update process of performing expectation calculation and maximization update steps until the change in the global parameter set converges to a preset tolerance threshold. The updated state transition probability matrix represents the mathematical probability boundary of the transitions between different market environment cycles.

[0009] Furthermore, the process of generating market state switching paths using Monte Carlo Markov chains based on the state transition probability matrix includes setting the total simulation time step of the backtesting period and the initial market state at the start time; in each subsequent time step iteration, extracting the single-row transition probability vector corresponding to the market state of the current step in the state transition probability matrix; constructing the cumulative probability distribution interval corresponding to the single-row transition probability vector; driving a pseudo-random number generator to generate random sample values, mapping and comparing the random sample values ​​to the cumulative probability distribution interval, determining the target implicit state of the next time step, and recording the target implicit states of all steps in time sequence to constitute the market state switching path.

[0010] Furthermore, the process of constructing a virtual price curve matrix by splicing historical volume and price data for corresponding state intervals under the market state switching path includes: randomly extracting continuous historical volume and price data segments with corresponding state labels for each continuous identical state interval in the market state switching path; performing boundary smoothing alignment at the time connection nodes of historical volume and price data segments from two adjacent different state intervals; and aggregating and assembling all the smoothed spliced ​​time series along the time axis into a virtual price curve matrix. The boundary smoothing alignment operation includes extracting breakpoint indicators between consecutive segments and inversely compensating for the price difference ratio at the breakpoints by incorporating it into the initial sequence fluctuation sequence of the next segment based on an exponential decay function.

[0011] This application also provides a data processing method applied to a logic generation and backtesting system driven by a financial knowledge graph for quantitative trading. The method includes: accessing market data for entity alignment and extraction; driving a graph database storage cluster connected to a data aggregation server to persistently store a four-dimensional financial knowledge graph covering macroeconomics, industry cycles, underlying assets, and trading rules; inputting the adjacency matrix and initial node feature matrix of the knowledge graph into a computational unit based on a relational graph convolutional network, calculating node embedding representation vectors using a multi-relational graph convolutional propagation model and extracting them as latent factors; using a hidden Markov model to divide the market environment into implicit market states, and estimating the state transition probability matrix using the expectation-maximization algorithm; using a Monte Carlo Markov chain to generate market state switching paths based on the state transition probability matrix, and concatenating historical price and volume data of corresponding state intervals under the market state switching paths to construct a virtual price curve matrix for backtesting calculations.

[0012] This application also provides a computer device, including a memory and a processor, wherein the processor executes a computer program stored in the memory to implement the steps of the above-described data processing method.

[0013] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the above-described data processing method.

[0014] The beneficial effects of this application are as follows: by transforming complex qualitative knowledge graph topological relationships into quantitative latent factors through relational graph convolutional networks, the dimension of the factor library is expanded; at the same time, the state transition simulation mechanism combining hidden Markov models and Monte Carlo Markov chains breaks the limitation of linear historical data, and performs automated stress testing on the robustness of logic parameters under extreme market state switching, reducing the tail risk and drawdown of logic in live applications. Attached Figure Description

[0015] Figure 1This is a schematic diagram of the structure of a logic generation and backtesting system driven by financial knowledge graphs for quantitative trading, provided in an embodiment of the present invention.

[0016] Figure 2 This is a flowchart of the data processing method provided in the embodiments of the present invention.

[0017] Figure 3 This is a schematic diagram of the structure of the multi-relationship graph convolution module provided in an embodiment of the present invention.

[0018] Figure 4 This is a schematic diagram of the implicit market state division module provided in an embodiment of the present invention.

[0019] Figure 5 This is a schematic diagram of the virtual price curve reconstruction module provided in an embodiment of the present invention.

[0020] Explanation of reference numerals in the attached figures:

[0021] 101 Data Aggregation Server, 102 Graph Database Storage Cluster, 103 Logical Inference Server, 301 Computing Unit, 401 Environment State Partitioning Module, 501 Virtual Price Curve Reconstruction Module, V 301 Multi-relation graph convolutional propagation model, V 302 Controller, V 401 Hidden Markov Model, V 501 Monte Carlo Markov chain resampling mechanism. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0023] The technical solutions of the various embodiments of this application can be combined with each other, but only if they are based on the ability of a person skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the scope of protection claimed by this application.

[0024] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, specific embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0025] like Figure 1As shown, this application provides a financial knowledge graph-driven logic generation and backtesting system for quantitative trading. The system's hardware support platform mainly includes a data aggregation server 101, a graph database storage cluster 102, and a logic inference server 103. At the hardware connectivity level, the data aggregation server 101, the graph database storage cluster 102, and the logic inference server 103 are all connected via a high-speed fiber optic network to ensure low-latency transmission of massive amounts of financial data and model parameters within the system. The system's software architecture is divided from bottom to top into a data layer, a graph computation layer, a logic generation layer, and a backtesting verification layer. Through the deployment of this basic hardware architecture, the system can provide robust computing power and storage support for the training of complex graph neural networks and concurrent backtesting of massive amounts of data.

[0026] In the data access and preprocessing stage, the system is configured with a bus module for internal or inter-server communication and a market data application programming interface (CAPI) access module. The data aggregation server 101 accesses external market data interfaces, macroeconomic databases, and policy text data sources through this interface module. The data aggregation server 101 is used to perform entity alignment and extraction on the acquired raw market data. At the network communication level, the data aggregation server 101 uses a gigabit Ethernet interface based on Transmission Control Protocol / Internet Protocol (TCP / IP) to capture external text news source data and structured financial statement data in real time. The data aggregation server 101 deploys a natural language processing (NLP) model to scan and identify unstructured financial entities in the external text news source data. This NLP model uses a domain-pre-trained language representation model based on a transformer architecture to extract dense word vectors of text feature words. In the feature alignment stage, the data aggregation server 101 performs feature disambiguation and mapping alignment between the identified unstructured financial entities and the standard core fields in the structured financial statement data. After alignment, the system assigns a globally unique entity identifier to each fused entity. The entity identifier can be a 64-bit universally unique identifier to ensure the uniqueness of data during system flow. The data aggregation server 101 synchronously transmits node data with globally unique entity identifiers to the graph database storage cluster 102 to construct the underlying node network of the knowledge graph. This preprocessing mechanism eliminates semantic ambiguity between multi-source heterogeneous data, providing a standardized, high-quality data foundation for graph operations.

[0027] Graph database storage cluster 102 utilizes an attribute graph model to persistently store multi-dimensional financial entities, constructing a four-dimensional financial knowledge graph encompassing macroeconomics, industry cycles, underlying assets, and trading rules. To ensure the throughput of massive data reads and writes, graph database storage cluster 102 adopts a distributed horizontal scaling architecture. In terms of data structure implementation, the graph database uses key-value pair-based attribute tables to record the specific attributes and values ​​of nodes. The knowledge graph consists of a set of nodes and a set of edges. Macroeconomic nodes include values ​​such as the consumer price index and the growth rate of money supply; industry cycle nodes record the prosperity indicators and inventory cycle data of various industries; underlying asset nodes cover the basic attributes and time-series price sequences of products such as stocks, options, futures, and exchange-traded funds; trading rule nodes record corresponding constraints such as price fluctuation limits, margin ratios, and transaction fees. Building upon this foundation, the knowledge graph's interaction edge set encompasses topological edges defining physical mapping relationships. These include first mapping edges representing the causal driving influence between macroeconomic indicator nodes and industry cycle nodes; second mapping edges representing the subordinate affiliation between industry cycle nodes and target asset nodes; third mapping edges representing cross-shareholding among target asset nodes; and fourth mapping edges representing the operational constraints imposed on target asset nodes by transaction rule nodes. Through real-time access to the data stream for incremental updates, the graph database storage cluster 102 enables the four-dimensional financial knowledge graph to reflect the overall characteristics and dynamic evolution logic of complex financial markets.

[0028] The logic inference server 103 serves as the system's computing core, internally equipped with a graphics processing unit (GPU) array to support large-scale parallel matrix operations. For example... Figure 3 As shown, the logic inference server 103 deploys a computational unit 301 based on a relational graph convolutional network. Due to the highly complex heterogeneous topology of financial knowledge graphs, conventional algorithms struggle to quantify nonlinear relationships. The controller V in the logic inference server 103... 302 The adjacency matrix and initial feature matrix of the knowledge graph are input to the computation unit 301 as input signals, and the multi-relation graph convolution propagation model V is used. 301 Latent factor features are extracted and aggregated layer by layer.

[0029] The computation unit 301 assigns an independent weight matrix to each relation type based on the different relation types r of the interaction edge set in the four-dimensional financial knowledge graph. Based on the topological connectivity record of the adjacency matrix, the computation unit 301 filters out the sets of neighbor nodes of each type for the target node i. It then performs linear transformations and aggregation calculations on the initial feature vectors of the different types of neighbor nodes using the corresponding weight matrices. The aggregation calculation result is then concatenated with the target node's own initial feature vector to obtain the updated node embedding representation vector.

[0030] In the feature extraction process, the multi-relation graph convolutional propagation model V 301 The system is configured to update the embedding representation vector of target node i in the (l+1)th hidden layer through the following mathematical operations: a non-linear activation operation is performed, which aggregates neighbor node features based on relation weights and concatenates the node's own features.

[0031]

[0032] in, h is the embedding representation vector of the target node i after transmission; i (l) The embedding representation vector of the target node i before transmission; The activation function is nonlinear; in this embodiment, it is preferably a modified linear unit function to introduce nonlinear feature representation capabilities. Represents the complete set of relation types in a knowledge graph; c represents the set of neighboring nodes of the target node i under a specific relation type r; i,r W is a normalization constant and the total number of elements configured as the set of neighbor nodes of this type, used to prevent gradient anomalies during the forward propagation of the network; r (l) It is a learnable weight matrix specifically designed for relation type r at layer l, used to capture the heterogeneous impact of different financial logics; h j (l) W0 is the embedding representation vector of neighbor node j before propagation; (l) It is a self-loop weight matrix that preserves the historical characteristics of the target node itself.

[0033] After completing multiple network iterations, the logic inference server 103 extracts and strips the node embedding representation vectors from the terminal output, identifying them as latent factors of the underlying asset in the current graph spatiotemporal context. By introducing the computation unit 301 to perform multi-relation graph convolution extraction, the system can calculate latent logical factors containing global logic across isolated price sequences, expanding the feature dimension of logic generation. Parameters such as the network hidden layer dimension can be adaptively adjusted according to actual computing resources. Those skilled in the art can also achieve the feature extraction purpose of this application using other graph neural network variants with structural graph feature extraction capabilities.

[0034] After extracting the implicit logical factors, a dynamic environment is established to verify the robustness of the logic. For example... Figure 4As shown, in the backtesting verification layer, the logic inference server 103 deploys an environment state partitioning module 401 corresponding to the market state transition simulation mechanism. The environment state partitioning module 401 extracts historical time series data of macroeconomic nodes and industry cycle nodes from the knowledge graph. After acquiring the multidimensional time series data, the environment state partitioning module 401 performs normal distribution standardization processing on it. A hidden Markov model V is then used. 401 The market environment is divided into several implicit market states Sk. In this embodiment, the total number of clusters is preset to 3, corresponding to the low volatility trend state, the high volatility oscillation state, and the liquidity tightening risk state, respectively.

[0035] During the model parameter estimation process, the environment state partitioning module 401 initializes the hidden Markov model V. 401 The global parameter set includes the initial state probability distribution vector, the state transition probability matrix between different states, and the observation emission probability distribution matrix corresponding to each hidden market state. Standardized multidimensional time-series data is input as the environmental observation sequence into this hidden Markov model V. 401 Then, the environmental state division module 401 estimates the state transition probability matrix of the model using the expectation-maximization algorithm.

[0036] The estimation process involves an iterative cycle of expectation calculation and maximization update steps until the change in the global parameter set converges to a preset tolerance threshold. In the expectation calculation step, based on the global parameter set of the current iteration, a forward-backward algorithm is used to calculate the single-point posterior probability of being in a specific latent market state at any given time point, and the joint state posterior probability of being in two specific latent market states at two consecutive time points, given a sequence of environmental observations. In the maximization update step, the state transition probability matrix between different states in the next iteration is updated by dividing the time-accumulated value of the joint state posterior probability along the time dimension by the time-accumulated value of the single-point posterior probability.

[0037] The updated state transition probability matrix is ​​a square matrix that represents the mathematical probability boundaries of transitions between different market environment cycles. The elements within this matrix follow the following mathematical constraints:

[0038]

[0039] Among them, a ij Represents the probability element value for transitioning from state i to state j; P is the conditional probability operator; S t This represents the market state at time point t; This represents the market state at time point t-1. By estimating the closed loop using the above parameters, the system can characterize the probability boundary of non-stationary market state transitions.

[0040] After obtaining the state transition probability matrix, as follows Figure 5 As shown, the virtual price curve reconstruction module 501 within the logic inference server 103 initiates the backtesting calculation process. Upon receiving the combination of logic parameters, the system utilizes the Monte Carlo Markov chain resampling mechanism V... 501 The system generates a large number of market state switching paths based on the state transition probability matrix. The virtual price curve reconstruction module 501 sets the total simulation time step of the backtesting period and the initial market state at the start time. In each subsequent time step iteration, a single-row transition probability vector corresponding to the market state of the current step is extracted from the state transition probability matrix. A cumulative probability distribution interval corresponding to this single-row transition probability vector is constructed. A pseudo-random number generator is driven to generate random sample values ​​with numerical distributions within this specific range. These random sample values ​​are mapped and compared to the cumulative probability distribution interval to determine the target implicit state for the next time step. The system records the target implicit states of all steps in chronological order to constitute the market state switching paths.

[0041] After the market state transition path is determined, the virtual price curve reconstruction module 501 stitches together historical volume and price data for the corresponding state intervals under that path to construct a virtual price curve matrix. For each consecutive identical state interval in the market state transition path, the system randomly extracts historical volume and price continuous data segments with corresponding state labels from the time series database. At the time connection node of historical volume and price continuous data segments of two adjacent different state intervals, the virtual price curve reconstruction module 501 performs a boundary smoothing alignment operation.

[0042] The boundary smoothing alignment operation involves extracting the closing price indicator at the end of the previous segment and the opening price indicator at the beginning of the next segment, and calculating the breakpoint price difference ratio between them. Based on the exponential decay function, the system incorporates this breakpoint price difference ratio into the initial sequence fluctuation sequence of the next segment as a reverse compensation. The decay compensation coefficient can be set to 0.05 to eliminate unnatural price jumps caused by forced abrupt changes in environmental states. The virtual price curve reconstruction module 501 aggregates and assembles all the smoothed time series along the time axis into a virtual price curve matrix. The logic inference server 103 finally performs verification on this virtual environment matrix, forcibly introducing a sudden change environment through a state transition simulation mechanism to test the risk adaptability of the logic parameters under extreme state switching.

[0043] In conjunction with the aforementioned logic generation and backtesting system driven by financial knowledge graphs for quantitative trading, this application also provides a data processing method. For example... Figure 2 As shown, this method includes a series of collaborative data interaction and model computation steps, specifically including:

[0044] Step S201: Control the data aggregation server 101 to access market data for entity alignment extraction.

[0045] The system's main control bus initiates a scheduling command, instructing the data aggregation server 101 to collect external unstructured and structured report data through the network interface. The embedded language representation model is used to extract entity features and calculate similarity, eliminating semantic ambiguity and assigning globally unique entity identifiers to qualified nodes. This step establishes a clean and unified data input standard.

[0046] Step S202: Drive the graph database storage cluster 102, which is connected to the data aggregation server 101, to persistently store a four-dimensional financial knowledge graph covering macroeconomics, industry cycles, underlying assets, and trading rules.

[0047] The system controls the distributed graph database engine to receive aligned node data and instantiate physical connections such as causal edges and related edges. This step not only completes the data storage but also builds a highly available knowledge graph foundation for graph network computing by maintaining the real-time topology.

[0048] Step S203: Instruct the logic reasoning server 103 to input the adjacency matrix and initial node feature matrix of the knowledge graph into the operation unit 301 based on the relational graph convolutional network, and then pass the multi-relational graph convolutional propagation model V. 301 Calculate the node embedding representation vector and extract it as a latent factor.

[0049] The main program allocates computing resources from the graphics processing unit array, and, considering the heterogeneous connectivity characteristics in the four-dimensional graph, assigns independent parameter matrices to perform neighbor feature transfer and nonlinear fusion. After multiple forward propagation calculations, the high-dimensional topological features are reduced to quantitative embedded vector signals, which are then output as the core logical factors.

[0050] Step S204, using the Hidden Markov Model V 401 The market environment is divided into implicit market states, and the state transition probability matrix is ​​estimated using the expectation-maximization algorithm.

[0051] The system extracts time series data from the macroeconomic dimension, preprocesses it, and then inputs it into a clustering model. Through repeated expectation calculations and probability coverage, the algorithm converges within a tolerance range, outputting a state transition probability matrix that reflects the objective laws governing the transitions between different market states.

[0052] Step S205, utilize the Monte Carlo Markov chain resampling mechanism V 501 Market state transition paths are generated based on the state transition probability matrix. Historical volume and price data of the corresponding state intervals are then spliced ​​together under the market state transition paths to construct a virtual price curve matrix for backtesting calculations.

[0053] Controller V 302 The system drives a resampling engine, using pseudo-random mapping to generate multiple Markov jump paths. When extracting real data segments corresponding to different states, the system introduces breakpoint difference calculation and exponential decay compensation logic to achieve smooth boundary alignment at segment interfaces. The generated virtual price curve matrix is ​​used to verify the resilience of the extracted latent factors.

[0054] This embodiment provides a computer device, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the data processing method described in the above embodiment.

[0055] Specifically, the computer device is the physical entity that carries the technical solution of this invention. Its hardware architecture provides the necessary computing power and data storage environment for the automated execution of the aforementioned method process. The processor can be a central processing unit (CPU), or a graphics processing unit (GPU), network processor (NP), or other application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other chips with data processing capabilities. In a preferred embodiment, a heterogeneous computing architecture combining a multi-core CPU and GPU acceleration can be used, or a processor integrating an AI acceleration unit can be employed to improve algorithm execution efficiency and inference speed. The memory includes high-speed random access memory (RAM) and non-volatile memory (such as hard disks, solid-state drives (SSDs), or flash memory). RAM is used to cache data, ensuring low-latency response for real-time monitoring; non-volatile memory is used to persistently store computer program instructions, historical architecture snapshots, design-state constraint rule bases, and long-term evaluation reports. Furthermore, the computer device includes a system bus and a communication interface. The system bus enables high-speed data transmission between the processor, memory, and peripherals, while the communication interface supports network interconnection between the device and external data acquisition sources, cloud platforms, or user terminals, ensuring real-time data acquisition and timely push of analysis results.

[0056] When a computer program is loaded and executed by a processor, its instruction sequence drives hardware resources to complete the various steps described in the foregoing embodiments. It should be understood that the computer device described in this embodiment can be an independently deployed server or workstation, a node in a distributed cluster, or a cloud-based virtualization instance. As long as it has the ability to execute the aforementioned computer program and achieve the corresponding technical effects, it falls within the protection scope of this invention. Furthermore, those skilled in the art will understand that a computer-readable storage medium (such as an optical disc, USB flash drive, or read-only memory) storing the aforementioned computer program also constitutes a complete implementation carrier of the technical solution of this invention. When this medium is connected to a computer device and read and executed, the data processing method of this invention can also be implemented.

[0057] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0058] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A financial knowledge graph-driven logic generation and backtesting system for quantitative trading, comprising a data aggregation server, a graph database storage cluster, and a logic inference server; characterized in that, The data aggregation server is used to access market data for entity alignment and extraction. The graph database storage cluster is communicatively connected to the data aggregation server and is used to persistently store a four-dimensional financial knowledge graph covering macroeconomics, industry cycles, underlying assets, and trading rules. The logic reasoning server is connected to the graph database storage cluster and configured to perform the following operations. The adjacency matrix and initial feature matrix of the four-dimensional financial knowledge graph are input into the operation unit based on the relational graph convolutional network. The node embedding representation vector is calculated and extracted as a latent factor through the multi-relational graph convolutional propagation model. Historical time series data are extracted from the four-dimensional financial knowledge graph. Hidden Markov models are used to divide the market environment into implicit market states, and the state transition probability matrix is ​​estimated by the expectation-maximization algorithm. Using a Monte Carlo Markov chain, a market state switching path is generated based on the state transition probability matrix. Historical volume and price data of the corresponding state intervals are then concatenated under the market state switching path to construct a virtual price curve matrix for backtesting calculations.

2. The financial knowledge graph-driven logic generation and backtesting system for quantitative trading as described in claim 1, characterized in that, The computation unit is configured to perform the following operations: assign an independent weight matrix to each relation type according to the different relation types of the interaction edge set in the four-dimensional financial knowledge graph; based on the adjacency matrix, select the set of neighbor nodes of each type of the target node; perform linear transformation and aggregation calculation on the initial feature vectors of different types of neighbor nodes using the corresponding weight matrix respectively; and concatenate the aggregation calculation result with the initial feature vector of the target node itself to obtain the updated node embedding representation vector.

3. The financial knowledge graph-driven logic generation and backtesting system for quantitative trading as described in claim 2, characterized in that, The multi-relation graph convolutional propagation model is configured to construct the updated node embedding representation vector through the following mathematical operations. in, h is the embedding representation vector of the target node i after propagation. i (l) Let be the embedding representation vector of the target node i before transmission. It is a non-linear activation function. For the complete set of relation types, Let c be the set of neighboring nodes of the type of the target node i under relation r. i,r W is a normalization constant and the total number of elements configured as the set of neighbor nodes of this type. r (l) For the weight matrix h of relation r j (l) Let W0 be the embedding representation vector of neighbor node j before propagation. (l) The target node itself is defined as its weight matrix; the logic inference server determines the node embedding representation vector output at the end as the latent factor.

4. The financial knowledge graph-driven logic generation and backtesting system for quantitative trading as described in claim 1, characterized in that, The Hidden Markov Model (HMM) is used to divide the market environment into implicit market states. Specifically, this involves acquiring multi-dimensional time series data of macroeconomic and industry cycles and standardizing them according to a normal distribution; setting a preset total number of clusters for the implicit market states; initializing the global parameter set of the HMM, which includes an initial state probability distribution vector, a state transition probability matrix between different states, and an observation emission probability distribution matrix corresponding to each implicit market state; and inputting the standardized multi-dimensional time series data as an environmental observation sequence into the HMM.

5. The financial knowledge graph-driven logic generation and backtesting system for quantitative trading as described in claim 4, characterized in that, The state transition probability matrix is ​​estimated using the expectation-maximization algorithm. Specifically, the following interactive update process is iteratively executed until the change in the global parameter set converges to a preset tolerance threshold. In the expected measurement step, based on the global parameter set of the current iteration round, the forward-backward algorithm is used to calculate the single-point posterior probability of being in a specific implicit market state at any time point under the given environmental observation sequence, and the joint state posterior probability of being in two specific implicit market states at two consecutive time points. In the maximization update step, the time-accumulated value of the joint state posterior probability along the time dimension is divided by the time-accumulated value of the single-point posterior probability to cover and update the state transition probability matrix between the different states in the next iteration round; the updated state transition probability matrix represents the mathematical probability boundary of mutual transformation between different market environment cycles.

6. The financial knowledge graph-driven logic generation and backtesting system for quantitative trading as described in claim 1, characterized in that, The process of generating market state switching paths using a Monte Carlo Markov chain based on the state transition probability matrix specifically includes setting the total simulation time step of the backtesting period and the initial market state at the start time; in each subsequent time step iteration, extracting the single-row transition probability vector corresponding to the market state of the current step in the state transition probability matrix; and constructing the cumulative probability distribution interval corresponding to the single-row transition probability vector. A pseudo-random number generator is driven to generate random sample values ​​with numerical distribution within a specific range. The random sample values ​​are mapped and compared to the cumulative probability distribution interval to determine the target implicit state of the next time step. The target implicit states of all steps are recorded in time sequence to form the market state switching path.

7. The financial knowledge graph-driven logic generation and backtesting system for quantitative trading as described in claim 6, characterized in that, The process of constructing a virtual price curve matrix by stitching together historical volume and price data for corresponding state intervals under the market state switching path specifically includes: For each consecutive identical state interval in the market state switching path, a historical continuous volume and price data segment with a corresponding state label is randomly extracted from the time series database; At the time connection node of two adjacent historical price and volume data segments in different state intervals, a boundary smoothing alignment operation is performed. The boundary smoothing alignment operation includes extracting the closing price indicator at the end of the previous segment and the opening price indicator at the beginning of the next segment, calculating the breakpoint price difference ratio between the two, and incorporating the breakpoint price difference ratio into the initial sequence fluctuation sequence of the next segment based on the exponential decay function to eliminate unnatural price gaps caused by forced abrupt changes in environmental state; and aggregating and assembling all the smoothed time series along the time axis into the virtual price curve matrix.

8. A data processing method applied to a financial knowledge graph-driven logic generation and backtesting system for quantitative trading, characterized in that, The following processing steps are included. Access market data for entity alignment and extraction; Persistent storage encompasses a four-dimensional financial knowledge graph covering macroeconomics, industry cycles, underlying assets, and trading rules; The adjacency matrix and initial feature matrix of the four-dimensional financial knowledge graph are input into the operation unit based on the relational graph convolutional network. The node embedding representation vector is calculated and extracted as a latent factor through the multi-relational graph convolutional propagation model. Hidden Markov Models are used to divide the market environment into implicit market states, and the state transition probability matrix is ​​estimated by the expectation-maximization algorithm. Using a Monte Carlo Markov chain, a market state switching path is generated based on the state transition probability matrix. Historical volume and price data of the corresponding state intervals are then concatenated under the market state switching path to construct a virtual price curve matrix for backtesting calculations.

9. A computer device comprising a memory and a processor, wherein the processor executes a computer program stored in the memory to implement the steps of the data processing method of claim 8.

10. A computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the steps of the data processing method as claimed in claim 8.