A multi-source policy signal standardized version snapshot and factor multiplexing method and system
Patent Information
- Application Number
- CN202611035555.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-13
- Publication Date
- 2026-09-25
AI Technical Summary
第一,策略信号缺少统一资产化结构
第一,本发明不是单一交易策略算法,而是将AI指标、技术指标、模型输出和用户因子统一转换为平台级信号/因子资产,解决了信号分散在脚本、模型或页面逻辑中的问题。
Smart Images

Figure CN122816680A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer information processing technology, specifically to a method and system for standardizing and reusing multi-source strategy signals through versioning snapshots and factor reuse. Background Technology
[0002] In AI strategy computing platforms, strategy signals come from a variety of sources, including template strategies, AI Indicators, machine learning models, traditional technical indicators, expression factors, Python custom factors, or visual workflow diagrams. The output formats of these sources vary: some output buy / sell signals (e.g., buy / sell), some output continuous scores, some output model predictions, some output standard factor values, and some also include component explanations and data quality fields.
[0003] Existing technologies cover automatic optimization of trading strategies, execution of algorithmic trading, testing of script strategies, AI quantitative research frameworks, and model version management. However, for the signal generation and snapshot reuse chain of product-oriented strategy computing systems, the following shortcomings still exist: First, strategy signals lack a unified asset-based structure. Current signals are typically limited to a single script, a single model output, a single strategy tester, or a specific trading scenario. They lack a unified standard contract that includes fields such as asset identifier (asset), timestamp (ts), rule (rule), resolution (resolution), signal (signal), score (score), factor value (factor_value), and prediction period (horizon).
[0004] Second, signal results are difficult to reuse across modules. If the front-end dashboard, formal offline evaluation, simulation run, report snapshot, and subsequent model training calculate signals separately, inconsistencies can easily arise due to differences in parameters, data windows, quality judgments, or algorithm versions.
[0005] Third, there is a lack of version and lineage mechanisms for source-oriented tracking. While existing solutions can manage model versions or policy parameters, they do not uniformly write algorithm versions, parameter hashes, data quality, component decomposition, workflow nodes, training runs, and run batches into signal or factor records.
[0006] Fourth, there is a lack of background pre-computation and hot-reading mechanisms. If AI signals are calculated in real time every time a user opens a page, it will increase the pressure on the Web API and lead to instability in simulation execution and report reading results.
[0007] Fifth, the lack of a unified conversion layer between model outputs, AI metrics, custom factors, and traditional technical metrics makes it difficult for model predictions to be transformed into factor assets that can be ranked, evaluated offline, combined, and retrained.
[0008] Therefore, it is essential to design a method and system for standardizing and reusing multi-source strategy signals through versioning snapshots and factor reuse. Summary of the Invention
[0009] The purpose of this invention is to provide a method and system for standardizing versioned snapshots and factor reuse of multi-source strategy signals. Through standardization conversion and version snapshot management, the unified assetization of multi-source signals is realized, ensuring that different downstream consumers obtain consistent results.
[0010] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A method for standardizing and reusing versioned snapshots of multi-source strategy signals and factor reuse includes: Step 1: Receive raw strategy signals from different sources, identify the source type and value type of each raw strategy signal, and convert raw strategy signals with different field names and direction expressions into unified candidate tuples containing asset identifier, timestamp, market resolution, prediction period, raw value, direction, data quality and lineage information according to the preset field alias mapping, direction convention mapping and asset mapping rules. Step 2: Execute the strategy signal generation logic on the unified candidate tuple to generate a standard signal record containing asset identifier, signal time, strategy rules, market resolution, discrete trading signals, continuous scores, confidence level, and component explanation data; Step 3: Normalize the set of strategy parameters used to generate standard signal records and calculate the parameter hash value. Use the parameter hash value and the algorithm version number corresponding to the standard signal record as the version identifier of the standard signal. Step 4: Construct a unique snapshot key by combining the asset identifier, strategy rule, market resolution, signal time, algorithm version number, and parameter hash value as fields, and write the standard signal record, version identifier, and data quality result into the strategy signal snapshot table in an idempotent manner according to the unique snapshot key to form a versioned signal snapshot. Step 5: When standard signal records need to be reused as factors, according to the selected standardization mode, the continuous scores or raw values in the standard signal records are converted into standard factor values, and model output factor records containing model identifier, prediction period, asset identifier, timestamp, factor name and standard factor values are generated and stored in the factor table. Step 6: When writing data in Step 4 and Step 5, the validity flag representing whether the data quality has passed and the lineage information related to the source are persisted to the strategy signal snapshot table and / or factor table. The validity flag is set according to the data quality result to control whether the downstream consumer can consume the record. Step 7: Through a unified application programming interface, responding to read requests from different consumer terminals such as front-end display, offline evaluation, simulation operation, or model training, query and return the corresponding persistent signal snapshot or factor record from the policy signal snapshot table and / or factor table, so that each consumer terminal can obtain the same result without having to re-execute the signal algorithm.
[0011] Furthermore, in step 1, the original strategy signals with different field names and directional expressions are uniformly converted into unified candidate tuples containing asset identifier, timestamp, market resolution, prediction period, original value, direction, data quality, and lineage information, specifically including: The original strategy signals from sources such as AI indicators, technical indicators, model predictions, Workflow node outputs, or custom factor outputs are merged into the original numerical fields of a unified candidate tuple through a preset field alias mapping table, with semantically similar fields being grouped together. The direction mapping table maps various direction commands to standard direction values of long, short, and flat.
[0012] Further, in step 4, the snapshot unique key is constructed as follows: The asset identifier, strategy rules, market resolution, signal time, algorithm version number, and parameter hash value are concatenated into a unique string. When a stored record with the same unique key exists, it is written in an update manner to ensure that the snapshot of the results generated by the same computing task is idempotent.
[0013] Furthermore, in step 5, the standardization mode includes the original value mode, the robust Z-score mode, the cross-section ranking mode, the truncated mode, or the direction mapping mode; The robust Z-score mode calculates the median and median absolute deviation of continuous scores within a specified window and then performs a standardized transformation to obtain a standard factor value. The cross-sectional ranking mode calculates the quantile ranking of continuous scores within the same time frame and asset pool, and then generates a standard factor value accordingly.
[0014] Furthermore, in step 6, the setting rule for the validity flag is as follows: when the data quality result indicates that the quality gate has not passed, the validity flag of the record is set to a non-tradable state; when the downstream simulation operation or transaction module consumes data, it only reads records with the validity flag set to a tradable state.
[0015] Furthermore, in step 6, the lineage information includes at least one or more of the following: workflow identifier that generated the signal or factor, workflow run identifier, source node identifier, and model version number, for subsequent tracing and reproduction of the signal's source.
[0016] Furthermore, in step 7, the unified application programming interface provides the signal snapshot or factor record pre-generated and written by the background worker process, which is returned directly as a hot result; if it has not been generated or is missing due to expiration, an asynchronous computing task is triggered to perform supplementary calculation, in order to avoid performing large-scale recalculation on the real-time request path.
[0017] This invention also provides a multi-source strategy signal standardization versioned snapshot and factor reuse system, applied to the above-mentioned multi-source strategy signal standardization versioned snapshot and factor reuse method, comprising: The strategy input adaptation layer is used to receive raw strategy signals from different sources, identify the source type and value type of each raw strategy signal, and convert raw strategy signals with different field names and direction expressions into unified candidate tuples containing asset identifier, timestamp, market resolution, prediction period, raw value, direction, data quality and lineage information according to preset field alias mapping, direction convention mapping and asset mapping rules. The signal generation layer is used to execute strategy signal generation logic on unified candidate tuples to generate standard signal records that include asset identifiers, signal time, strategy rules, market resolution, discrete trading signals, continuous scores, confidence levels, and component interpretation data. The versioned snapshot layer is used to normalize the set of strategy parameters used to generate standard signal records and calculate parameter hash values. The parameter hash values and the algorithm version number corresponding to the standard signal record are used together as version identifiers. The layer also constructs a unique snapshot key by combining asset identifier, strategy rule, market resolution, signal time, algorithm version number, and parameter hash value as fields. Based on this unique snapshot key, the standard signal record, version identifier, and data quality results are idempotently written into the strategy signal snapshot table to form a versioned signal snapshot. The factor reuse layer is used to convert continuous scores or raw values in standard signal records into standard factor values according to the selected standardization mode when standard signal records need to be reused as factors. It also generates model output factor records containing model identifier, prediction period, asset identifier, timestamp, factor name and standard factor value, and stores them in the factor table. The traceability record layer is used to persist validity flags, which characterize whether the data quality has passed, and lineage information related to the source, to the strategy signal snapshot table and / or factor table when writing data to the versioned snapshot layer and factor reuse layer; wherein, the validity flag is set according to the data quality results to control whether the downstream consumer can consume the record; The hot read interface layer is used to respond to read requests from different consumer terminals such as front-end display, offline evaluation, simulation operation or model training through a unified application programming interface. It queries and returns the corresponding persistent signal snapshot or factor record from the policy signal snapshot table and / or factor table, so that each consumer terminal can obtain the same result without having to re-execute the signal algorithm.
[0018] In summary, the present invention has at least one of the following beneficial technical effects: First, this invention is not a single trading strategy algorithm, but rather a unified conversion of AI indicators, technical indicators, model outputs, and user factors into platform-level signal / factor assets, solving the problem of signals being scattered in scripts, models, or page logic.
[0019] Second, this invention enables the front-end display, formal offline evaluation, simulation operation, and report snapshot to read the same signal result by using the algorithm version number (algo_version), parameter hash value (params_hash), and snapshot unique key, thus avoiding the difference in caliber caused by repeated calculations from multiple entry points.
[0020] Third, this invention synchronously binds data quality access control, component decomposition, and lineage information to signal or factor records, making the strategy signal not just a buying or selling direction, but an interpretable, traceable, and auditable data object.
[0021] Fourth, this invention supports background pre-calculation and hot result reading, reducing the pressure of real-time calculation of Web API, and is suitable for productization scenarios with multiple users, multiple strategies, and multiple assets.
[0022] Fifth, this invention enables model outputs to be factorized, saved, and reused, allowing model prediction results to be incorporated into information coefficient (IC) analysis, cross-sectional ranking, combined weights, offline evaluation reports, and subsequent model training, forming long-term accumulative data assets.
[0023] Sixth, this invention retains scalability. Different algorithms can replace specific scoring methods, and different databases can replace specific storage methods, but as long as they follow the rules of signal standardization, version keys, unique keys, quality binding, and lineage, they can be connected to the platform's processing chain. Attached Figure Description
[0024] Figure 1 This is a schematic diagram of the method flow of the present invention; Figure 2 This is a schematic diagram of the system structure of the present invention. Detailed Implementation
[0025] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.
[0026] To facilitate understanding, key terms and symbols used in this invention are first explained using convention. This invention can be described using the following symbols and rules, without being limited to a specific programming language or database implementation.
[0027] Asset identifier: A, such as CL (crude oil), GC (gold) or stock code.
[0028] Signal timestamp: T, corresponding to bar_ts or ts.
[0029] Policy rules (R): R, such as ai_indicator_v1, ma_cross, etc.
[0030] Algorithm version: V, such as ai_indicator_v1.1.0.
[0031] Normalized parameter set (Parameters): P, the JSON of parameters serialized in a fixed field order.
[0032] Param's Hash: H, H = hash(V + P + R + resolution).
[0033] Data quality outcome (Quality): Q, which includes data_quality_passed and quality_issues.
[0034] Discrete trading signal (Signal): S, which can be long, short, flat, or 1, -1, 0.
[0035] Standard Factor Value (FV): A standardized value derived from score, prediction, or factor_value.
[0036] Lineage information: L, including workflow_id, run_id, source_node_id, model_version, etc.
[0037] Unified Candidate Tuple (U): Composed of asset, ts, horizon, source_type, value_kind, raw_value, direction, quality, and lineage.
[0038] Raw value to be normalized: Y, which can come from score, prediction, raw_prediction, or factor_value.
[0039] Median (M): a rolling window or cross-sectional median used for robust Z-score normalization.
[0040] Median Absolute Deviation (D_mad): Used to reduce the impact of extreme values.
[0041] Standardize Mode: mode, which can be raw, zscore, rank, clip, or direction.
[0042] Confidence (C): Derived from algorithm output or component coverage and quality status.
[0043] Effectiveness Flag: E, determined by a combination of is_valid and valid_for_trade.
[0044] Epsilon: epsilon is used for denominator protection.
[0045] Normalized truncation limit (Zmax): zmax is used for clipping to stabilize the output range.
[0046] This invention provides a method and system for standardized versioning snapshots and factor reuse of multi-source policy signals, applicable to multi-asset AI policy computation platforms. Policies within the platform may originate from template policies, AI indicators (AIIndicator), machine learning models, technical indicators, expression factors, or user-defined factors. The output formats of these sources vary; without a unified agreement, the platform cannot guarantee that front-end dashboards, offline evaluations, simulation runs, and report readings will yield the same results.
[0047] This invention defines a strategy signal as an assetizable data object that retains not only the signal itself, but also the algorithm version, parameter hash, data quality, component decomposition, and source batch information required to compute the signal, thus forming a traceable and reusable data asset.
[0048] Method Implementation Examples: Figure 1 This is a flowchart illustrating a method for standardizing and reusing multi-source strategy signals through versioning snapshots and factor reuse, provided in an embodiment of the present invention. Figure 1 As shown, the method specifically includes the following steps: Step 1: Receive raw strategy signals from different sources, identify the source type (source_type) and value type (value_kind) of each raw strategy signal, and convert the raw strategy signals with different field names and direction expressions into unified candidate tuples containing asset identifier, timestamp, market resolution, prediction period, original value, direction, data quality and lineage information according to the preset field alias mapping, direction convention mapping and asset mapping rules.
[0049] The system first receives multi-asset market data, feature matrices, strategy template parameters, model prediction outputs, and user-defined factor outputs, and categorizes them according to their source type (source_type) into AI indicators (indicator_signal), model predictions (model_prediction), factor values (factor_value), and workflow node outputs (workflow_node_output). Subsequently, it performs asset mapping, time bucket alignment, field alias mapping, direction convention unification, and quality status binding.
[0050] The key rule of the integration process (the unified rule for multi-source inputs) is as follows: A mapping function is used to map raw results from different sources into a unified candidate tuple U. Inputs with different field names but the same semantics are merged through a field alias mapping (alias_map). For example, semantically identical fields such as predicted value, raw predicted value, continuous score, and custom factor value can all be merged into the raw value field of the unified candidate tuple U. Directional expressions such as buy, sell, hold, long, short, and flat are merged into standard directional values of 1, -1, and 0, representing long, short, and flat, respectively, through a direction mapping (direction_map). Simultaneously, workflow node outputs establish source tracking relationships with node output fields through workflow identifier (workflow_id), run identifier (run_id), and source node identifier (source_node_id). If an input lacks a non-essential field, the system retains it as a null value and a record validity flag (is_valid); if duplicate inputs exist for the same asset, at the same time, and from the same source, deduplication or updates are performed based on the algorithm version, parameter hash, and generation time.
[0051] For example, in one specific embodiment, step 1 involves uniformly converting the original policy signals with different field names and directional expressions, specifically including: The original policy signals, which are sourced from AI indicators, technical indicators, model predictions, Workflow node outputs, or custom factor outputs, are merged into the original numerical fields of a unified candidate tuple through a pre-defined field alias mapping table, with semantically similar fields being grouped together. In addition, a direction mapping table is used to uniformly map diverse direction commands into standard long, short, and flat direction values.
[0052] Step 2: Execute the strategy signal generation logic on the unified candidate tuple to generate a standard signal record containing asset identifier, signal time, strategy rules, market resolution, discrete trading signals, continuous scores, confidence level, and component explanation data.
[0053] The signal generation layer executes specific signal algorithms. Taking the AI Indicator as an example, the system first calculates several normalized components: trend component x_trend = clip((EMA_fast-EMA_slow) / max(ATR, epsilon), -zmax, zmax) / zmax, momentum component x_momentum = clip(log(C_t / C_{tm}) / max(vol_m, epsilon), -zmax, zmax) / zmax, breakout component x_breakout = clip((C_t-mid_N) / max(range_N, epsilon), -zmax, zmax) / zmax, and RSI component x_rsi = (RSI_t-50) / 50. The system then merges the scores by weights to obtain a continuous score: score = sum(w_i*x_i) / max(sum(|w_i|), epsilon). When score ≥ the long threshold (theta_long), direction = 1 (representing a long position); when score ≤ the short threshold (theta_short), direction = -1 (representing a short position); otherwise, direction = 0 (representing closing a position or no direction). Traditional technical indicators can directly output direction; the model adapter uses the predicted value as the raw value; the workflow node adapter selects specified fields from the node output and includes lineage information. The system also saves component explanation data (components_jsonb) to record the value, weight, threshold, and triggering reason of each component, facilitating subsequent explanation, backtesting reproduction, and auditing.
[0054] Step 3: Normalize the set of strategy parameters used to generate standard signal records and calculate the parameter hash value (params_hash). Use the parameter hash value and the algorithm version number (algo_version) corresponding to the standard signal record as the version identifier of the standard signal.
[0055] The system generates parameter hash values (params_hash) based on the algorithm version and normalized parameters. Together with the algorithm version number (algo_version), these hash values constitute the signal version identifier, enabling signal records to be traced back to specific algorithm versions, parameter windows, thresholds, and component weights.
[0056] Step 4: Construct a unique snapshot key by combining the asset identifier, strategy rule, market resolution, signal time, algorithm version number, and parameter hash value as fields. Then, write the standard signal record, version identifier, and data quality result into the strategy signal snapshot table in an idempotent manner based on the unique snapshot key to form a versioned signal snapshot.
[0057] Step 4 follows the following snapshot unique key rule: A unique identifier for a strategy signal snapshot is the combination of asset identifier (A) + strategy rule (R) + market resolution + signal time (T) + algorithm version (V) + parameter hash (H). If the same unique key is calculated again, the signal, continuous score, component explanation data, and complete payload of that record are updated instead of generating duplicate signals, thus ensuring the idempotency of snapshots generated by the same computational task.
[0058] In one specific embodiment, a snapshot unique key is constructed by concatenating the asset identifier, strategy rule, market resolution, signal time, algorithm version number, and parameter hash value into a unique string. When a stored record with the same unique key exists, it is written in an update manner.
[0059] Step 5: When standard signal records need to be reused as factors, the continuous scores or raw values in the standard signal records are converted into standard factor values (factor_value) according to the selected standardization mode, and model output factor records containing model identifier, prediction period, asset identifier, timestamp, factor name and standard factor value are generated and stored in the factor table.
[0060] The system converts continuous scores, predictions, or raw values into standardized factor values (FVs) based on a standardization mode. The factor transformation rules are as follows: when the input contains continuous scores but lacks standardized factor values, the continuous scores are used as candidate factor values; when the input contains predictions but lacks standardized factor values, the predictions are used as candidate factor values; when scaling is required, different transformation methods are used depending on the mode. Standardization modes include: Raw mode: directly preserves the original numerical value (raw_value).
[0061] Robust Z-score model (zscore): Calculate Z = clip((YM) / max(D_mad, epsilon), -zmax, zmax), and set factor_value = Z / zmax.
[0062] Cross-sectional ranking mode (rank): Calculate percentile_rank(Y) - 0.5 for all assets within the same timestamp (ts), prediction period (horizon), and asset pool (universe).
[0063] Clip mode: Restricts the original value (Y) to its upper and lower bounds.
[0064] Direction mapping mode: Maps long, short, and flat positions to 1, -1, and 0 respectively.
[0065] The generated model output factor records are stored according to the unique key rule for model output factors, that is: model_id + horizon + asset + ts + factor_name is used as the unique identification combination of model output factors. If the same model, the same prediction period, the same asset, the same time, and the same factor name are written again, an idempotent update (upsert) is performed.
[0066] Step 6: When writing data in Step 4 and Step 5, the validity flag representing whether the data quality has passed, as well as the lineage information related to the source, are persisted to the strategy signal snapshot table and / or factor table. The validity flag is set according to the data quality results to control whether the downstream consumer can consume the record.
[0067] Step 6 follows these quality binding rules: If the data quality result (Q) indicates that the data quality gate has failed (data_quality_passed=false), the record can still retain the quality diagnostic information and payload, but the tradable flag (valid_for_trade) in the validity flag should be set to false to prevent misuse during simulation or deployment processes. If the original value is missing or cannot be converted, the system will set the record's validity flag (is_valid) to false. Source tracing rules require that each reusable signal or factor be associated with at least the algorithm version (algo_version), parameter hash (params_hash), or model identifier (model_id), and may further be associated with lineage information.
[0068] In one specific embodiment, the validity flag is set according to the following rule: when the data quality result indicates that the quality threshold has not been passed, the validity flag of the record is set to a non-tradable state; when the downstream simulation run or transaction module consumes data, it only reads records with the validity flag set to a tradable state. Lineage information includes at least one or more of the following: workflow identifier (workflow_id), workflow run identifier (run_id), source node identifier (source_node_id), and model version number (model_version), used for subsequent tracing and reproduction of the signal's source.
[0069] Step 7: Through a unified application programming interface (API), respond to read requests from different consumer terminals such as front-end display, offline evaluation, simulation operation, or model training, query and return the corresponding persistent signal snapshot or factor record from the policy signal snapshot table and / or factor table, so that each consumer terminal can obtain the same result without having to re-execute the signal algorithm.
[0070] The reuse layer does not reinterpret the strategy logic, but instead reads the persistent results. The front-end dashboard reads the latest hot snapshot, the formal offline evaluation reads historical signals under the same rules, the simulation run reads the latest signals to generate target positions, the report snapshot reads component explanations and quality fields, and subsequent model training can read persistent factors as input features. The consumer does not reinterpret the input source, nor does it repeatedly implement the strategy algorithm. It only reads standard signals or standard factors through a unified interface. The standard output contract rule it follows is: before writing any signal or factor from any source, one or more of the following should be generated: asset identifier (asset), timestamp (ts), prediction period (horizon), source type (source_type), continuous score (score), standard factor value (factor_value), confidence (confidence), algorithm version (algo_version), parameter hash (params_hash), record validity flag (is_valid), tradable flag (valid_for_trade), and lineage information (lineage_jsonb).
[0071] In one specific embodiment, a unified application programming interface provides signal snapshots or factor records that are pre-generated and written by a background worker process and returned directly as hot results; if they have not yet been generated or are missing due to expiration, an asynchronous computation task is triggered to perform supplementary computation, in order to avoid performing large-scale recomputation on the real-time request path.
[0072] The execution entities, data dependencies, and output relationships of the above steps are shown in Table 1.
[0073] Table 1. Execution Entity, Data Dependencies, and Output Relationships
[0074] Table 2 provides an adaptive summary and explanation of the key rules of this invention, their application in the system, and the corresponding method steps.
[0075] Table 2 Key rules of the present invention, their application in the system, and corresponding methods.
[0076] System Implementation Example: Figure 2 This is a structural block diagram of a multi-source strategy signal standardization versioning snapshot and factor reuse system provided in an embodiment of the present invention. Figure 2 As shown, the system is used to implement the above method embodiments, specifically including: The strategy input adaptation layer receives raw strategy signals from different sources, identifies the source type and value type of each raw strategy signal, and, based on preset field alias mapping, direction convention mapping, and asset mapping rules, uniformly converts raw strategy signals with different field names and direction expressions into unified candidate tuples containing asset identifier, timestamp, market resolution, prediction period, original value, direction, data quality, and lineage information. This layer corresponds to step 1 in the method embodiment.
[0077] The signal generation layer is used to execute strategy signal generation logic on the unified candidate tuples to generate standard signal records containing asset identifiers, signal times, strategy rules, market resolution, discrete trading signals, continuous scores, confidence levels, and component explanation data. This layer corresponds to step 2 in the method embodiment.
[0078] The versioned snapshot layer is used to normalize the set of strategy parameters used to generate the standard signal record and calculate the parameter hash value. The parameter hash value and the algorithm version number corresponding to the standard signal record are used together as a version identifier. Furthermore, a unique snapshot key is constructed using the asset identifier, strategy rule, market resolution, signal time, algorithm version number, and parameter hash value as field combinations. Based on this unique snapshot key, the standard signal record, version identifier, and data quality result are idempotently written into the strategy signal snapshot table to form a versioned signal snapshot. This layer corresponds to steps 3 and 4 in the method embodiment.
[0079] In one specific embodiment, the versioned snapshot layer constructs a unique snapshot key by concatenating the asset identifier, strategy rule, market resolution, signal time, algorithm version number, and parameter hash value into a unique string. When there is a stored record with the same unique key, it is written in an update manner to ensure that the snapshots of the results generated by the same computing task are idempotent.
[0080] The factor reuse layer is used to convert continuous scores or raw values in the standard signal records into standard factor values according to the selected standardization mode when the standard signal records need to be reused as factors, and to generate model output factor records containing model identifier, prediction period, asset identifier, timestamp, factor name and the standard factor values, and store them in the factor table. This layer corresponds to step 5 in the method embodiment.
[0081] In one specific embodiment, in the factor reuse layer, the factor table is established with a joint unique index based on model identifier, prediction period, asset identifier, timestamp and factor name, which is used to perform idempotent updates for repeated calculations of the same factor.
[0082] A traceability record layer is used to persist a validity flag indicating whether the data quality has passed, along with source-related lineage information, to the strategy signal snapshot table and / or the factor table when data is written to the versioned snapshot layer and the factor reuse layer. The validity flag is set based on the data quality result to control whether the downstream consumer can consume the record. This layer corresponds to step 6 in the method embodiment.
[0083] The hot-read interface layer is used to respond to read requests from different consumer terminals such as front-end display, offline evaluation, simulation operation, or model training through a unified application programming interface. It queries and returns the corresponding persisted signal snapshots or factor records from the policy signal snapshot table and / or the factor table, allowing each consumer terminal to obtain the same result without having to re-execute the signal algorithm. This layer corresponds to step 7 in the method embodiment.
[0084] The data transmission and processing relationships between different layers / modules of the system have been described in detail in the "Step Execution Subject, Data Dependency and Output Relationship Table" of the aforementioned method embodiment.
[0085] The following describes the key data structures in the system, which are used in steps 1 to 7 of the method embodiment.
[0086] In one embodiment, the design of the policy signal snapshot table is shown in Table 3.
[0087] Table 3 Strategy Signal Snapshot Table
[0088] This table establishes a joint unique constraint on normalized_symbol, rule, resolution, bar_ts, algo_version, and params_hash.
[0089] In another embodiment, the design of the model output factor table is shown in Table 4.
[0090] Table 4 Model Output Factors
[0091] This table establishes a joint unique constraint on model_id, horizon, asset, ts, and factor_name.
[0092] The table structure described above is not limited to a specific database type; it can be implemented using relational databases, time-series databases, or other persistent storage.
[0093] Specific application examples: To facilitate understanding, several specific application examples are given below.
[0094] Example 1: Generation of AI Indicator Signal Snapshots The platform configures an AI indicator for single-product CL strategies. The algorithm version (algo_version) is ai_indicator_v1.1.0, and the parameters include EMA fast line 12, EMA slow line 36, momentum window 8, breakout window 20, RSI window 14, bullish threshold 0.22, bearish threshold -0.22, and the weights of trend, momentum, breakout, and RSI components.
[0095] The system reads the 60m candlestick chart of CL from the market hotspot, first calculating EMA_fast, EMA_slow, ATR, short-term momentum, breakout range, and RSI, and then normalizing them according to four component categories: trend, momentum, breakout, and RSI. The system obtains a comprehensive score using the continuous score fusion formula: score = sum(w_i*x_i) / max(sum(|w_i|), epsilon). A long signal is generated when the score is higher than the bullish threshold, a short signal is generated when the score is lower than the bearish threshold, otherwise the flat signal or the current position direction is maintained. The system writes this signal to the strategy signal snapshot table and saves the component explanation data (components), complete parameters (params), payload, continuous score (score), direction (direction), confidence level (confidence), data quality pass flag (data_quality_passed), record validity flag (is_valid), and tradable flag (valid_for_trade). The front-end dashboard, simulation run, and reporting modules all read this snapshot, instead of implementing their own AI Indicator algorithms.
[0096] Example 2: Model Output Reuse as Standard Factor The upstream model outputs predicted values for a batch of assets. The model factor node in the system identifies the predicted value field and marks the source type as model_prediction and the value type as predicted. If the standard factor value is missing, the predicted value is used as a candidate original value Y.
[0097] If the standardization mode (standardize_mode) is set to cross-sectional ranking mode (rank), then percentile_rank(Y)-0.5 is calculated for all assets based on the same timestamp (ts), prediction period (horizon), and asset pool identifier (universe_id), forming the standard factor value (factor_value). If the robust Z-score mode (zscore) is set, then Z is calculated based on the median M and median absolute deviation D_mad in the rolling window, and Z / zmax is used as the standard factor. The system then writes the result to the model output factor table, with fields including model identifier (model_id), prediction period (horizon), asset identifier (asset), timestamp (ts), predicted value (prediction), factor name (factor_name), standard factor value (factor_value), workflow identifier (workflow_id), run identifier (run_id), source node identifier (source_node_id), lineage information (lineage_jsonb), record validity flag (is_valid), and factor usable from_ts.
[0098] Subsequent cross-sectional ranking strategies can directly read this factor for grouping, the reporting module can calculate IC and quantile returns, and another model can also use this factor as an input feature. Thus, the model output is no longer a one-off experimental result, but a manageable, reusable, and traceable standard factor asset.
[0099] Example 3: Background Pre-calculation and Hot Reading The background worker process periodically identifies the list of strategies and assets that need to be pre-computed, generates signals in batches and writes them to the strategy signal snapshot table, and records statistical indicators such as the number of computation requests (requested_count), ready number (ready_count), missing number (missing_count), stale number (stale_count), completed number (computed_count), cache item number (cache_item_count), and elapsed time (elapsed_ms) for this pre-computed process.
[0100] When the front-end page is opened or the simulation runs to check the policy status, the system prioritizes reading the most recent hot snapshot. If the hot snapshot has expired or is missing, it then decides whether to trigger an asynchronous computation task to perform a recalculation based on permissions and load policies. This approach ensures that the web user request path does not bear a large-scale recalculation task and guarantees that the page, simulation run, and report read consistent results within the same time window.
[0101] The correspondence between the system parameters of the present invention and its engineering implementation can be found in Table 5.
[0102] Table 5. Correspondence between system parameters and engineering implementation
[0103] In terms of engineering implementation, this invention is not limited to specific file names or programming languages; Table 6 is only used as an illustration of modular implementation. Table 6 Modular Implementation Table
[0104] Embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0105] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0106] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0107] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1The steps of the function specified in one or more boxes.
[0108] Contents not described in detail in this specification are prior art known to those skilled in the art. It is hereby indicated that the above description is intended to help those skilled in the art understand this invention, but does not limit the scope of protection of this invention. Any equivalent substitutions, modifications, improvements, or simplifications of the above descriptions that do not depart from the essential content of this invention fall within the scope of protection of this invention.
Claims
1. A method for standardizing and reusing versioned snapshots and factors of multi-source strategy signals, characterized in that, include: Step 1: Receive raw strategy signals from different sources, identify the source type and value type of each raw strategy signal, and convert raw strategy signals with different field names and direction expressions into unified candidate tuples containing asset identifier, timestamp, market resolution, prediction period, raw value, direction, data quality and lineage information according to the preset field alias mapping, direction convention mapping and asset mapping rules. Step 2: Execute the strategy signal generation logic on the unified candidate tuple to generate a standard signal record containing asset identifier, signal time, strategy rules, market resolution, discrete trading signals, continuous scores, confidence level, and component explanation data; Step 3: Normalize the set of strategy parameters used to generate standard signal records and calculate the parameter hash value. Use the parameter hash value and the algorithm version number corresponding to the standard signal record as the version identifier of the standard signal. Step 4: Construct a unique snapshot key by combining the asset identifier, strategy rule, market resolution, signal time, algorithm version number, and parameter hash value as fields, and write the standard signal record, version identifier, and data quality result into the strategy signal snapshot table in an idempotent manner according to the unique snapshot key to form a versioned signal snapshot. Step 5: When standard signal records need to be reused as factors, according to the selected standardization mode, the continuous scores or raw values in the standard signal records are converted into standard factor values, and model output factor records containing model identifier, prediction period, asset identifier, timestamp, factor name and standard factor values are generated and stored in the factor table. Step 6: When writing data in Step 4 and Step 5, the validity flag representing whether the data quality has passed and the lineage information related to the source are persisted to the strategy signal snapshot table and / or factor table. The validity flag is set according to the data quality result to control whether the downstream consumer can consume the record. Step 7: Through a unified application programming interface, responding to read requests from different consumer terminals such as front-end display, offline evaluation, simulation operation, or model training, query and return the corresponding persistent signal snapshot or factor record from the policy signal snapshot table and / or factor table, so that each consumer terminal can obtain the same result without having to re-execute the signal algorithm.
2. The method for standardizing and reusing versioned snapshots and factors of multi-source strategy signals according to claim 1, characterized in that, In step 1, the original strategy signals with different field names and directional expressions are uniformly converted into unified candidate tuples containing asset identifier, timestamp, market resolution, prediction period, original value, direction, data quality, and lineage information, specifically including: The original strategy signals from sources such as AI indicators, technical indicators, model predictions, Workflow node outputs, or custom factor outputs are merged into the original numerical fields of a unified candidate tuple through a preset field alias mapping table, with semantically similar fields being grouped together. The direction mapping table maps various direction commands to standard direction values of long, short, and flat.
3. The method for standardizing and reusing versioned snapshots and factors of multi-source strategy signals according to claim 1, characterized in that, In step 4, a snapshot unique key is constructed, specifically as follows: The asset identifier, strategy rules, market resolution, signal time, algorithm version number, and parameter hash value are concatenated into a unique string. When a stored record with the same unique key exists, it is written in an update manner to ensure that the snapshot of the results generated by the same computing task is idempotent.
4. The method for standardizing and reusing versioned snapshots and factors of multi-source strategy signals according to claim 3, characterized in that, In step 5, the standardization mode includes the original value mode, the robust Z-score mode, the cross-section ranking mode, the truncated mode, or the direction mapping mode; The robust Z-score mode calculates the median and median absolute deviation of continuous scores within a specified window and then performs a standardized transformation to obtain a standard factor value. The cross-sectional ranking mode calculates the quantile ranking of continuous scores within the same time frame and asset pool, and then generates a standard factor value accordingly.
5. The method for standardizing and reusing versioned snapshots and factors of multi-source strategy signals according to claim 1, characterized in that, In step 6, the setting rule for the validity flag is as follows: when the data quality result indicates that the quality gate has not passed, the validity flag of the record is set to a non-tradable state; when the downstream simulation operation or transaction module consumes data, it only reads records with the validity flag set to a tradable state.
6. The method for standardizing and reusing multi-source strategy signals through versioning snapshots and factor reuse according to claim 1, characterized in that, In step 6, the lineage information includes at least one or more of the following: workflow identifier that generated the signal or factor, workflow run identifier, source node identifier, and model version number, for subsequent tracing and reproduction of the signal's source.
7. The method for standardizing and reusing multi-source strategy signals through versioning snapshots and factor reuse according to claim 1, characterized in that, In step 7, the unified application programming interface provides the signal snapshot or factor record pre-generated and written by the background worker process, which is returned directly as a hot result; if it has not been generated or is missing due to expiration, an asynchronous computing task is triggered to perform supplementary calculation, in order to avoid performing large-scale recalculation on the real-time request path.
8. A multi-source strategy signal standardization versioning snapshot and factor reuse system, applied to the multi-source strategy signal standardization versioning snapshot and factor reuse method according to any one of claims 1-7, characterized in that, include: The strategy input adaptation layer is used to receive raw strategy signals from different sources, identify the source type and value type of each raw strategy signal, and convert raw strategy signals with different field names and direction expressions into unified candidate tuples containing asset identifier, timestamp, market resolution, prediction period, raw value, direction, data quality and lineage information according to preset field alias mapping, direction convention mapping and asset mapping rules. The signal generation layer is used to execute strategy signal generation logic on unified candidate tuples to generate standard signal records that include asset identifiers, signal time, strategy rules, market resolution, discrete trading signals, continuous scores, confidence levels, and component interpretation data. The versioned snapshot layer is used to normalize the set of strategy parameters used to generate standard signal records and calculate parameter hash values. The parameter hash values and the algorithm version number corresponding to the standard signal record are used together as version identifiers. The layer also constructs a unique snapshot key by combining asset identifier, strategy rule, market resolution, signal time, algorithm version number, and parameter hash value as fields. Based on this unique snapshot key, the standard signal record, version identifier, and data quality results are idempotently written into the strategy signal snapshot table to form a versioned signal snapshot. The factor reuse layer is used to convert continuous scores or raw values in standard signal records into standard factor values according to the selected standardization mode when standard signal records need to be reused as factors. It also generates model output factor records containing model identifier, prediction period, asset identifier, timestamp, factor name and standard factor value, and stores them in the factor table. The traceability record layer is used to persist validity flags, which characterize whether the data quality has passed, and lineage information related to the source, to the strategy signal snapshot table and / or factor table when writing data to the versioned snapshot layer and factor reuse layer; wherein, the validity flag is set according to the data quality results to control whether the downstream consumer can consume the record; The hot read interface layer is used to respond to read requests from different consumer terminals such as front-end display, offline evaluation, simulation operation or model training through a unified application programming interface. It queries and returns the corresponding persistent signal snapshot or factor record from the policy signal snapshot table and / or factor table, so that each consumer terminal can obtain the same result without having to re-execute the signal algorithm.