Strategy service-based training data generation method and system and storage medium
Through the training data generation method based on strategy services, the greedy algorithm is used to synthesize multi-dimensional K-line data of financial data in real time, solving the problems of large storage space, long update time and lack of flexibility in the existing technology, and achieving efficient and flexible data processing and fast data access.
Patent Information
- Application Number
- CN202510164496.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-06
AI Technical Summary
When dealing with the multi-dimensional K-line data demand, existing financial data processing methods have problems such as large storage space, long data update time, lack of flexibility and low model training efficiency.
The training data generation method based on strategy services is adopted, by saving the K-line data of the basic frequency, and using greedy algorithms to select sub-frequency combinations from the ordered decreasing array, the K-line data of the target frequency is synthesized in real time, and the newly generated data is updated to the cache.
It significantly saves storage space, improves data processing efficiency and flexibility, shortens data access response time, and supports the calculation of complex data items to meet the multi-dimensional data needs of financial model training.
Smart Images

Figure CN120106883A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of thin-walled parts processing, and in particular to a training data generation method, system and storage medium based on strategy services. Background Art
[0002] In the quantitative analysis and strategy development of the financial market, K-line data (also known as candlestick chart data) is one of the core data forms. K-line data contains key information such as opening price, highest price, lowest price, closing price, etc., and is an important basis for evaluating market trends and formulating trading strategies. In the training of financial models, K-line data of different time frequencies are usually required, such as 1 minute, 5 minutes, 15 minutes, 1 hour, daily frequency, etc.
[0003] However, existing financial data processing methods have many shortcomings in dealing with multi-dimensional K-line data needs. On the one hand, in order to cover all possible frequency requirements, it is often necessary to pre-store a large amount of K-line data with different frequencies, which not only takes up a lot of storage space, but also takes a lot of time and resources when updating data. On the other hand, when a new data dimension is needed, existing methods often need to re-encode and generate data, which lacks flexibility. In addition, frequently reading data from disk also greatly reduces the efficiency of model training. Therefore, we propose a training data generation method, system and storage medium based on policy services. Summary of the invention
[0004] The purpose of the present invention is to provide a training data generation method, system and storage medium based on policy service, which solves the problems raised in the background technology.
[0005] To achieve the above object, the present invention provides the following technical solution: a training data generation method based on policy service, comprising the following method steps:
[0006] Step 1: Save the K-line data of the basic frequency, which includes 1 second, 1 minute, 1 hour and daily frequency data;
[0007] Step 2: Represent the K-line data of each basic frequency in a unified dimension, where the dimension of 1 second frequency is 1, the dimension of 1 minute frequency is 60, the dimension of 1 hour frequency is 3600, and the dimension of daily frequency is 14400;
[0008] Step 3: Maintain an ordered decreasing array F containing the existing K-line frequency dimensions;
[0009] Step 4: According to the target frequency N, a sub-frequency combination is selected from the array F through a greedy algorithm to generate an ordered non-increasing array f, where the sub-frequency combination satisfies ∑f[i]=N;
[0010] Step 5: Based on the array f and the cached K-line data, synthesize the K-line data of the target frequency in real time, and update the newly generated K-line data to the cache.
[0011] As a preferred implementation of the present invention, the greedy algorithm in step 4 specifically includes:
[0012] In each iteration, the first maximum frequency t that is less than or equal to the current residual value is selected from array F, t is added to array f, and the residual value is updated to N=Nt until the residual value is 0.
[0013] As a preferred embodiment of the present invention, the K-line data of the synthetic target frequency in step 5 includes:
[0014] According to the K-line timestamps corresponding to each sub-frequency in the array f, the opening price, highest price, lowest price, closing price, trading volume, transaction amount and derived indicators are calculated by formula.
[0015] As a preferred implementation of the present invention, the derivative indicators include volume-weighted average price and moving average price, and the calculation formulas thereof are respectively:
[0016]
[0017] Where g(i) is the timestamp offset function, defined as:
[0018]
[0019] The present invention also relates to a training data generation system based on policy services, comprising:
[0020] Basic data storage module, used to store K-line data of basic frequency;
[0021] Dimension index module, used to maintain the ordered decreasing array F and generate the sub-frequency combination array f;
[0022] Real-time aggregation module, dynamically synthesizes the K-line of the target frequency based on the array f and cache data;
[0023] The cache update module updates the newly generated K-line data to the memory cache.
[0024] As a preferred embodiment of the present invention, the real-time aggregation module supports the synthesis of complex data items including volume-weighted average price and moving average price.
[0025] The present invention also relates to a computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement the method according to any one of claims 1 to 4.
[0026] Compared with the prior art, the present invention has the following beneficial effects:
[0027] The present invention abandons the traditional practice of pre-storing a large amount of K-line data of different frequencies, and instead only stores the K-line data of the basic frequency. This transformation not only greatly saves storage space and reduces storage costs, but also avoids maintenance complexity and data update delays that may be caused by data redundancy;
[0028] The present invention can quickly and accurately select the appropriate sub-frequency combination from the stored basic frequency data to synthesize the K-line data of the target frequency in real time. In this process, there is no need for cumbersome data conversion or re-encoding, which greatly improves the efficiency of data processing. At the same time, since the data of any target frequency can be synthesized in real time, the present invention also significantly improves the flexibility of data processing, allowing users to quickly obtain the data of the required frequency as needed without worrying about data loss or delay;
[0029] The present invention not only supports the synthesis of basic data items such as opening price, highest price, lowest price, closing price, trading volume, and transaction amount, but also extends to the calculation of complex data items such as volume-weighted average price (VWAP), moving average price (MA), etc. This feature enables the present invention to fully meet the needs of multi-dimensional data in financial model training, and provides richer and more accurate data support for quantitative analysis and strategy development;
[0030] By updating the newly generated K-line data to the memory cache, the present invention further improves the speed of data access. When the K-line data of a certain frequency needs to be accessed, the system can directly read it from the cache without synthesizing it again or accessing the disk, thereby significantly shortening the response time of data access.
[0031] The system architecture of the present invention is clear, the responsibilities of each module are clearly defined, and it is easy to expand and maintain. With the continuous development of the financial market and the continuous changes in user needs, the system can easily add new basic frequency data or synthesis algorithms to adapt to new application scenarios and data needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:
[0033] Figure 1 It is a schematic diagram of the process of the present invention;
[0034] Figure 2 It is a project processing flow chart of the present invention. DETAILED DESCRIPTION
[0035] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the present invention is further explained below in conjunction with specific implementation methods.
[0036] A training data generation method based on a policy service includes the following steps:
[0037] Step 1: Save the K-line data of the basic frequency, which includes 1 second, 1 minute, 1 hour and daily frequency data;
[0038] Step 2: Represent the K-line data of each basic frequency in a unified dimension, where the dimension of 1 second frequency is 1, the dimension of 1 minute frequency is 60, the dimension of 1 hour frequency is 3600, and the dimension of daily frequency is 14400;
[0039] Step 3: Maintain an ordered decreasing array F containing the existing K-line frequency dimensions;
[0040] Step 4: According to the target frequency N, a sub-frequency combination is selected from the array F through a greedy algorithm to generate an ordered non-increasing array f, where the sub-frequency combination satisfies ∑f[i]=N;
[0041] Step 5: Based on the array f and the cached K-line data, the K-line data of the target frequency is synthesized in real time, and the newly generated K-line data is updated to the cache;
[0042] The greedy algorithm described in step 4 specifically includes:
[0043] In each round of iteration, the first maximum frequency t that is less than or equal to the current residual value is selected from the array F, t is added to the array f, and the residual value is updated to N=Nt until the residual value is 0. The K-line data of the synthetic target frequency in step 5 includes:
[0044] According to the K-line timestamps corresponding to each sub-frequency in the array f, the opening price, highest price, lowest price, closing price, trading volume, transaction amount and derived indicators are calculated by formula.
[0045] The following is the algorithm formula of the present invention:
[0046] Assume that the timestamp of the K-line is T. To conveniently represent the timestamp, set the function g as:
[0047]
[0048] (1) Opening price
[0049]
[0050] (2) Highest Price
[0051]
[0052] (3) Closing price
[0053]
[0054] (4) Lowest price
[0055]
[0056] (5) Trading volume
[0057]
[0058] (6) Transaction volume
[0059]
[0060] (7) vwap (Volume Weighted Average Price, volume weighted average price)
[0061]
[0062] (8) man (Moving Average, the moving average price of the closing prices of n candlesticks, which can actually be MA5, MA20, etc.)
[0063]
[0064] A training data generation system based on policy service, characterized by comprising:
[0065] Basic data storage module, used to store K-line data of basic frequency;
[0066] Dimension index module, used to maintain the ordered decreasing array F and generate the sub-frequency combination array f;
[0067] Real-time aggregation module, dynamically synthesizes the K-line of the target frequency based on the array f and cache data;
[0068] Cache update module, updates the newly generated K-line data to the memory cache;
[0069] The real-time aggregation module supports the synthesis of complex data items including volume-weighted average price and moving average price.
[0070] A computer-readable storage medium stores computer program instructions, and when the instructions are executed by a processor, the above method steps are implemented.
[0071] Example
[0072] Take the synthetic 12-minute K-line (N=720) as an example:
[0073] 1. Select sub-frequency combination f = [300, 300, 60, 60] from F = [1, 60, 300];
[0074] 2. According to the timestamp offset function g(i), locate the start time of each sub-K line;
[0075] 3. Aggregate basic data items such as opening price and highest price according to the formula, and calculate VWAP and MA;
[0076] 4. Store the generated 12-minute K-line into the cache to reduce subsequent repeated calculations.
[0077] The above shows and describes the basic principles and main features of the present invention and the advantages of the present invention. It is obvious to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic features of the present invention. Therefore, no matter from which point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description, and it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims are included in the present invention.
[0078] In addition, it should be understood that although the present specification is described according to implementation modes, not every implementation mode contains only one independent technical solution. This description of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment may also be appropriately combined to form other implementation modes that can be understood by those skilled in the art.
Claims
1. A training data generation method based on policy service, characterized in that: The method comprises the following steps: Step 1: Save the K-line data of the basic frequency, which includes 1 second, 1 minute, 1 hour and daily frequency data; Step 2: Represent the K-line data of each basic frequency in a unified dimension, where the dimension of 1 second frequency is 1, the dimension of 1 minute frequency is 60, the dimension of 1 hour frequency is 3600, and the dimension of daily frequency is 14400; Step 3: Maintain an ordered decreasing array F containing the existing K-line frequency dimensions; Step 4: According to the target frequency N, a sub-frequency combination is selected from the array F through a greedy algorithm to generate an ordered non-increasing array f, where the sub-frequency combination satisfies ∑f[i]=N; Step 5: Based on the array f and the cached K-line data, synthesize the K-line data of the target frequency in real time, and update the newly generated K-line data to the cache.
2. The method for generating training data based on policy service according to claim 1, characterized in that: The greedy algorithm described in step 4 specifically includes: In each iteration, the first maximum frequency t that is less than or equal to the current residual value is selected from array F, t is added to array f, and the residual value is updated to N=Nt until the residual value is 0.
3. The method for generating training data based on policy service according to claim 1, characterized in that: The K-line data of the synthetic target frequency described in step 5 include: According to the K-line timestamps corresponding to each sub-frequency in the array f, the opening price, highest price, lowest price, closing price, trading volume, transaction amount and derived indicators are calculated through formulas.
4. The method for generating training data based on policy service according to claim 1, characterized in that: The derivative indicators include volume-weighted average price and moving average price, and their calculation formulas are: Where g(i) is the timestamp offset function, defined as:
5. A training data generation system based on policy service, characterized in that: include: Basic data storage module, used to store K-line data of basic frequency; Dimension index module, used to maintain the ordered decreasing array F and generate the sub-frequency combination array f; Real-time aggregation module, dynamically synthesizes the K-line of the target frequency based on the array f and cache data; The cache update module updates the newly generated K-line data to the memory cache.
6. The training data generation system based on policy service according to claim 5, characterized in that: The real-time aggregation module supports the synthesis of complex data items including volume-weighted average price and moving average price.
7. A computer-readable storage medium, characterized in that: Computer program instructions are stored, and when the instructions are executed by a processor, the method according to any one of claims 1 to 4 is implemented.