Historical market data storage system and storage and query method based on file index
Through a historical market data storage system based on file index, the problem of massive historical market data storage and query in the financial industry has been solved, efficient data storage and rapid query are achieved, and the performance of data services has been significantly improved.
Patent Information
- Application Number
- CN202510083778.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-16
AI Technical Summary
The financial industry is facing the problem of storing and querying massive historical market data. Current technologies such as shared memory, K-V databases and relational databases perform poorly in high concurrency and large-scale data scenarios.
A historical quotient data storage system based on file index is adopted, including an environment management module, a data preprocessing module, a data storage module and a data query module. The original data is normalized into fixed-length time series data through data preprocessing and stored in binary blocks, and quickly querying is performed in combination with query hash tables and global indexes.
It significantly improves the storage efficiency and query response speed of historical market data, solves the problems of "difficulty in storage and slow query" in traditional storage solutions, and provides more powerful data service support for the financial industry.
Smart Images

Figure CN120011314A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of financial industry data processing, and in particular to a historical market data storage system based on file indexing and a storage and query method. Background Art
[0002] At present, in the financial securities industry, with the development of the capital market and the increase in the number of listed companies, the scale of historical market data generated daily has shown explosive growth. At the same time, the growth of the investor group has also increased the demand for storage and query of massive historical market data. In a single trading market, the number of trading targets is usually in the tens of thousands. Based on the minute-level market data of each target, about 60,000 records can be generated each year. The total amount of data in the entire market can easily reach billions. For example, the 30-year trading history of the Shanghai and Shenzhen Stock Exchanges has accumulated more than 10 billion data, and the amount of data will further increase exponentially with the expansion of the market and the growth of investor demand. Faced with such a huge data scale, how to achieve low-cost and high-efficiency storage and meet the needs of accurate query and rapid response has become a technical problem that the financial industry urgently needs to break through in the era of big data.
[0003] Current mainstream storage methods such as shared memory, KV databases (such as Redis) and relational databases (such as MySQL) all have limitations: the storage scale of shared memory is limited; KV databases have high memory storage costs and insufficient support for complex queries in large-scale time series data scenarios; relational databases perform poorly when processing tens of billions of time series data under high concurrency requests. Summary of the invention
[0004] In order to help solve the above technical problems, the present application provides a historical market data storage system based on file indexing and a storage and query method.
[0005] In the first aspect, the present application provides a historical market data storage system based on file index, which adopts the following technical solution: A historical market data storage system based on file index, wherein the system comprises: An environment management module, the environment management module is used to maintain the environment information and container information list; A data preprocessing module, which is used to normalize the original market data generated by daily stock market transactions into fixed-length time series data; A data storage module, the data storage module is used to store the time series data into a disk; The data query module is composed of a query hash table and a global index of the historical data of each transaction object to provide fast query and high concurrent access.
[0006] Preferably, the environment management module is used to create an environment management file, and to store a file header, a container name and a container information list in a binary manner.
[0007] Preferably, the data preprocessing module is used to normalize and organize the data in the order of increasing transaction subject ID numbers and increasing transaction times.
[0008] Preferably, the data storage module stores data files in the form of binary blocks, the first data file is used to save the file header, field list, and to maintain a global index table of historical data of the transaction object, and the remaining data files are used to store data blocks to reduce redundant data.
[0009] Preferably, the data query module includes the query hash table and a global index of historical data of each transaction target. The query hash table is stored in the physical memory and organized by the transaction target ID. The Value corresponding to the ID records the first data file address, data update time and hot data table; the global index is located in the first data file of each transaction target and is a sparse index structure.
[0010] In a second aspect, the present application provides a method for storing historical market data based on file indexes, which adopts the following technical solutions: A method for storing historical market data based on file index, wherein the data storage method is performed by the historical market data storage system based on file index according to the first aspect, and the data storage method comprises: Step A1: converting the data to be stored into fixed-length normalized data through the data preprocessing module; Step A2: Based on the normalized data, create an environmental management file for storing environmental information and a container information list; Step A3: Create a data file for storing data and update the container information list in the environment file; Step A4: Initialize the query module; Step A5: Load the historical market data of the transaction target; Step A6: Repeat steps A4 to A5 to complete the loading of historical market data for all trading targets.
[0011] Preferably, the historical market data of the loaded transaction subject includes: Step A51: confirm the transaction subject data to be loaded; Step A52: Write data according to the read-write consistency algorithm.
[0012] Preferably, writing data in accordance with a read-write consistency algorithm includes: Step A521: modify the read / write flag of the data file. Before writing data, first increase the start update flag value by 1; Step A522: Update the data, fill the data block according to the field list information, so that the data is stored incrementally according to the timestamp; Step A523: Update the data file header, global index table, container information list, and hotspot data table according to the modification time and data address; Step A524: After the data update is completed, the update end flag value is increased by 1; Step A525: Update completed.
[0013] In a third aspect, the present application provides a method for querying historical market data based on file index, which adopts the following technical solution: A method for querying historical market data based on file index, wherein the data query method is executed by the historical market data storage system based on file index according to the first aspect, and the data query method comprises: Step B1: construct a query hash table according to the container information list in the environment management file; Step B2: by querying the hash table, find the first data file corresponding to the transaction target, and quickly retrieve the hot data offset position; Step B3: If the query hits the hot data table, in compliance with the read-write consistency algorithm, extract the data according to the offset position of the hot data, and parse and calculate the data content through the field information list; if the hot data is not hit, first read the global index stored in the data file, find the corresponding data offset information in the global index according to the query requirements, and retrieve the data and parse it.
[0014] Preferably, the reading process of the read-write consistency algorithm includes: Step B31: read and record the update start identification value and the update end identification value; Step B32: Determine whether the start update flag value is equal to the update end flag: if they are equal, execute step B33; if they are not equal, return to step B31; Step B33: If the requested data has obtained the offset address in the hotspot data table, check whether the update time recorded in the query hash table is consistent with the update time in the data file header: if the check passes, directly execute step B35; if the check fails, execute step B34; Step B34: confirm the offset through the global index; Step B35: Read required data; Step B36: Read the update start identification value and the update end identification value again; Step B37: Determine whether the update start identification value and the update end identification value are equal to the values recorded in step B31. If they are equal, go to step B38; if they are not equal, go to step B31; Step B38: Reading completed.
[0015] To sum up, the system of the present application solves the problem of "difficult storage and slow query" of historical market data in traditional storage solutions through the design of efficient file structure and query index, combined with the improvement and optimization of data preprocessing, storage module and query module, and significantly improves the storage efficiency and query response speed of historical market data, providing more powerful technical support for data services in the financial industry. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 A schematic block diagram of an embodiment of the environmental management document of the present application; Figure 2 A schematic block diagram of an embodiment of a data file of the present application; Figure 3 A flowchart of an embodiment of a write process of a read-write consistency algorithm of the present application; Figure 4 A flowchart of an embodiment of a read process of a read-write consistency algorithm of the present application; Figure 5 This is a schematic diagram of the index query for this application. DETAILED DESCRIPTION
[0017] The present application is further described below in conjunction with the accompanying drawings, and the structure and principle of the present application are very clear to people in the field. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0018] Figure 1 A schematic block diagram of an embodiment of the environmental management document of the present application; Figure 2 A schematic block diagram of an embodiment of a data file of the present application; Figure 3 A flowchart of an embodiment of a write process of a read-write consistency algorithm of the present application; Figure 4 A flowchart of an embodiment of a read process of a read-write consistency algorithm of the present application; Figure 5 This is a schematic diagram of the index query for this application.
[0019] The present application proposes a historical market data storage system based on file indexing, which includes an environment management module, a data preprocessing module, a data storage module and a data query module. The environment management module is used to maintain environment information and container information lists, and is used to create environment management files for storing file headers, container names and container information lists in binary mode; the data preprocessing module is used to normalize the original market data generated by daily stock market transactions into fixed-length time series data, and is used to normalize the data in the order of increasing transaction target ID numbers and increasing transaction times; the data storage module is used to store time series data to disk and store data files in the form of binary blocks, the first data file is used to save the file header and field list, and is used to maintain the global index table of the historical data of the transaction target, and the remaining data files are used to store data blocks to reduce redundant data; the data query module is composed of a query hash table and a global index of the historical data of each transaction target for fast query and high concurrent access, the data query module includes a query hash table and a global index of the historical data of each transaction target, the query hash table is stored in physical memory, and is based on the transaction target ID The Value corresponding to the organization ID records the address of the first data file, data update time, and hot data table; the global index is located in the first data file of each transaction target and is a sparse index structure.
[0020] The data storage method of the present application comprises the following steps: Step A1: The data preprocessing module generates normalized data.
[0021] Step A2: Create an environmental management file based on the data generated in Step A1.
[0022] Figure 1 This is a schematic diagram of an embodiment of the environmental management file of this application. The specific design of the environmental management file is as follows: A21: Directory Management: (1) Set a basic directory for each environment. The environment management files are stored in the basic directory, and the file names are defined by the user.
[0023] A22: File structure: (1) The environmental management file includes a fixed-length file header and a container information list.
[0024] A23: File header design: The file header stores necessary environment information. For specific fields, refer to the following table: Field Name illustrate File Identifier Environmental management file identifier, indicating that the file type is an environmental management file System version Version number, for example "1.0.0" Total length of file header Environmental management file header length, fixed length Creation time Environmental management file creation time Container list quantity The number of cells in the container information list Container information length The total length of the container information list Start writing identifier Read-write consistency flag (starting value) End write identifier Read-write consistency flag (end value) Reserved seat Reserved space for expansion A24: Container information list: A241: Located after the file header, each container information has a fixed length.
[0025] A242: The container list is stored in order using the container name (transaction target ID) as the keyword.
[0026] A25: Storage format: All content in the environment management file is stored in binary format.
[0027] Step A3: Create a data file to store data, and update the container information list in the environment file.
[0028] Figure 2 This is a schematic block diagram of an embodiment of a data file of the present application. The data file design specification is as follows: A31: File naming and management: Data files are named with numbers starting from 0. A maximum capacity is set for each file. When the capacity is exceeded, a new data file is automatically created with increasing numbers.
[0029] A32: Data file structure: A321: File header: a fixed area that stores file-related information (such as checksum, read / write flags, update time, etc.).
[0030] A322: Field list information: stored only in the first data file, including field name, offset, field type, etc.
[0031] A323: Global index table: stored in the first data file, provides time-based sparse indexing, and flexibly adjusts the time granularity through a hierarchical mechanism to achieve efficient time range queries.
[0032] A324: Data block: Data is stored in chronological order. Each piece of historical market data corresponds to a data block, and the data content strictly corresponds to the field list. Data blocks are arranged in ascending order of timestamps, covering data from the earliest to the latest.
[0033] A33: Storage format: The file header, field list information, and data blocks are stored in binary format.
[0034] Step A4: Initialize the query module.
[0035] A41: Query hash table initialization: The mapping relationship between the transaction target ID and the first data file address is stored in the query hash table.
[0036] A42: Hotspot data table initialization: Fill in regular indexes for the hot data table of the transaction target (the data entries maintained in the hot data table are fixed, except for specific regular indexes that will reside in the address table, the other indexes will be eliminated and updated according to the query situation), and the other indexes of the data table are set to empty during initialization.
[0037] Step A5: Load the historical market data of the transaction target A51: Load data for a single transaction object (see Figure 5 ): The data block corresponding to the target transaction target is located by querying the hash table (for example, 000001.min, that is, the transaction target is the minute-level historical data of 000001. For the convenience of management and query, the present invention can be implemented according to different time granularities of historical market conditions and the combination of transaction targets for storage).
[0038] A52: Write data following the read-write consistency algorithm: A521: Modify the read and write flags of the data file, fill the data block according to the field list information, and ensure that the data is stored in increments according to the timestamp.
[0039] A522: Update the data file header, global index table, container information list, and hot data table based on the modification time, data address, and other information.
[0040] A523: Figure 3 This is a flow chart of an embodiment of the writing process of the read-write consistency algorithm of the present application. The specific steps of the process of writing A521 and A522 data in accordance with the read-write consistency are as follows: Step A521: modify the read / write flag of the data file. Before writing data, first increase the start update flag value by 1; Step A522: Update the data, fill the data block according to the field list information, so that the data is stored incrementally according to the timestamp; Step A523: Update the data file header, global index table, container information list, and hotspot data table according to the modification time and data address; Step A524: After the data update is completed, the update end flag value is increased by 1; Step A525: Update completed.
[0041] Step A6: Repeat steps A4 to A5 to complete the loading of historical market data for all trading targets.
[0042] The data query method of the present application comprises the following steps: Step B1: construct a query hash table according to the container information list in the environment management file; Step B2: by querying the hash table, find the first data file corresponding to the transaction target, and quickly retrieve the hot data offset position; Step B3: If the query hits the hot data table, in compliance with the read-write consistency algorithm, extract the data according to the offset position of the hot data, and parse and calculate the data content through the field information list; if the hot data is not hit, first read the global index stored in the data file, find the corresponding data offset information in the global index according to the query requirements, and retrieve the data and parse it.
[0043] Specifically, it can include parsing the query request and confirming the range of historical market data to be obtained according to the request. Look up the query hash table, specifically locate the first data file, data modification time, and hot data table of the transaction target in the query hash table. And try to hit the requested data in the hot data table to obtain the relative offset position of the data. Then locate the data file and read the data. During the data reading process, according to the read-write consistency algorithm, check whether the read-write flag in the data file meets the reading conditions. If it meets the conditions, read directly, otherwise wait.
[0044] Figure 4 This is a flowchart of an embodiment of the reading process of the read-write consistency algorithm of the present application, and the specific steps are: Step B31: read the start update identification value (startSeq) and the end update identification value (endSeq) and record them; Step B32: Determine whether startSeq is equal to endSeq: if they are equal, go to step B33; If they are not equal, return to step B31: Step B33: If the requested data has obtained the offset address in the hot data table, check whether the update time recorded in the query hash table is consistent with the update time in the data file header: if the check is passed, directly jump to step B35; if the check fails, go to step B34; Step B34: confirm the offset through the global index; Step B35: Read required data; Step B36: read startSeq and endSeq again; Step B37: Determine whether startSeq and endSeq are equal to the values recorded in step B31. If they are equal, go to step B38; if they are not equal, go to step B31; Step B38: Reading completed.
[0045] In addition, the method also includes updating the hotspot data table of the query hash table, as follows: 1. If the index expires, initialize the hot data table corresponding to the transaction target and update the timestamp synchronously.
[0046] 2. Eliminate and update the indexes in the hot data table.
Claims
1. A historical market data storage system based on file index, characterized in that: The system comprises: An environment management module, the environment management module is used to maintain the environment information and container information list; A data preprocessing module, which is used to normalize the original market data generated by daily stock market transactions into fixed-length time series data; A data storage module, the data storage module is used to store the time series data into a disk; The data query module is composed of a query hash table and a global index of the historical data of each transaction object to provide fast query and high concurrent access.
2. The historical market data storage system based on file index according to claim 1 is characterized in that: The environment management module is used to create an environment management file and to store a file header, a container name and a container information list in a binary manner.
3. The historical market data storage system based on file index according to claim 2 is characterized in that: The data preprocessing module is used to normalize and organize the data in the order of increasing transaction target ID number and increasing transaction time.
4. The historical market data storage system based on file index according to claim 3 is characterized in that: The data storage module stores data files in the form of binary blocks. The first data file is used to save the file header, field list, and to maintain the global index table of historical data of the transaction object. The remaining data files are used to store data blocks to reduce redundant data.
5. The historical market data storage system based on file index according to claim 4 is characterized in that: The data query module includes the query hash table and the global index of the historical data of each transaction target. The query hash table is stored in the physical memory and organized according to the transaction target ID. The Value corresponding to the ID records the first data file address, data update time and hot data table; the global index is located in the first data file of each transaction target and is a sparse index structure.
6. A method for storing historical market data based on file index, characterized in that: The data storage method is performed by the file index-based historical market data storage system according to claim 5, and the data storage method includes: Step A1: converting the data to be stored into fixed-length normalized data through the data preprocessing module; Step A2: creating an environmental management file for storing environmental information and a container information list based on the standardized data; Step A3: Create a data file for storing data and update the container information list in the environment file; Step A4: Initialize the query module; Step A5: Load the historical market data of the transaction target; Step A6: Repeat steps A4 to A5 to complete the loading of historical market data for all trading targets.
7. The data storage method according to claim 6, characterized in that: The historical market data of the loaded transaction subject includes: Step A51: confirm the transaction subject data to be loaded; Step A52: Write data according to the read-write consistency algorithm.
8. The data storage method according to claim 7, characterized in that: Writing data in accordance with the read-write consistency algorithm includes: Step A521: modify the read / write flag of the data file. Before writing data, first increase the start update flag value by 1; Step A522: Update the data, fill the data block according to the field list information, so that the data is stored incrementally according to the timestamp; Step A523: Update the data file header, global index table, container information list, and hotspot data table according to the modification time and data address; Step A524: After the data update is completed, the update end flag value is increased by 1; Step A525: Update completed.
9. A method for querying historical market data based on file index, characterized in that: The data query method is executed by the historical market data storage system based on file index according to claim 5, and the data query method includes: Step B1: construct a query hash table according to the container information list in the environment management file; Step B2: by querying the hash table, find the first data file corresponding to the transaction target, and quickly retrieve the hot data offset position; Step B3: If the query hits the hot data table, in compliance with the read-write consistency algorithm, extract the data according to the offset position of the hot data, and parse and calculate the data content through the field information list; if the hot data is not hit, first read the global index stored in the data file, find the corresponding data offset information in the global index according to the query requirements, and retrieve the data and parse it.
10. The data query method according to claim 9, characterized in that: The reading process of the read-write consistency algorithm includes: Step B31: read and record the update start identification value and the update end identification value; Step B32: Determine whether the start update flag value is equal to the update end flag: if they are equal, execute step B33; if they are not equal, return to step B31; Step B33: If the requested data has obtained the offset address in the hotspot data table, check whether the update time recorded in the query hash table is consistent with the update time in the data file header: if the check passes, directly execute step B35; if the check fails, execute step B34; Step B34: confirm the offset through the global index; Step B35: Read required data; Step B36: Read the update start identification value and the update end identification value again; Step B37: Determine whether the update start identification value and the update end identification value are equal to the values recorded in step B31. If they are equal, go to step B38; if they are not equal, go to step B31; Step B38: Reading completed.