Financial market data-oriented distributed memory database system and implementation method thereof

By adopting hierarchical data storage, table-level multi-version concurrency control and shared memory persistence technology in the in-memory database system, combined with lock-free hashmap and chain array storage structure, the problems of high concurrent read and write performance bottlenecks and persistence delays in financial market data processing are solved, efficient read and write performance and fast persistence are achieved, and large-scale real-time data distribution is supported.

CN120179649APending Publication Date: 2025-06-20WIND INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510303099.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

When processing financial market data, existing in-memory databases have problems such as high concurrent read and write performance bottlenecks, persistence delay, limited large-scale distributed scaling capabilities, and insufficient support for hierarchical timing data.

Method used

A distributed memory database system is designed, using a hierarchical data storage module to store financial market data through a tree table structure, and using a table-level multi-version concurrent control module and a shared memory persistence module, combining a lock-free hashmap and chain array storage structure to achieve high concurrent read and write and fast persistence. At the same time, the real-time data distribution module realizes large-scale real-time data distribution through TCP control channels and multicast mechanisms.

Benefits of technology

It significantly improves the system's read and write performance, supports millions of accesses per second, realizes millisecond data persistence, and supports large-scale real-time data distribution, perfectly adapts to the hierarchical storage needs of financial market data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179649A_ABST
    Figure CN120179649A_ABST
Patent Text Reader

Abstract

According to the financial market data-oriented distributed memory database system and the implementation method thereof, a hierarchical data storage module and a table-level multi-version concurrency control module establish indexes by constructing a tree table structure and using a lock-free hashmap to finally generate a chain array; the problem that the concurrent processing capacity is affected by lock competition caused by limitation of a single-thread model is solved, a hierarchical data storage module and a shared memory persistence module store binlogs through a tree table structure and a shared memory, and the problem that data is not persistent due to storage of key values is solved; the real-time data distribution module solves the problem that the distributed expansion capability is limited due to point-to-point synchronization through a multicast and TCP hybrid transmission mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and particularly to a distributed in-memory database system for financial market data and its implementation method. Background Art

[0002] With the development of the financial market, financial market data shows characteristics such as large data volume, frequent updates, and high access concurrency. Traditional disk databases are difficult to meet these requirements due to I / O bottlenecks, while existing in-memory databases, although having improved performance, still have the following limitations:

[0003] 1. Analysis of existing technologies.

[0004] Redis: Although it supports multiple data types, its single-threaded model limits its concurrent processing ability;

[0005] Memcached: Only supports simple key-value pair storage and does not support data persistence;

[0006] Other in-memory databases: Generally have problems such as read-write performance bottlenecks and high persistence latency.

[0007] 2. Deficiencies of existing technologies.

[0008] The read-write performance cannot meet the access requirements of millions of times per second;

[0009] The persistence mechanism causes high latency;

[0010] The distributed expansion ability is limited;

[0011] Insufficient support for hierarchical time-series data. Summary of the Invention

[0012] The objective of the technical solution of the present invention is to design an in-memory database system and its implementation method that can simultaneously meet the requirements of high-concurrency read-write, fast persistence, real-time distribution, and hierarchical data storage.

[0013] The technical solution of the present invention provides an implementation method for a distributed in-memory database for financial market data, including the following steps:

[0014] The hierarchical data storage module creates a main table object, adds column information of the main table one by one, adds a first sub-table column to the main table object, returns the first empty sub-table object, adds column information of the first sub-table column one by one to the empty sub-table object to complete the definition of the first sub-table, adds a second sub-table column to the sub-table object, returns the second empty sub-table object, and adds column information of the second sub-table column one by one to the second empty sub-table object to complete the definition of the second sub-table, until the table nested hierarchical data table stored in the form of a tree corresponding to all financial market data is completed;

[0015] The table-level multi-version concurrency control module initializes and generates the first initialization memory block for the table-nested hierarchical data table according to a preset size. When adding records, the appended part is added to the initialization memory block, and at the same time, an index is established using a lock-free hashmap. When the first initialization memory block is full, a second initialization memory block with a size of the (exponent + 1)th power of the preset size is allocated and chained behind the first initialization memory block until it is chained behind the third initialization memory block with an exponent equal to the predefined exponent. Each subsequent block has the same size as the third initialization memory block to construct an innovative chained array storage structure;

[0016] When there are too many holes in the initialization memory block and the memory utilization rate is low due to deleting records without moving any memory or other operations, the table version switching logic is triggered to create a new table copy, and the data in the initialization memory block is reorganized and written into the new table copy;

[0017] The shared memory persistence module manages the entire shared memory in a page table manner to generate a log storage area, and generates binlog log data for the processing of records by the table-level multi-version concurrency control module. The binlog log data is continuously stored in the log storage area in an appended manner;

[0018] When the data in the log area meets one of the preset conditions that the log is greater than the preset quantity, the log exceeds N times the snapshot data, or the time reaches a certain threshold, the snapshot operation is triggered. The snapshot operation is to replay a certain number of logs based on the previous snapshot, replace the old snapshot data, and delete the replayed logs after the replacement is successful;

[0019] The real-time data distribution module establishes a TCP control channel at the downstream node, synchronizes the hierarchical data table, database identifier in the chained array, sequence number of the current binlog log data, schema information, and multicast configuration parameters through the TCP channel. The downstream node accesses the multicast according to the multicast configuration parameters and starts receiving multicast packets. After receiving the first multicast packet, the log event number of the multicast packet is updated, and the established packet compensation channel completes the data before the current multicast packet. Then, the multicast event is replayed on this basis to achieve synchronous update, and packet loss detection is performed by detecting the sequence number and log number in each multicast packet. If packet loss is found, compensation is performed through the TCP channel.

[0020] Preferably, the column information includes information such as column name, column data type, and size.

[0021] Preferably, the preset size is 2 to the Nth power of records.

[0022] The technical solution of the present invention also provides a distributed in-memory database system for financial market data, which adopts the implementation method of the distributed in-memory database for financial market data as described above. The distributed in-memory database system for financial market data includes:

[0023] The hierarchical data storage module is used to convert financial market data into a nested hierarchical data table stored in the form of a tree to implement a multi-level table structure design;

[0024] The table-level multi-version concurrency control module is used to initialize and generate a first initialization memory block for the nested hierarchical data table according to a preset size. When adding records, the appended part is added to the initialization memory block, and at the same time, an index is established using a lock-free hashmap. When the first initialization memory block is full, a second initialization memory block with a size of the (exponent + 1)th power of the preset size is allocated and chained behind the first initialization memory block until it is chained behind the third initialization memory block with an exponent equal to the predefined exponent. Each subsequent block has the same size as the third initialization memory block to construct an innovative chained array storage structure; when there are too many holes in the initialization memory block and the memory utilization rate is low due to deleting records without moving any memory or other operations, the table version switching logic is triggered to create a new table copy, and the data in the initialization memory block is reorganized and written into the new table copy;

[0025] The shared memory persistence module is used to manage the entire shared memory in the form of a page table to generate a log storage area, and generate binlog log data for the processing of records by the table-level multi-version concurrency control module. The binlog log data is continuously stored in the log storage area in an appended manner; when the data in the log area meets one of the preset conditions that the log is greater than the preset quantity, the log exceeds N times the snapshot data, or the time reaches a certain threshold, a snapshot operation is triggered. The snapshot operation is to replay a certain number of logs based on the previous snapshot to replace the old snapshot data. After the replacement is successful, the replayed logs are deleted;

[0026] The real-time data distribution module is used to establish a TCP control channel at the downstream node, synchronize the hierarchical data table, database identifiers in the chained array, the sequence number of the current binlog log data, schema information, and multicast configuration parameters through the TCP channel. The downstream node accesses the multicast according to the multicast configuration parameters and starts receiving multicast packets. After receiving the first multicast packet, the log event number of the multicast packet is updated, and the established packet compensation channel completes the data before the current multicast packet, and then starts playing back the multicast event on this basis to achieve synchronous update, and performs packet loss detection by detecting the sequence number and log number in each multicast packet. If packet loss is found, compensation is performed through the TCP channel.

[0027] The technical solution of the present invention proposes a distributed in-memory database system for financial market data and its implementation method. The hierarchical data storage module and the table-level multi-version concurrency control module establish an index by constructing a tree table structure and using a lock-free hashmap to finally generate a chained array, which solves the problem that the competition of locks caused by the single-threaded model limit affects the concurrency processing ability. The hierarchical data storage module and the shared memory persistence module solve the problem of data non-persistence caused by key-value pair storage through the tree table structure and the shared memory to store binlog logs. The real-time data distribution module solves the problem of limited distributed expansion ability caused by point-to-point synchronization through the multicast and TCP hybrid transmission mechanism. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 is the hierarchical structure of the database;

[0029] Figure 2 is the chained array;

[0030] Figure 3 is the data shared memory address space. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.

[0032] An implementation method of a distributed in-memory database for financial market data includes the following steps:

[0033] As Figure 1 shown, the hierarchical data storage module creates a main table object, adds the column information of the main table one by one, adds the first sub-table column to the main table object, returns the first empty sub-table object, and adds the column information of the first sub-table column to the empty sub-table object one by one to complete the definition of the first sub-table (primary sub-table). Add the second sub-table column to the sub-table object, return the second empty sub-table object, and add the column information of the second sub-table column to the second empty sub-table object one by one to complete the definition of the second sub-table (secondary sub-table), until the nested hierarchical data table in the form of a tree corresponding to all financial market data is completed. The column information includes column name, column data type, size and other information. The hierarchical data table is used as metadata and is used when reading and writing data subsequently. For the sub-tables of the same column, the table structures must be the same, and the hierarchical data storage module supports infinite-level table nesting.

[0034] As Figure 2As shown in the figure, the table-level multi-version concurrency control module initializes and generates the first initialized memory block for the table-nested hierarchical data table according to a preset size. The preset size should be 2 to the power of N records. When adding records, the appended part is added to the initialized memory block, and at the same time, a lock-free hashmap is used to establish an index. When the first initialized memory block is full, a second initialized memory block with a size of 2 to the power of N+1 is allocated and chained behind the first initialized memory block until it is chained behind the third initialized memory block with a predefined size of 2 to the power of M. Each subsequent block has the same size as the third initialized memory block to construct an innovative chained array storage structure.

[0035] When deleting records, mark-and-delete is adopted, which simply marks this record as invalid without involving any memory movement. When there are too many memory holes due to deletions or other operations and the memory utilization rate is low, a table version switching logic is triggered. First, a new table copy is created, then the data of the old table is reorganized and written into the new table copy, and finally the version is updated to make the new version visible. All subsequent users use the new version table, and the old version table will be released after the reference count reaches 0.

[0036] Under such an allocation strategy, to locate the Index-th record, the following algorithm can be used to achieve a time complexity of O(1), thus greatly improving the random read efficiency.

[0037] 1) uint32_t b = 32 - __builtin_clz((index >> n) + 1) - 1;

[0038] 2) If b is less than m, the record is in the b-th block, and the specific position in the b-th block is index - (1 << (b + n)) + (1 << n);

[0039] 3) If b is greater than or equal to m, the record is in the (index - (((1 << m) - 1) << n)) / (1 << (m + n - 1))-th block, and the specific position in the b-th block is (index - (((1 << m) - 1) << n)) % (1 << (m + n - 1)).

[0040] To improve the read and write performance of the in-memory database, the competition of locks must be reduced, and locks should be used as little as possible or even completely avoided. The table-level multi-version concurrency control module adopts the table-level multi-version concurrency control technology, abandons strong consistency, and only satisfies eventual consistency. The write operation of data is locked to ensure data correctness, but the read operation is not locked. The final read consistency is satisfied by quickly pushing after data changes. At the same time, a chained array is used to avoid frequent memory movement and improve the reading performance.

[0041] Such as Figure 3As shown, the shared memory persistence module manages the entire shared memory in the form of a page table to generate a log storage area. Each database has a total of 64 page entries, each page entry has 32 pages, and the size of each page is 64MB. All subsequent paging operations are in units of 64MB, that is, the design concept of huge pages is used here. The shared memory adopts a dual-region design, and the virtual address space is divided into 2 memory heaps. One is the log data area, which uses page entries 0 - 31, and the other is the snapshot data area, which uses page entries 32 - 63.

[0042] The shared memory persistence module generates binlog log data from the processing of records by the table-level multi-version concurrency control module. The binlog log data is continuously stored in the log storage area in an append manner. When the data in the log area meets the preset conditions, a snapshot operation will be triggered. The preset conditions are that the log is greater than the preset quantity, the log exceeds N times the snapshot data, or the time reaches a certain threshold, whichever is met first. The snapshot operation is based on the previous snapshot, replays a certain number of 64MB logs, and then replaces the old snapshot data. After the replacement is successful, the replayed logs will be deleted. Both the log area and the snapshot area will be used in a circular rolling manner.

[0043] This embodiment uses shared memory to store binlog logs, manages the log memory through a set of shared memory management frameworks, and ensures that the memory is not overly occupied through a regular full-database snapshot mechanism.

[0044] The real-time data distribution module establishes a TCP control channel at the downstream node, synchronizes the hierarchical data table, the database identifier in the chained array, the sequence number of the current binlog log data, the schema information, and the multicast configuration parameters through the TCP channel. The downstream node accesses the multicast according to the multicast configuration parameters and starts receiving multicast packets. After receiving the first multicast packet, it updates the log event number of the multicast packet, and the established retransmission channel completes the data before the current multicast packet. Then, it starts replaying the multicast events on this basis to achieve synchronous update, and performs packet loss detection by detecting the sequence number and log number in each multicast packet. If packet loss is found, compensation is performed through the TCP channel.

[0045] A distributed in-memory database system for financial market data includes:

[0046] A hierarchical data storage module for converting financial market data into a table-nested hierarchical data table stored in the form of a tree to achieve a multi-level table structure design.

[0047] The table-level multi-version concurrency control module is used to initialize and generate the first initialized memory block from the table nested hierarchical data table stored in the form of a tree according to a preset size. The preset size should be 2 to the power of N records. When adding records, the appended part is added to the initialized memory block, and at the same time, a lock-free hashmap is used to establish an index. When the first initialized memory block is full, a second initialized memory block with a size of 2 to the power of N+1 is allocated and chained behind the first initialized memory block until it is chained behind the third initialized memory block with a predefined size of 2 to the power of M. Each subsequent block has the same size as the third initialized memory block to construct an innovative chained array storage structure. At the same time, an intelligent version switching mechanism is adopted. When there are too many memory holes due to deletion or other operations and the memory utilization rate is not high, a table version switching logic is triggered. First, a new table copy is created, then the data of the old table is reorganized and written into the new table copy, and finally the version is updated to make the new version visible. All subsequent users use the new version table, and the old version table will be released after the reference count reaches 0.

[0048] The shared memory persistence module is used to generate binlog log data from the processing of records by the table-level multi-version concurrency control module and store it in the log storage area managed in a page table manner in the shared memory. When the data in the log area meets the preset conditions, a snapshot operation will be triggered. The preset conditions are that when the log is greater than the preset quantity, the log exceeds N times the snapshot data, or the time reaches a certain threshold. The snapshot operation is based on the previous snapshot, replays a certain number of 64MB logs, and then replaces the old snapshot data. After the replacement is successful, the replayed logs will be deleted. Both the log area and the snapshot area will be used in a circular rolling manner.

[0049] The real-time data distribution module is used to establish a TCP control channel at the downstream node and synchronize the hierarchical data table, the database identifier in the chained array, the sequence number of the current binlog log data, the schema information, and the multicast configuration parameters through the TCP channel. The downstream node accesses the multicast according to the multicast configuration parameters and starts receiving multicast packets. After receiving the first multicast packet, the log event number of the multicast packet is updated, and the established retransmission channel supplements the data before the current multicast packet, and then starts playing back the multicast events on this basis to achieve synchronous update. And by detecting the sequence number and log number in each multicast packet, packet loss detection is performed. If packet loss is found, compensation is performed through the TCP channel.

[0050] The embodiments of the present invention have the following beneficial effects:

[0051] 1. Significantly improves the system read and write performance and supports millions of accesses per second;

[0052] 2. Achieves millisecond-level data persistence;

[0053] 3. Supports large-scale real-time data distribution;

[0054] 4. Perfectly adapt to the hierarchical storage requirements of financial market data.

Claims

1. A method for implementing a distributed memory database for financial market data, characterized in that: The following steps are involved: The hierarchical data storage module creates a main table object, adds column information of the main table one by one, adds the first sub-table column to the main table object, returns the first empty sub-table object, adds the column information of the first sub-table column one by one to the empty sub-table object to complete the definition of the first sub-table, adds the second sub-table column to the sub-table object, returns the second empty sub-table object, adds the column information of the second sub-table column one by one to the second empty sub-table object to complete the definition of the second sub-table, until the table nesting hierarchical data table corresponding to all financial market data stored in the form of a tree is completed; The table-level multi-version concurrency control module initializes the nested hierarchical data table according to the preset size to generate the first initialization memory block. When adding records, the additional part is added to the initialization memory block, and the index is established using a lock-free hashmap. When the first initialization memory block is full, the second initialization memory block with the preset index plus 1 is allocated and linked to the first initialization memory block until the third initialization memory block with the same index as the pre-defined index is linked. Each subsequent block is equal to the size of the third initialization memory block to build an innovative chain array storage structure. When there are too many holes in the initialization memory block and the memory utilization is low due to deleting records without moving any memory or other operations, the table version switching logic is triggered, a new table copy is created, and the data in the initialization memory block is reorganized and written to the new table copy; The shared memory persistence module manages the entire shared memory in a page table manner to generate a log storage area, and generates binlog log data by processing the records in the table-level multi-version concurrency control module. The binlog log data is continuously stored in the log storage area in an appended manner. When the log area data meets one of the preset conditions, that is, the number of logs is greater than the preset number, or the number of logs exceeds N times the snapshot data, or the time reaches a certain threshold, the snapshot operation is triggered. The snapshot operation is based on the last snapshot, replaying a certain number of logs to replace the old snapshot data. After the replacement is successful, the replayed logs are deleted; The real-time data distribution module establishes a TCP control channel in the downstream node, and synchronizes the hierarchical data table, the database identifier in the linked array, the serial number of the current binlog log data, the schema information and the multicast configuration parameters through the TCP channel. The downstream node accesses the multicast according to the multicast configuration parameters and starts to receive the multicast packet. After receiving the first multicast packet, the log event number of the multicast packet is updated, and the established packet replenishment channel is used to replenish the data before the current multicast packet. Then, on this basis, the multicast event is played back to achieve synchronous update, and packet loss detection is performed by detecting the serial number and log number in each multicast packet. If packet loss is found, compensation is performed through the TCP channel.

2. The method for implementing a distributed memory database for financial market data according to claim 1, characterized in that: The column information includes information such as column name, column data type and size.

3. The method for implementing a distributed memory database for financial market data according to claim 1, characterized in that: The preset size is 2 to the power of N records.

4. A distributed memory database system for financial market data, characterized by: The distributed memory database system for financial market data according to claim 1 is implemented, wherein the distributed memory database system for financial market data comprises: The hierarchical data storage module is used to convert financial market data into a table nested hierarchical data table stored in the form of a tree to achieve a multi-level table structure design; The table-level multi-version concurrency control module is used to initialize and generate the first initialization memory block according to the preset size for the nested hierarchical data table. When adding records, the additional part is added to the initialization memory block, and the index is established using a lock-free hashmap. When the first initialization memory block is full, a second initialization memory block with a preset index plus 1 is allocated and linked to the first initialization memory block until it is linked to a third initialization memory block with an index equal to a predefined index. Each subsequent block is equal to the size of the third initialization memory block to build an innovative chain array storage structure; when there are too many holes in the initialization memory block and the memory utilization rate is low due to deleting records without moving any memory or other operations, the table version switching logic is triggered to create a new table copy, and the data in the initialization memory block is reorganized and written to the new table copy; The shared memory persistence module is used to manage the entire shared memory in a page table manner to generate a log storage area, and the table-level multi-version concurrency control module processes the records to generate binlog log data, which is continuously stored in the log storage area in an appended manner. When the log area data meets one of the preset conditions, such as the number of logs is greater than the preset number, the number of logs exceeds N times the snapshot data, or the time reaches a certain threshold, the snapshot operation is triggered. The snapshot operation is based on the last snapshot, replays a certain number of logs, replaces the old snapshot data, and deletes the replayed logs after the replacement is successful. The real-time data distribution module is used to establish a TCP control channel in the downstream node, and synchronize the hierarchical data table, the database identifier in the linked array, the serial number of the current binlog log data, the schema information and the multicast configuration parameters through the TCP channel. The downstream node accesses the multicast according to the multicast configuration parameters and starts to receive the multicast packet. After receiving the first multicast packet, the log event number of the multicast packet is updated, and the established packet replenishment channel is used to replenish the data before the current multicast packet. Then, on this basis, the multicast event is played back to achieve synchronous update, and packet loss detection is performed by detecting the sequence number and log number in each multicast packet. If packet loss is found, compensation is performed through the TCP channel.