Self-adaptive data segmentation and query optimization system based on multi-core parallel computing
By adopting the adaptive data segmentation and query optimization method of multi-core parallel computing in the database system, the problem of increasing query speed and response time in large-scale data processing in traditional databases is solved, and efficient query processing and data security are achieved.
Patent Information
- Application Number
- CN202510334057.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-04
AI Technical Summary
When traditional database systems process large-scale data, query speed and response time may increase significantly, especially when facing high concurrent requests, bottlenecks are prone to occur, which cannot effectively meet the requirements for data query efficiency in modern applications, and fail to make full use of the parallel computing capabilities of multi-core processors.
Adaptive data segmentation and query optimization system based on multi-core parallel computing is adopted. By separating the data query service from the main database, using the parallel processing capabilities of multi-core CPUs, the data tables in the segmented table are processed in parallel, combining data indexing, caching mechanisms and load balancing, query performance is optimized and high concurrency support is provided.
It significantly improves query processing speed, reduces response time, reduces system delay, improves user experience, and ensures that data writing and query processing are not interfered with each other and data security.
Smart Images

Figure CN120256464A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of database management, and more specifically, particularly relates to an adaptive data segmentation and query optimization system based on multi-core parallel computing. Background Art
[0002] With the rapid development of information technology, especially in the context of the wide application of big data, cloud computing, and distributed systems, traditional single-machine databases often cannot meet the requirements of massive data storage and real-time query.
[0003] When traditional database systems process large-scale data, the query speed and response time may increase significantly. Especially when facing high-concurrency requests, bottlenecks are likely to occur. These problems result in the inability of traditional databases to effectively meet the requirements for data query efficiency in modern applications. With the increase in users and devices, front-end application programs face an increasingly high query load, causing traditional databases to be unable to quickly respond to a large number of concurrent query requests, thereby affecting the performance of application programs and the user experience.
[0004] In database systems, read-write separation is a common technical means to improve performance and reliability. However, in existing database architectures, read-write separation often fails to provide effective performance optimization during large-scale data queries. Traditional architectures cannot effectively support data segmentation and query processing mechanisms under high concurrency. Current database systems usually do not fully utilize the parallel computing capabilities of multi-core processors, resulting in inefficient allocation of computing resources and low processing efficiency when performing complex queries. Especially in distributed systems, when data is synchronized among multiple nodes, how to ensure data real-time and consistency, and avoid data redundancy and latency issues, is the key to achieving efficient queries. Summary of the Invention
[0005] In view of the above or existing problems of the adaptive data segmentation and query optimization system based on multi-core parallel computing, the present invention is proposed.
[0006] To solve the above technical problems, the present invention provides the following technical solutions:
[0007] An embodiment of the present invention provides an adaptive data segmentation and query optimization system based on multi-core parallel computing, including: a query database server, which is used to separate data query services from the main database through a local customized database. After data segmentation, it utilizes the parallel processing ability of a multi-core CPU to perform parallel processing on the data tables in the segmented table set, generating result subsets, and the result subsets are finally merged into a result merge set and returned to the application server; a customized database, which is used to support data table segmentation and parallel processing by synchronizing data changes in the main database in real time, utilizes multi-core CPU parallel computing to accelerate queries, optimizes query performance through data indexing, caching mechanisms, and load balancing, and provides high-concurrency support; an adaptive data control unit, which is used for a CPU core number scanner to detect the number of CPU cores of the query database server and send the detected result value Ncpu to the database segmentation controller; the database segmentation controller detects the total number of data records Ndata of the data tables to be processed in the customized database, and then calculates the number of data records Nnumber in each segmented data table as Nnumber = Ndata / Ncpu; the merge controller merges the Ncpu result sub-tables in the result subsets to generate a result merge set; a data customization management unit, which is used for the main database monitoring and data acquisition controller to monitor data changes in the main database in real time, and collect and store the changed data records in the main database into the update data set; the customized data update controller monitors the status of the update data set in real time. When it detects that there are data records arriving in the update data set, it reads the data records in real time and starts the customized database update process to complete the update of the data records in the customized database; the main database, which is used to bear the write load from the front-end application server and the load generated by the customized data of the query database server; the write database server, which is used to provide the front-end application server with the service of writing data into the main database; the backup database, which is used to back up the data of the main database in real time to ensure the security of system data.
[0008] As a preferred solution of the adaptive data segmentation and query optimization system based on multi-core parallel computing according to the present invention, wherein: after implementing data segmentation, utilizing the parallel processing ability of a multi-core CPU to perform parallel processing on the data tables in the segmented table set, generating result subsets, includes:
[0009] After data segmentation, the query database server utilizes its multi-core CPU to perform parallel processing on the segmented sub-tables. Each CPU core processes the data of one sub-table and executes the corresponding query operation; each CPU core receives one sub-table and independently executes the query operation; after each CPU core finishes processing its own sub-table, it generates a result subset. If the i-th CPU core processes the i-th sub-table, then the result generated by this core is R i , where i ∈ [1, Ncpu].
[0010] As a preferred solution of the adaptive data segmentation and query optimization system based on multi-core parallel computing according to the present invention, the system supports data table segmentation and parallel processing by real-time synchronizing data changes in the main database, accelerates queries using multi-core CPU parallel computing, and optimizes query performance through data indexing, caching mechanisms, and load balancing in the customized database, and provides high concurrency support, including:
[0011] In the case of a large amount of data, the query database achieves load balancing by evenly dividing the data table into multiple sub-tables according to the number of records. Given a data table T with a total number of records Ndata, if there are Ncpu CPU cores, the system divides the data table into Ncpu sub-tables, and each sub-table contains approximately Ndata / Ncpu records;
[0012] Each segmented sub-table is assigned to a CPU core for independent query operations. Through parallel processing, after the data table is segmented, the query tasks are distributed to multiple CPU cores in parallel, and each CPU core is responsible for processing the query tasks of the sub-table assigned to it; the query results generated by each CPU core are subsets, and finally, the multiple result subsets are integrated into the final query result by the merge controller and returned to the application program.
[0013] As a preferred solution of the adaptive data segmentation and query optimization system based on multi-core parallel computing according to the present invention, the CPU core number scanner detects the number of CPU cores of the query database server and sends the detected result value Ncpu to the database segmentation controller, including:
[0014] According to the available number of CPU cores, the database segmentation controller splits the data table into an appropriate number of sub-tables. If there are more CPU cores, the system divides the data table into more sub-tables to enhance the effect of parallel processing and thus improve query performance.
[0015] As a preferred solution of the adaptive data segmentation and query optimization system based on multi-core parallel computing according to the present invention, the database segmentation controller detects the total number of data records Ndata of the data table to be processed in the customized database, and then calculates the number of data records Nnumber in each segmented data table as Nnumber = Ndata / Ncpu, including:
[0016] If the total number of records in the data table is set as Ndata and the number of CPU cores of the query database server is Ncpu, then the number of data records Nnumber in each sub-table can be calculated by the following formula:
[0017]
[0018] If Ndata is an even number, the database segmentation controller equally segments the data tables to be processed in the customized database into Ncpu data tables according to the natural order of data records based on the Nnumber value; if Ndata is an odd number, then Nnumber = (Ndata - 1) / Ncpu, and the database segmentation controller equally segments the first Ndata - 1 data records in the data tables to be processed in the customized database into Ncpu data tables.
[0019] As a preferred embodiment of the adaptive data segmentation and query optimization system based on multi-core parallel computing according to the present invention, wherein: the merging controller merges the Ncpu result sub-tables in the result subset to generate a result merging set, including:
[0020] If the query result requires sorting according to a certain field, the merging controller will perform sorting during the merging of each result subset to ensure that the final result is sorted according to the specified conditions;
[0021] If the query includes a deduplication operation, the merging controller will check for duplicate data and remove duplicate records during the merging to ensure that each record appears only once in the final result;
[0022] Once the data sorting and deduplication operations are completed, the merging controller integrates all result subsets, merges each result subset in order, and forms a large result set.
[0023] As a preferred embodiment of the adaptive data segmentation and query optimization system based on multi-core parallel computing according to the present invention, wherein: the main database monitoring and data acquisition controller monitors the changes in the data in the main database in real time, and collects and stores the changed data records in the main database into the update data set, including:
[0024] The main database monitoring and data acquisition controller monitors any changes in the data in the database in real time by listening to the change log or trigger mechanism of the database. The database trigger automatically captures the data modification events in the main database and notifies the monitoring and acquisition controller that data changes have occurred;
[0025] When data changes are detected, the monitoring controller will extract the changed data records through database query operations. For inserted and updated records, the monitoring controller collects the specific content of the changed data.
[0026] As a preferred embodiment of the adaptive data segmentation and query optimization system based on multi-core parallel computing according to the present invention, wherein: the customized data update controller monitors the status of the update data set in real time. When it detects that there are data records arriving in the update data set, it reads the data records in real time and starts the customized database update process to complete the update of the data records in the customized database, including:
[0027] The customized data update controller monitors the status of the updated data set in real time. After detecting the arrival of new data records, it immediately reads and triggers the update process. According to the type of data records, the controller performs corresponding operations on the customized database to ensure data consistency and integrity.
[0028] As a preferred solution of the adaptive data segmentation and query optimization system based on multi-core parallel computing according to the present invention, wherein: the load borne from the front-end application server for writing and the load generated by querying the customized data of the database server includes:
[0029] When data changes occur in the main database, the query database server needs to synchronously update its local customized database. Whenever data changes occur in the main database, the write database server receives a synchronization request from the query database server and is responsible for transferring the latest data into the customized database.
[0030] As a preferred solution of the adaptive data segmentation and query optimization system based on multi-core parallel computing according to the present invention, wherein:
[0031] Providing the main database data writing service for the front-end application server, including:
[0032] The front-end application server sends a write request to the write database server through the API. The request contains the data that the user needs to write into the database. The write request includes the fields and value timestamps of the data record, the user ID, and the operation type.
[0033] The beneficial effects of the present invention are as follows: Through adaptive data cutting and multi-core parallel processing, the system can cut a large data table into multiple sub-tables and utilize the multi-core CPU for parallel processing, greatly improving the query processing speed. This significantly improves the query efficiency and reduces the response time when facing massive data and high-concurrency requests. By separating the query database server from the customized database and combining the deployment of multiple query database servers, it can effectively share high-concurrency access requests, reduce the load on the front-end application server, further reduce the system latency, and improve the user experience. By separating the write database server from the main database, it ensures that data writing and query processing do not interfere with each other, avoids the impact of write operations on query efficiency, and improves the writing performance. At the same time, the real-time synchronization mechanism of the backup database ensures data security. Description of the Drawings
[0034] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0035] Figure 1 This is a schematic diagram of an adaptive data segmentation and query optimization system structure based on multi-core parallel computing provided by an embodiment of the present invention. Detailed implementation manners
[0036] To make the above objects, features, and advantages of the present invention more apparent and understandable, the following detailed description of the specific implementation manners of the present invention will be given with reference to the accompanying drawings of the specification.
[0037] In the following description, many specific details are set forth to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0038] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that can be included in at least one implementation manner of the present invention. The appearances of "in one embodiment" in different places in this specification do not all refer to the same embodiment, nor are they separate or alternative embodiments that exclude each other with other embodiments.
[0039] Embodiment
[0040] S1: Separate the data query service from the main database through the local customized database. After implementing data segmentation, utilize the parallel processing ability of the multi-core CPU to perform parallel processing on the data tables in the segmented table set, generate result subsets, and finally merge the result subsets into a result merge set and return it to the application server.
[0041] Preferably, after data segmentation, the query database server utilizes its multi-core CPU to perform parallel processing on the segmented sub-tables. Each CPU core processes the data of one sub-table and executes the corresponding query operation; each CPU core receives one sub-table and independently executes the query operation; after each CPU core finishes processing its own sub-table, it generates a result subset. If the i-th CPU core processes the i-th sub-table, then the result generated by this core is R i , where i ∈ [1, Ncpu].
[0042] Furthermore, assume that Ndata = 1000 pieces of data and the query database server has 4 CPU cores (Ncpu = 4).
[0043] The data table will be segmented into 4 sub-tables, and each sub-table roughly contains Ndata / Ncpu = 1000 / 4 = 250 records (if Ndata is odd, make a slight adjustment to ensure that the last sub-table contains the remaining data).
[0044] The data distribution of each sub - table is as follows:
[0045] Sub - table 1: Records 1 to 250
[0046] Sub - table 2: Records 251 to 500
[0047] Sub - table 3: Records 501 to 750
[0048] Sub - table 4: Records 751 to 1000
[0049] Each CPU core of the query database server (e.g., Core 1, Core 2, Core 3, Core 4) receives its corresponding sub - table data (Sub - table 1 to Sub - table 4) and performs a query operation. The query operation can be any database operation, such as retrieving data of a specific field, aggregate calculation, filtering records that meet the conditions, etc.
[0050] Core 1: Processes Sub - table 1 (Records 1 to 250), performs a query operation and generates a result subset R1;
[0051] Core 2: Processes Sub - table 2 (Records 251 to 500), performs a query operation and generates a result subset R2;
[0052] Core 3: Processes Sub - table 3 (Records 501 to 750), performs a query operation and generates a result subset R3;
[0053] Core 4: Processes Sub - table 4 (Records 751 to 1000), performs a query operation and generates a result subset R4;
[0054] Each CPU core performs corresponding operations according to the query conditions and independently generates query results. After all CPU cores complete the query operations and generate result subsets, the merge controller merges all result subsets R1, R2, R3, R4 into the final query result set and returns it to the front - end application.
[0055] S2: By real - time synchronizing data changes in the main database, it supports data table partitioning and parallel processing, uses multi - core CPU parallel computing to accelerate queries, and the customized database optimizes query performance through data indexing, caching mechanisms, and load balancing, and provides high - concurrency support.
[0056] Preferably, in the case of a large amount of data, the query database achieves load balancing by evenly partitioning the data table into multiple sub - tables according to the number of records. Given a data table T with a total number of records Ndata, if there are Ncpu CPU cores, the system divides the data table into Ncpu sub - tables, and each sub - table contains approximately Ndata / Ncpu records;
[0057] Each split sub-table is assigned to a CPU core for independent query operations. Through parallel processing, after the data table is split, the query tasks are distributed to multiple CPU cores in parallel. Each CPU core is responsible for processing the query tasks of the sub-table assigned to it; the query results generated by each CPU core are subsets, and finally the multiple result subsets are integrated into the final query results through the merge controller and returned to the application.
[0058] Furthermore, the data access frequency weight factor (W_i) and field complexity factor (C_j) are introduced to optimize the segmentation formula to:
[0059]
[0060] Weight factor calculation:
[0061]
[0062] C j =1+log2(field type complexity) (e.g. INT=1, JSON=3, BLOB=5)
[0063] High-frequency cores are assigned to complex data tables with higher weights, especially those fields involving complex query conditions; low-frequency cores are assigned to simpler field data to ensure that CPU resources are used reasonably;
[0064] The system dynamically adjusts the data segmentation strategy based on the real-time monitored CPU frequency to ensure that load balancing can be maintained even when the CPU frequency is adjusted:
[0065] Load balancing is improved by 30%: Since the system dynamically adjusts the data distribution, the load of each CPU core is more balanced, avoiding the load imbalance between cores. In particular, high-frequency CPU cores can take on more computing tasks, while low-frequency cores take on lighter tasks. In this way, the overall computing power of the system is fully utilized and the load balancing is significantly improved;
[0066] Energy efficiency ratio optimized by 15%: The system adapts the CPU's dynamic frequency adjustment to ensure that high-frequency cores can achieve maximum performance when executing computationally intensive queries, while low-frequency cores maintain lower power consumption when performing simple tasks. In this way, the system not only improves performance but also effectively reduces energy consumption, thereby optimizing energy efficiency.
[0067] Further, suppose there is a data table T in the database, and its total number of records is Ndata. Suppose there are NcpuNcpuNcpu available CPU cores in the system, and the goal is to divide the data table T into Ncpu sub-tables, each of which contains about Ndata / Ncpu records.
[0068] The records in the data table T are sequentially divided into multiple sub - tables T1, T2, ..., T Ncpu , and the number of records in each sub - table is approximately the same. The record volume N Ti of each sub - table is close to Ndata / Ncpu, and the record distribution is uniform.
[0069] Assume that the starting record number of each sub - table is: sub - table T1 starts from record 1 and ends at Ndata / Ncpu. Sub - table T2 starts from record Ndata / Ncpu + 1 and ends at 2×Ndata / Ncpu, and so on until the last sub - table T Ncpu .
[0070] Each CPU core is assigned a sub - table Ti. For example: core 1 processes sub - table T1, core 2 processes sub - table T2, and so on until core Ncpu processes sub - table T Ncpu . Each CPU core independently processes the query tasks of the assigned sub - table Ti. For example: core 1 is responsible for querying the data in sub - table T1,
[0071] core 2 is responsible for querying the data in sub - table T2, and so on until all cores execute the query tasks.
[0072] When each CPU core is processing, it performs operations related to the query, such as filtering records from the sub - table according to the query conditions and sorting the query results.
[0073] After each CPU core completes the query, it generates a result subset R1, R2, ..., RNcpu for its own processed query results, where each result subset contains the query results of the corresponding sub - table.
[0074] After all CPU cores have finished processing, the query results need to be merged into the final query result. For this purpose, the system sets up a merge controller, which is responsible for: merging the result subsets R1, R2, ..., RNcpu generated by all cores; for aggregation queries, the merge controller needs to merge the calculation results of all result subsets; for sorting queries, the merge controller may need to re - sort multiple sorted results.
[0075] S3: The CPU core number scanner detects the number of CPU cores of the query database server and sends the detected result value Ncpu to the database segmentation controller; the database segmentation controller detects the total number of data records Ndata in the data table to be processed in the customized database, and then calculates the number of data records Nnumber in each segmented data table = Ndata / Ncpu; the merge controller merges the Ncpu result sub - tables in the result subsets to generate a result merge set.
[0076] Preferably, according to the number of available CPU cores, the database splitting controller splits the data table into an appropriate number of sub-tables. If there are more CPU cores, the system will split the data table into more sub-tables to enhance the effect of parallel processing, thereby improving the query performance.
[0077] Preferably, set the total number of records in the data table as Ndata, and the number of CPU cores of the query database server as Ncpu. Then the number of data records Nnumber in each sub-table can be calculated by the following formula:
[0078]
[0079] If Ndata is an even number, the database splitting controller equally splits the data table to be processed in the customized database into Ncpu data tables according to the natural order of the data records according to the Nnumber value; if Ndata is an odd number, then Nnumber = (Ndata - 1) / Ncpu, and the database splitting controller equally splits the first Ndata - 1 data records in the data table to be processed in the customized database into Ncpu data tables according to the Nnumber value.
[0080] Preferably, if the query result requires sorting by a certain field, the merging controller will perform sorting when merging each result subset to ensure that the final result is sorted according to the specified conditions;
[0081] If the query contains a deduplication operation, the merging controller will check for duplicate data and remove duplicate records during merging to ensure that each record appears only once in the final result;
[0082] Once the data sorting and deduplication operations are completed, the merging controller integrates all result subsets, merges each result subset in order, and forms a large result set.
[0083] Further, assume that there is a data table orders with a total number of records Ndata = 1,001,00, and the number of CPU cores of the query database server is Ncpu = 10. Since Ndata = 1001000N is an odd number, we use the odd splitting method:
[0084] Nnumber = (Ndata - 1) / Ncpu = 100,099
[0085] Each sub-table will contain 100,099 records; the first 100,099 records in the data table orders are evenly divided into 10 sub-tables, with each sub-table containing 100,099 records. The last remaining 1 record will be processed separately. It can be optionally assigned to the last sub-table or placed in a dedicated temporary sub-table. The system assigns these 10 sub-tables to 10 CPU cores for parallel processing. Core 1 processes sub-table 1, core 2 processes sub-table 2,
[0086] Furthermore, set some key performance metric thresholds and parameters for the splitting strategy, which will serve as the basis for the system's adaptive adjustment.
[0087] Set the minimum value of the cache hit rate, for example, set it to 0.8. If the cache hit rate is lower than this value, it indicates poor cache performance and the need to optimize the data splitting strategy. Set the threshold for memory bandwidth occupancy, for example, 80%. If the memory bandwidth exceeds this threshold, it indicates that the memory bandwidth may become a bottleneck and the need to adjust the strategy. Set the basic splitting granularity of the data table, for example, each sub-table is split into 1000 records. Set the maximum number of concurrent threads according to the number of CPU cores in the system, usually set to Ncpu. The system monitors in real time by collecting the performance metrics of the system and dynamically adjusts the data splitting strategy based on this information.
[0088] Collect the following performance metrics through the system interface: CPU core frequency, L3 cache hit rate, and memory bandwidth occupancy
[0089] Performance score calculation: The comprehensive score is calculated from performance metrics in multiple dimensions and aims to reflect the overall performance state of the system. The calculation formula is as follows:
[0090]
[0091] where: α, β, γ are the weight coefficients of each indicator, representing the importance of different indicators, f current is the actual frequency of the current CPU core, f base is the base frequency, cache_hit_rate is the L3 cache hit rate, mem_bandwidth is the current usage rate of the memory bandwidth, and mem_bandwidthmax is the maximum bandwidth of the memory;
[0092] Based on the real-time performance score, the system dynamically adjusts the data splitting strategy to optimize the query processing efficiency.
[0093] Relationship between the dynamic adjustment strategy and the score
[0094] Excellent performance (score 0.9 - 1.0): Sufficient system resources, maintain the current splitting strategy, and appropriately increase the number of concurrent threads (such as increasing by 10%).
[0095] Good performance (score 0.7 - 0.9): The system load is low. Fine-tune the locality optimization parameters as needed, such as adjusting data locality (e.g., preferentially allocate adjacent records to the same sub-table).
[0096] Medium performance (score 0.5 - 0.7): The system load is medium. Reduce the segmentation granularity, divide large data tables into smaller sub-tables, and reduce the persistence frequency of sub-tables.
[0097] Performance bottleneck (score < 0.5): The system load is heavy. Trigger an alarm, forcefully reduce the number of concurrent threads, and prioritize ensuring core business queries.
[0098] S4: The main database monitoring and data acquisition controller monitors the changes in the data in the main database in real time, collects and stores the changed data records in the main database into the update dataset; the customized data update controller monitors the status of the update dataset in real time. When it detects that there are data records arriving in the update dataset, it reads the data records in real time and starts the customized database update process to complete the update of the data records in the customized database.
[0099] Preferably, the main database monitoring and data acquisition controller monitors any changes in the data in the database in real time by listening to the change log or trigger mechanism of the database. The database trigger automatically captures the data modification events in the main database and notifies the monitoring and acquisition controller that data changes have occurred;
[0100] When detecting data changes, the monitoring controller will extract the changed data records through database query operations. For inserted and updated records, the monitoring controller collects the specific content of the changed data.
[0101] Preferably, the customized data update controller monitors the status of the update dataset in real time. After detecting the arrival of new data records, it immediately reads and triggers the update process. According to the data record type, the controller performs corresponding operations on the customized database to ensure data consistency and integrity.
[0102] S5: Bear the write load from the front-end application server and the load generated by querying the customized data of the database server.
[0103] Preferably, when the data in the main database changes, the query database server needs to synchronously update its local customized database. Whenever the data in the main database changes, the write database server receives a synchronization request from the query database server and is responsible for transferring the latest data to the customized database.
[0104] Further, when the primary database changes, the write database server will pass the change information to the query database server through database change logs, triggers, or a dedicated data synchronization mechanism. The write database server will monitor the change logs of the primary database, capture data changes, and generate change events. Set a trigger on the primary database, and when the data changes, the trigger will automatically notify the query database server.
[0105] Another common method is to use a message queue system. The write database server sends the change events to the message queue, and the query database server will retrieve the updated events from the queue. After receiving the change notification, the query database server will synchronize and update according to the latest data of the primary database; the query database server will extract the latest data of the primary database. For insert operations, the query database server will insert the new data record into the customized database; for update operations, the query database server will update the corresponding record in the customized database; for delete operations, the query database server will delete the corresponding record in the customized database.
[0106] S6: Provide the primary database data writing service for the front-end application server.
[0107] Preferably, the front-end application server sends a write request to the write database server through the API. The request contains the data that the user needs to write into the database. The write request contains the fields and value timestamps of the data record, the user ID, and the operation type.
[0108] Further, the front-end application server generates a write request through the operations or events input by the user. When the user submits an order, updates information, or deletes a certain piece of data in the application interface, the front-end will generate a JSON-format request containing the user ID, data fields and values, operation type, and timestamp.
[0109] The front-end sends the request to the backend API gateway through an API call, and the API gateway will perform the following operations:
[0110] Check whether the request data is complete and valid. For example, ensure that fields such as timestamps and user IDs exist and conform to the expected format. Confirm whether the user ID provided in the request is legal to ensure that only authorized users can initiate write operations. The API gateway forwards the verified request to the write database server for further processing.
[0111] After receiving the request, the write database server will perform the following operations:
[0112] Parse the data record according to the field content in the request, and determine the corresponding database operation based on the operation type (such as insert, update, or delete). Insert a new record into the database table. The system generates a unique record ID (such as an order number) and inserts the data into the target table. Update the existing record according to the data and conditions (such as the primary key ID or unique identifier) in the request. The updated data may be a modification of a certain field or a replacement of the entire record. Delete the target record according to the identifier in the request. Store the timestamp together with the data to ensure the timeliness of the record. The timestamp can be used to determine the update time of the data or create an audit record for the data.
[0113] In the description of the present invention, it should be noted that the terms "first", "second", and "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.
[0114] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0115] In several embodiments provided by the present invention, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed mutual coupling or direct coupling or communication connection can be through some communication interfaces. The indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.
[0116] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0117] In addition, in each embodiment of the present invention, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0118] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than to limit it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
[0119] In addition, although the operations of the method of the present invention are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all of the shown operations must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.
Claims
1. An adaptive data segmentation and query optimization system based on multi-core parallel computing, characterized in that Including: A query database server, which is used to separate the data query service from the main database through a local customized database. After data segmentation is implemented, it utilizes the parallel processing ability of a multi-core CPU to perform parallel processing on the data tables in the segmented table set, generating result subsets, and finally merging the result subsets into a result merge set and returning it to the application server; A customized database, which is used to support data table segmentation and parallel processing by real-time synchronizing data changes in the main database, utilizes multi-core CPU parallel computing to accelerate queries, optimizes query performance through data indexing, caching mechanisms, and load balancing, and provides high-concurrency support; An adaptive data control unit, which is used for the CPU core number scanner to detect the number of CPU cores of the query database server and send the detected result value Ncpu to the database segmentation controller; the database segmentation controller detects the total number of data records Ndata of the data tables to be processed in the customized database, and then calculates the number of data records Nnumber in each segmented data table as Nnumber = Ndata / Ncpu; the merge controller merges the Ncpu result sub-tables in the result subset to generate a result merge set; A data customization management unit, which is used for the main database monitoring and data collection controller to monitor data changes in the main database in real time, collect and store the data records that have changed in the main database into the update data set; the customized data update controller monitors the status of the update data set in real time. When it detects that there are data records arriving in the update data set, it reads the data records in real time and starts the customized database update process to complete the update of the data records in the customized database; The main database, which is used to bear the write load from the front-end application server and the load generated by the customized data of the query database server; A write database server, which is used to provide the front-end application server with the service of writing main database data; A backup database, which is used to back up the data of the main database in real time to ensure the security of system data.
2. The adaptive data segmentation and query optimization system based on multi-core parallel computing according to claim 1, wherein After implementing data segmentation, utilizing the parallel processing ability of a multi-core CPU to perform parallel processing on the data tables in the segmented table set, generating result subsets, including: After data partitioning, the query database server uses its multi-core CPU to process the partitioned sub-tables in parallel. Each CPU core processes the data of one sub-table and executes the corresponding query operation; each CPU core receives a sub-table and independently executes the query operation; after each CPU core finishes processing its own sub-table, it generates a result subset. If the i-th CPU core processes the i-th sub-table, then the result generated by this core is Ri i , where i ∈ [1, Ncpu].
3. The adaptive data segmentation and query optimization system based on multi-core parallel computing according to claim 1, characterized in that, By real-time synchronizing data changes in the main database, supporting data table segmentation and parallel processing, utilizing multi-core CPU parallel computing to accelerate queries, the customized database optimizes query performance through data indexing, caching mechanisms, and load balancing, and provides high-concurrency support, including: In the case of a large amount of data, the query database achieves load balancing by evenly dividing the data table into multiple sub-tables according to the number of records. Given a data table T with a total number of records Ndata, if there are Ncpu CPU cores, the system divides the data table into Ncpu sub-tables, and each sub-table contains approximately Ndata / Ncpu records; Each of the segmented sub - tables is assigned to a CPU core for independent query operations. Through parallel processing, after the data table is segmented, the query tasks are distributed to multiple CPU cores in parallel, and each CPU core is responsible for processing the query tasks of the sub - table assigned to it; the query results generated by each CPU core are subsets, and finally, the merge controller integrates multiple result subsets into the final query result and returns it to the application program.
4. The adaptive data segmentation and query optimization system based on multi-core parallel computing according to claim 1, wherein The CPU core number scanner detects the number of CPU cores of the query database server and sends the detected result value Ncpu to the database segmentation controller, including: According to the available number of CPU cores, the database segmentation controller splits the data table into an appropriate number of sub - tables. If there are more CPU cores, the system will split the data table into more sub - tables to enhance the effect of parallel processing, thereby improving the query performance.
5. The adaptive data segmentation and query optimization system based on multi-core parallel computing according to claim 1, characterized in that, The database segmentation controller detects the total number of data records Ndata of the data table to be processed in the customized database, and then calculates the number of data records Nnumber in each segmented data table as Nnumber = Ndata / Ncpu, including: Set the total number of records in the data table as Ndata and the number of CPU cores of the query database server as Ncpu. Then the number of data records Nnumber in each sub - table can be calculated by the following formula: If Ndata is even, the database segmentation controller evenly splits the data table to be processed in the customized database into Ncpu data tables according to the Nnumber value in the natural order of data records; if Ndata is odd, then Nnumber = (Ndata - 1) / Ncpu, and the database segmentation controller evenly splits the first Ndata - 1 data records in the data table to be processed in the customized database into Ncpu data tables.
6. The adaptive data segmentation and query optimization system based on multi-core parallel computing according to claim 1, wherein The merge controller merges the Ncpu result sub - tables in the result subset to generate a result merge set, including: If the query result requires sorting by a certain field, the merge controller will sort when merging each result subset to ensure that the final result is sorted according to the specified conditions; If the query contains a deduplication operation, the merge controller will check for duplicate data and remove duplicate records during the merge to ensure that each record appears only once in the final result; Once the data sorting and deduplication operations are completed, the merge controller integrates all result subsets, merges each result subset in order to form a large result set.
7. The adaptive data segmentation and query optimization system based on multi-core parallel computing according to claim 1, wherein The main database monitoring and data acquisition controller monitors the changes in the data of the main database in real - time and collects and stores the data records that have changed in the main database into the update dataset, including: The main database monitoring and data acquisition controller monitors any changes in the data in the database in real - time by listening to the change log or trigger mechanism of the database. The database trigger automatically captures the data modification events in the main database and notifies the monitoring and acquisition controller that data changes have occurred; When data changes are detected, the monitoring controller will extract the changed data records through database query operations. For inserted and updated records, the monitoring controller collects the specific content of the changed data.
8. The adaptive data segmentation and query optimization system based on multi-core parallel computing according to claim 1, characterized in that, The customized data update controller monitors the status of the update data set in real time. When it detects that a data record arrives in the update data set, it reads the data record in real time and starts the customized database update process to complete the update of the data record in the customized database, including: The customized data update controller monitors the status of the update data set in real time. After detecting the arrival of a new data record, it immediately reads and triggers the update process. According to the data record type, the controller performs corresponding operations on the customized database to ensure data consistency and integrity.
9. The adaptive data segmentation and query optimization system based on multi-core parallel computing according to claim 1, characterized in that It undertakes the write load from the front-end application server and the load generated by querying the customized data of the database server, including: When the data in the main database changes, the query database server needs to synchronously update its local customized database. Whenever the data in the main database changes, the write database server receives a synchronization request from the query database server and is responsible for transferring the latest data to the customized database.
10. The adaptive data segmentation and query optimization system based on multi-core parallel computing according to claim 1, characterized in that, It provides the main database data write service for the front-end application server, including: The front-end application server sends a write request to the write database server through the API. The request contains the data that the user needs to write into the database. The write request includes the fields and value timestamps of the data record, the user ID, and the operation type.