Data processing method and device, computer equipment and storage medium

By configuring persistent memory in the database system and allocating appropriately sized buffers, the problem of low performance of the database system when processing large data volumes is solved, and more efficient data operation and response speed is achieved.

CN120234352APending Publication Date: 2025-07-01NEW H3C BIG DATA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510381078.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

When processing large data volumes, the database system relies on disk I/O operations, resulting in slow data reading and writing speeds, affecting performance and response speed.

Method used

Configure persistent memory in the database system, allocate memory through buffers with a memory threshold size, and use persistent memory as a temporary buffer for data operations. When the data volume is large, allocate larger buffers to reduce disk I/O overhead.

Benefits of technology

By utilizing persistent memory, the efficiency of data operations is significantly improved, and the performance and response speed of the database system are improved, especially when processing large amounts of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234352A_ABST
    Figure CN120234352A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of databases, and discloses a data processing method and device, computer equipment and a storage medium, the method comprises the following steps: in response to a request for performing data operation on to-be-processed data, determining the data volume of the to-be-processed data; under the condition that the data size of the to-be-processed data is smaller than a memory threshold value, data operation is conducted on the to-be-processed data through the first buffer area, and the persistent memory comprises the first buffer area with the memory threshold value; and under the condition that the data volume of the to-be-processed data is greater than the memory threshold value, distributing a second buffer area matched with the data volume of the to-be-processed data from the persistent memory, and performing data operation on the to-be-processed data by utilizing the second buffer area. According to the method, the persistent memory is used as a temporary buffer area to process the to-be-processed data, so that the performance and response speed of a database system can be improved, and the efficiency of memory intensive data operation can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of databases, and particularly to a data processing method, apparatus, computer device, and storage medium. Background Art

[0002] A database system can provide data storage for an enterprise and uniformly manage the data. In the actual application process, data operations on the database system are often required, such as querying, index building, sorting, file operations, etc.

[0003] In the related art, the database system uses disk storage and relies on disk I / O (input / output) operations, and the data reading and writing speeds are slow. In the case of a large amount of data, this method will have performance bottlenecks, affecting the performance and response speed of the entire database system. Summary of the Invention

[0004] In view of this, the present invention provides a data processing method, apparatus, computer device, and storage medium to solve the problem of low data operation efficiency of the database system.

[0005] In a first aspect, the present invention provides a data processing method applied to a database system. The database system is configured with persistent memory, and the persistent memory includes a first buffer with a memory threshold size, including:

[0006] Responding to a request for performing a data operation on the data to be processed, determining the data volume of the data to be processed;

[0007] When the data volume of the data to be processed is less than the memory threshold, using the first buffer to perform a data operation on the data to be processed;

[0008] When the data volume of the data to be processed is greater than the memory threshold, allocating a second buffer from the persistent memory that matches the data volume size of the data to be processed, and using the second buffer to perform a data operation on the data to be processed.

[0009] In some optional embodiments, the using the second buffer to perform a data operation on the data to be processed includes:

[0010] Writing the data to be processed into the second buffer;

[0011] Dividing the data to be processed in the second buffer into multiple data blocks, and performing data operations on each data block in parallel to generate data operation results corresponding to each data block;

[0012] Summarize the data operation results corresponding to each of the data blocks to generate the data operation result of the data to be processed.

[0013] In some alternative embodiments, the method further includes:

[0014] When the data volume of the data to be processed is greater than the memory threshold, create N child processes, where N is a positive integer;

[0015] The step of dividing the data to be processed in the second buffer into multiple data blocks and performing data operations on each data block in parallel to generate the data operation results corresponding to each of the data blocks includes:

[0016] Divide the data to be processed in the second buffer into N data blocks, and the N data blocks correspond to the N child processes one by one;

[0017] Use the N child processes to perform data operations on the corresponding data blocks in parallel to obtain the data operation results corresponding to each of the child processes.

[0018] In some alternative embodiments, the step of creating N child processes includes:

[0019] Obtain the pre-set maximum number of processes;

[0020] Create N child processes according to the maximum number of processes; N is less than or equal to the maximum number of processes.

[0021] In some alternative embodiments, after performing data operations on the data to be processed using the second buffer, the method further includes:

[0022] Release the second buffer.

[0023] In some alternative embodiments, the method further includes:

[0024] Create a corresponding main process for the client connecting to the database system;

[0025] Allocate a first buffer with a size equal to the memory threshold for the client from the persistent memory based on the main process.

[0026] In some alternative embodiments, the method further includes:

[0027] After the client disconnects from the database system, release the first buffer allocated for the client.

[0028] Second aspect, the present invention provides a data processing device, which is applied to a database system. The database system is configured with persistent memory, and the persistent memory includes a first buffer with a memory threshold size, including:

[0029] A request processing module, configured to determine the data volume of the data to be processed in response to a request for performing a data operation on the data to be processed;

[0030] A first processing module, configured to perform a data operation on the data to be processed by using the first buffer when the data volume of the data to be processed is less than the memory threshold;

[0031] A second processing module, configured to allocate a second buffer matching the data volume of the data to be processed from the persistent memory and perform a data operation on the data to be processed by using the second buffer when the data volume of the data to be processed is greater than the memory threshold.

[0032] Third aspect, the present invention provides a computer device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the data processing method according to the first aspect or any corresponding embodiment thereof.

[0033] Fourth aspect, the present invention provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to execute the data processing method according to the first aspect or any corresponding embodiment thereof.

[0034] Fifth aspect, the present invention provides a computer program product, including computer instructions, and the computer instructions are used to cause a computer to execute the data processing method according to the first aspect or any corresponding embodiment thereof.

[0035] The present invention configures persistent memory for a database system. When a data operation needs to be performed on data to be processed, the persistent memory is used as a temporary buffer to process the data to be processed. Even when the data volume is large, disk I / O overhead can be reduced, and the performance and response speed of the database system can be greatly improved. By using a preset memory threshold to allocate a first buffer of a corresponding size, data operations can be completed in a timely and fast manner by using the persistent memory; when the data volume of the data to be processed exceeds the memory threshold, a second buffer can also be allocated separately, and a larger memory space in the persistent memory is used to perform data operations on the data to be processed, which can effectively improve the efficiency of memory-intensive data operations. Description of the Drawings

[0036] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the related art, the following will briefly introduce the accompanying drawings required for the description of the specific embodiments or the related art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0037] Figure 1 It is a schematic process diagram of sorting in the PostgreSQL database;

[0038] Figure 2 It is a schematic flowchart of the data processing method according to the embodiment of the present invention;

[0039] Figure 3 It is a schematic architecture diagram of a database system according to the embodiment of the present invention;

[0040] Figure 4 It is a schematic flowchart of another data processing method according to the embodiment of the present invention;

[0041] Figure 5 It is a schematic flowchart of sorting data according to the embodiment of the present invention;

[0042] Figure 6 It is a block diagram of the structure of a data processing device according to the embodiment of the present invention;

[0043] Figure 7 It is a schematic hardware structure diagram of a computer device according to the embodiment of the present invention. Specific Embodiments

[0044] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0045] PostgreSQL (Postgres Structured Query Language) is an extensible open-source object-relational database management system (ORDBMS). The features and functions of PostgreSQL include:

[0046] 1. Open Source: PostgreSQL is an open-source project that can be freely obtained, used, and modified.

[0047] 2. Object-Relational Database: While supporting the relational data model, PostgreSQL also supports user-defined data types and complex data structures, making it a powerful object-relational database.

[0048] 3. High Reliability: PostgreSQL adopts the Multi-Version Concurrency Control (MVCC) mechanism, which can provide highly reliable data consistency and transaction processing. In addition, it also has a failover and recovery mechanism to ensure data security and reliability.

[0049] 4. High Performance: PostgreSQL uses various optimization techniques, including query optimization, indexing, and replication, and supports parallel query and parallel database operations to provide excellent performance.

[0050] 5. Scalability: PostgreSQL provides many flexible scalability options, including partitioning, tablespaces, and a plugin architecture. This enables it to adapt to applications of different scales and requirements.

[0051] 6. Comprehensive Features: PostgreSQL provides many advanced features, such as full transaction support, multi-version control, replication, logical replication, full-text search, geospatial data, JSON, and XML support, etc.

[0052] 7. Multi-Language Support: PostgreSQL provides support for multiple programming languages, including C / C++, Python, Java, Perl, Ruby, etc., enabling developers to interact with the database using their preferred language.

[0053] Due to the excellent features of PostgreSQL in terms of data reliability, scalability, and good SQL standard compatibility, etc., PostgreSQL is widely used in various applications and projects, and can be used from small personal websites to large-scale enterprise solutions and cloud computing environments. It is a feature-rich, powerful, and highly customizable database management system that has been widely used in data-intensive enterprise applications, web applications, geographic information systems, data warehouses, financial data management, and other fields.

[0054] Currently, when data operations such as sorting and table joining are required on the PostgreSQL database, scheduling is performed according to the set work_mem parameter (working memory parameter, a memory threshold). Taking sorting as an example, if the memory space required for sorting is less than the work_mem parameter, in-memory sorting will be used; otherwise, external sorting using disk will be employed.

[0055] Figure 1 A schematic diagram of the process of sorting in the PostgreSQL database is shown. As Figure 1 shown, the sorting process includes the following steps:

[0056] Query statement: After starting PostgreSQL, enter a query statement containing a sorting operation, which may involve sorting all / part of the data;

[0057] Statement parsing: First, the query statement is parsed by PostgreSQL to create a basic query plan;

[0058] Plan execution: PostgreSQL executes the basic query plan to generate data blocks to be processed subsequently;

[0059] Read data to be sorted: Parse the configuration file postgresql.conf to obtain the work_mem parameter; for smaller data blocks, they can be loaded into memory, and then the Executor sorts the data blocks based on the sorting requirements; for larger data blocks, it is required to read data from disk and perform sorting operations on disk using a series of operators, save the results of in-memory sorting to disk, and perform a multi-way merge operation when all blocks are processed;

[0060] After sorting is completed, the query results are returned to the user and displayed or dumped to the required output.

[0061] Limited by the memory size, the work_mem parameter is generally in the order of megabytes, such as 64MB, etc. The memory allocation space for data operations is small, and memory-intensive operations cannot be performed on larger data sets; moreover, data operations with a large amount of data rely on disk I / O operations, the data reading and writing speed is slow, the performance is low, and the demand for temporary disk space is large, resulting in high management and maintenance costs for the database system.

[0062] Persistent memory (PMEM) is a new type of storage technology that provides a storage layer between traditional memory and traditional storage. Compared with traditional memory such as DRAM (Dynamic Random Access Memory), persistent memory not only has fast read and write speeds and volatility (data is lost after power failure), but also has the ability of data persistence, which can save data in a non-volatile medium that will not be lost when power is off. This enables it to achieve a balance between data persistence and data access speed.

[0063] Persistent memory usually appears in the form of memory modules. Compared with traditional memory, persistent memory has a larger capacity than traditional DRAM (the maximum capacity of a single persistent memory can reach 512GB), and the access speed of persistent memory can even be comparable to that of traditional memory, much faster than traditional storage media (such as hard disks or solid-state drives). Therefore, persistent memory can not only randomly access data as fast as memory, but also be used as storage to store various data structures and files. Utilizing this feature, persistent memory can be applied to various scenarios, such as big data storage, database technology, and cloud computing. Due to its large capacity, fast speed, and ability to persistently store data, it can improve system performance and data security, meeting the needs of the new generation of data storage, processing, and management.

[0064] The data processing method provided by the embodiments of the present invention adds persistent memory to the database system and uses this persistent memory to execute memory-intensive data operations. Since the access speed of persistent memory is much faster than that of disks, it can greatly improve the efficiency of data operations, thereby improving the performance and response speed of the entire database system.

[0065] In this embodiment, a data processing method is provided. Figure 2 It is a flowchart of the data processing method according to the embodiments of the present invention. Among them, this data processing method can be used in a database system, such as the management server of a database system, etc. And this database system is configured with persistent memory.

[0066] In this embodiment, the database system is transformed based on persistent memory, and the persistent memory is mounted to the database system so that this persistent memory can be used when performing data operations. For example, this database system can be PostgreSQL, MySQL, etc., and this embodiment does not make any limitations in this regard.

[0067] Figure 3 Shows a schematic diagram of an architecture of a database system. As Figure 3As shown, the database system is equipped with a disk storage system, which can store DB (Data Base) files, WAL (Write-Ahead Logging) files, etc.; and, the database system is also equipped with persistent memory, and the persistent memory is provided with a Persistent Cache, which can store frequently accessed data to improve system performance.

[0068] Moreover, parameters for partitioning memory, namely memory thresholds, are pre-configured for the persistent memory, and the memory thresholds are used to represent the memory required for data operations. Taking the database system as PostgreSQL as an example, the memory threshold can be set based on the work_mem parameter.

[0069] In this embodiment, the memory threshold of the persistent memory can be set relatively large, which can be a value close to the GB level. For example, the memory threshold is not less than 512MB, or the memory threshold is not less than 1GB. That is, the memory threshold of the persistent memory in this embodiment is much larger than the megabyte-level threshold (such as 64MB, etc.) used in traditional database systems.

[0070] To be able to implement data operations (such as query, sorting, etc.) on the database system based on the persistent memory, the persistent memory needs to be configured.

[0071] Specifically, in the operating system of the database (usually the Linux operating system), the persistent memory can be mounted to the file system in a manner similar to ordinary memory, and the file system interface is used to operate the persistent memory. Moreover, the memory management and allocation logic of the database system can be reformed, including the management of memory pools, memory blocks, and memory pages, as well as the use of APIs (Application Programming Interfaces) for interacting with the persistent memory, so as to support the allocation of the memory required for data operations in the persistent memory.

[0072] Among them, the APIs provided by PMDK (Persistent Memory Development Kit) can be used to manage the PMEM memory pool and allocate memory.

[0073] When starting the database system, the path of the persistent memory can be configured to start the persistent memory. For example, the '-M' parameter can be used to specify the path of the persistent memory, which is convenient for using the persistent memory as a temporary buffer to store the data required during the data operation process.

[0074] In this embodiment, for memory-intensive data operations of the database system, a certain amount of memory is required. These data operations include sorting, table association, etc. For these data operations, persistent memory can be allocated with corresponding buffers, namely the first buffer; moreover, the size of the space of the first buffer is consistent with a preset memory threshold.

[0075] Specifically, the preset memory threshold can be obtained, and then based on this memory threshold, a memory space of a corresponding size can be allocated from the persistent memory as a buffer for subsequent use, namely the first buffer. Generally, the first buffer is a continuous space allocated in the persistent memory.

[0076] For example, if the memory threshold is 1 GB, then a first buffer of 1 GB in size can be allocated from the persistent memory.

[0077] Step S201: In response to a request for performing a data operation on the data to be processed, determine the data volume of the data to be processed.

[0078] In this embodiment, when the client needs to perform a data operation on some data in the database system, a corresponding request will be initiated. For the sake of description, the data corresponding to this request is referred to as the data to be processed. Moreover, the data volume size of the data to be processed can be determined.

[0079] For example, to complete the current task of the client, it is necessary to sort 10 GB of data in the database. Then the 10 GB of data is the data to be processed, and its data volume is 10 GB.

[0080] Step S202: When the data volume of the data to be processed is less than the memory threshold, use the first buffer to perform a data operation on the data to be processed.

[0081] In this embodiment, if the data volume of the data to be processed is less than the memory threshold, it means that all the data to be processed can be written into the first buffer of the persistent memory. Therefore, at this time, the first buffer can be directly used to perform a data operation on the data to be processed.

[0082] Specifically, directly write the data to be processed into the first buffer, and then perform corresponding data operations (such as sorting) on the data to be processed in the first buffer. Since the first buffer is a part of the memory space in the persistent memory, directly processing the data to be processed based on the first buffer can ensure efficiency; moreover, the memory threshold can be relatively large, which can meet the requirements of most data operations.

[0083] Step S203: When the data volume of the data to be processed is greater than the memory threshold, allocate a second buffer from the persistent memory that matches the data volume size of the data to be processed, and use the second buffer to perform a data operation on the data to be processed.

[0084] Among them, if the amount of data to be processed is large and greater than the memory threshold, the first buffer cannot store all the data to be processed at this time. If the data to be processed is still processed based on the first buffer, only a part of the data can be processed; if the traditional method is used for processing, the data to be processed needs to be temporarily stored in a disk storage system such as a solid-state drive, resulting in slow data reading and writing speeds and low performance, which affects the execution efficiency of data operations.

[0085] In this embodiment, since the persistent memory space is large, there is generally additional unused memory space after allocating the first buffer. If the amount of data to be processed is greater than the memory threshold, a buffer of a size matching the amount of data to be processed, that is, a second buffer, can be allocated from the persistent memory; generally, the size of the second buffer is the same as the amount of data to be processed, so that all the data to be processed can be written into the second buffer. It can be understood that compared with the first buffer, the second buffer has a larger memory space.

[0086] In this case, the second buffer can be used to perform data operations on the data to be processed. For example, the data to be processed can be loaded into the second buffer of the persistent memory by using the PMDK interface, and then the corresponding data operations are performed on the data to be processed in the second buffer. Finally, the data operation result of the data to be processed can be obtained, and subsequent processing can be performed based on the data operation result, such as sending it to the corresponding client, etc.

[0087] The data processing method provided in this embodiment configures persistent memory for the database system. When data operations need to be performed on the data to be processed, the persistent memory is used as a temporary buffer to process the data to be processed. When the amount of data is large, it can also reduce the disk I / O overhead, and can greatly improve the performance and response speed of the database system. Allocating the first buffer of the corresponding size by using the preset memory threshold can complete data operations in time and quickly by using the persistent memory; when the amount of data to be processed exceeds the memory threshold, a second buffer can also be allocated separately, and the larger memory space in the persistent memory is used to perform data operations on the data to be processed, which can effectively improve the efficiency of memory-intensive data operations.

[0088] Another data processing method is provided in this embodiment, which can be used in a database system, such as a management server of a database system, etc. Figure 4 It is a flowchart of the data processing method according to the embodiment of the present invention, as Figure 4 shown, and the process includes the following steps.

[0089] Step S401, obtain the memory threshold, and allocate a first buffer with the size of the memory threshold from the persistent memory.

[0090] In some alternative embodiments, the above step S401, "allocate a first buffer of the memory threshold size from the persistent memory", may include steps A1 to A2.

[0091] Step A1: Create a corresponding main process for the client connecting to the database system.

[0092] Step A2: Based on the main process, allocate a first buffer of the memory threshold size for the client from the persistent memory.

[0093] In this embodiment, the database system can support access by multiple clients. As Figure 3 shown, the database system can connect to M clients and allow multiple of them to access the database system simultaneously. When a client accesses the database system, it will establish a connection with the database system. For example, the client can initiate a session connection request to establish a session connection with the database system, so as to query the data in the database system.

[0094] Specifically, for the client connected to the database system, create a process for it to perform subsequent data operations. For ease of description, the process created for the client is called the main process. The main process is used to execute SQL statements related to the client.

[0095] Moreover, the main process can, based on the function interface of the database system, apply for a buffer for data operation, i.e., the first buffer, for the client from the persistent memory.

[0096] It can be understood that if there are currently multiple clients all connected to the database system, corresponding main processes can be created for each client, and then a first buffer of the memory threshold size can be allocated for each client. And to ensure that a first buffer can be allocated for each client, the memory threshold cannot be set too large.

[0097] For example, the maximum access amount allowed by the database system can be pre-configured, that is, how many clients are allowed to access the database system simultaneously. Then the memory threshold cannot exceed the ratio of the size of the persistent memory to the maximum access amount, so as to enable a first buffer to be allocated for each client.

[0098] In this embodiment, after the client establishes a connection with the database system, allocate the first buffer for the client, that is, allocate the first buffer when the persistent memory is needed, which can avoid the memory space of the persistent memory being occupied invalidly.

[0099] Optionally, the method further includes: after the client disconnects from the database system, release the first buffer allocated for the client.

[0100] In this embodiment, after the database system completes the data operations required by the client, if the client no longer needs to access the database system, the connection between the two can be disconnected. At this time, the database system can release the first buffer previously allocated for the client to release the memory space corresponding to the first buffer for subsequent data operations. In addition, the main process of the client can also be terminated.

[0101] In this embodiment, after the client disconnects the connection, releasing the corresponding first buffer in the persistent memory can effectively improve the space utilization rate of the persistent memory.

[0102] Step S402: In response to a request for performing data operations on the data to be processed, determine the data volume of the data to be processed.

[0103] For details, please refer to Figure 2 Step S201 of the illustrated embodiment, which will not be elaborated here.

[0104] Step S403: When the data volume of the data to be processed is less than the memory threshold, use the first buffer to perform data operations on the data to be processed.

[0105] For details, please refer to Figure 2 Step S202 of the illustrated embodiment, which will not be elaborated here.

[0106] Step S404: When the data volume of the data to be processed is greater than the memory threshold, allocate a second buffer from the persistent memory that matches the data volume of the data to be processed, and use the second buffer to perform data operations on the data to be processed.

[0107] Specifically, the above step S404 "using the second buffer to perform data operations on the data to be processed" includes steps S4041 to S4043.

[0108] Step S4041: Write the data to be processed into the second buffer.

[0109] Specifically, if the data volume of the data to be processed is greater than the memory threshold, the data to be processed can be added to the second buffer based on the corresponding function interface of the database system, so that the data to be processed can be written into the second buffer.

[0110] Among them, since the space of persistent memory is generally large, after allocating the first buffer for each client, there is generally enough free space left; moreover, the possibility that multiple clients process data exceeding the memory threshold simultaneously is relatively low. Therefore, generally, the free space of persistent memory is greater than the data volume of the data to be processed. That is, at this time, a second buffer with the data volume size of the data to be processed can be allocated for the corresponding client. For example, if the data volume of the data to be processed is 20GB, then 20GB of continuous memory space in the persistent memory can be used as the second buffer, and the data to be processed is written into the second buffer.

[0111] If the free space of persistent memory is insufficient, that is, the free space of persistent memory is less than the data volume of the data to be processed, then it is possible to wait for the release of the persistent memory space, or, based on the traditional processing method, use the first buffer with the memory threshold size to perform data operations. At this time, the data to be processed needs to be temporarily stored on the disk, and the disk is temporarily used to complete the current data operation.

[0112] Step S4042: Divide the data to be processed in the second buffer into multiple data blocks, and perform data operations on each data block in parallel to generate data operation results corresponding to each data block.

[0113] In this embodiment, when the data volume of the data to be processed is large, to improve the processing efficiency of the data to be processed, the data to be processed is divided into multiple data blocks, so that corresponding data operations can be performed on each data block respectively, and each data block can be processed in parallel, effectively improving the processing efficiency.

[0114] Moreover, for each data block, the data operation result of the corresponding data block can be obtained after the data operation. It can be understood that the data operation results of each data block are the results of intermediate processing.

[0115] Optionally, the method further includes the following step B1.

[0116] Step B1: Create N child processes when the data volume of the data to be processed is greater than the memory threshold. N represents the number of child processes, N is a positive integer, and N≥2.

[0117] Moreover, the above step S4042 "Divide the data to be processed in the second buffer into multiple data blocks, and perform data operations on each data block in parallel to generate data operation results corresponding to each data block" can include the following steps B21 to B22.

[0118] Step B21: Divide the data to be processed in the second buffer into N data blocks. Among them, the N data blocks correspond to the N child processes one by one.

[0119] Step B22: Use N child processes to perform data operations on corresponding data blocks in parallel to obtain the data operation results corresponding to each child process.

[0120] In this embodiment, if the amount of data to be processed is greater than the memory threshold, multiple child processes are created for this data operation, and the number of child processes is N. For example, the main process created for the client can generate N child processes based on functions such as the fork function.

[0121] When performing data operations on the data to be processed subsequently, the data to be processed in the second buffer can be divided into corresponding numbers of data blocks, that is, N data blocks, based on the number N of child processes. That is, there is a one-to-one correspondence between child processes and data blocks. For each data block, the corresponding child process can be used to perform data operations on it to obtain the data operation results of each data block.

[0122] For example, based on the PMDK interface, the memory corresponding to each data block in the second buffer can be mapped to the address space of the child process so that each child process can perform data operations on the corresponding data block.

[0123] Optionally, the above step B1, "Create N child processes", may include steps B11 to B12.

[0124] Step B11: Obtain the preset maximum number of processes.

[0125] Step B12: Create N child processes according to the maximum number of processes; N is less than or equal to the maximum number of processes.

[0126] In this embodiment, a parameter that limits the maximum number of processes, that is, the maximum number of processes, is preset. For example, the value of the maximum number of processes can be set based on the max_parallel_workers parameter. When creating child processes, the number N of child processes cannot exceed the maximum number of processes to avoid affecting the processing efficiency due to too many child processes.

[0127] Generally, the number N of child processes is equal to the maximum number of processes. For example, if the preset maximum number of processes is 10, then 10 child processes are created when the amount of data to be processed exceeds the memory threshold.

[0128] In addition, in order to avoid a sub-process that needs to process too much data, resulting in a high load on some sub-processes, the data to be processed is generally divided equally, that is, the data to be processed is divided into N data blocks of equal size. For example, if the memory threshold is 1GB and the maximum number of processes is 10, if the amount of data to be processed is 20GB, the data to be processed can be divided into 10 data blocks, each of which is 2GB in size, and the main process creates 10 sub-processes, so that each sub-process can perform the required data operations on the corresponding data blocks in parallel.

[0129] In this embodiment, by creating N subprocesses, parallel processing can be simply and conveniently implemented. Moreover, by limiting the size of N based on the maximum number of processes, it is possible to avoid affecting the processing efficiency due to too many subprocesses.

[0130] Step S4043, summarizing the data operation results corresponding to each data block to generate the data operation results of the data to be processed.

[0131] In this embodiment, after determining the data operation results of each data block, for example, after obtaining the data operation results of the corresponding data block based on N sub-processes, these data operation results are summarized to finally obtain the data processing results of the entire data to be processed. The main process may perform the summary processing.

[0132] Optionally, after step S404 of “using the second buffer to perform data operations on the data to be processed”, the method may further include: releasing the second buffer.

[0133] In this embodiment, similar to the above process of releasing the first buffer, after the data operation of the to-be-processed data is completed using the second buffer, the second buffer occupying a larger memory space can be released so that the memory space corresponding to the second buffer can be reused later.

[0134] For example, after step B22 "using N sub-processes to perform data operations on the corresponding data blocks in parallel to obtain data operation results corresponding to each sub-process", each sub-process can be terminated; and after the main process summarizes the data processing results of the data to be processed and sends them to the downstream, for example, after sending them to the corresponding client, there is no need to use the second buffer to perform other processing on the data to be processed, so the second buffer can be released.

[0135] It is understandable that even if the client still has a session connection with the database system, the second buffer can be released in time after it is no longer used. In other words, the second buffer can be released first, and then the corresponding first buffer can be released after the client disconnects.

[0136] In this embodiment, using persistent memory to perform corresponding data operations on the data to be processed can improve the overall performance and response speed of the database system. Moreover, when the amount of data to be processed exceeds the memory threshold, the data to be processed is divided into multiple data blocks, and then each data block can be processed in parallel, so that the data operation task can be completed quickly even when the amount of data is large, ensuring the efficiency of memory-intensive data operations.

[0137] Since sorting is a common operation in the database system, and operations such as queries also involve sorting operations, this embodiment takes the data operation as the sorting operation as an example to illustrate how to complete the sorting operation based on the method provided in this embodiment. As Figure 5 shown, the process of sorting data includes steps S501 to S513.

[0138] Step S501, configure persistent memory for the database system, and set the memory threshold and the maximum number of processes.

[0139] In this embodiment, the database system uses PostgreSQL, and sets the memory threshold and the maximum number of processes based on the work_mem parameter and the max_parallel_workers parameter respectively. In this embodiment, work_mem = 1GB, max_parallel_workers = 10, that is, the memory threshold is 1GB, and the maximum number of processes is 10.

[0140] Step S502, start the database system and the persistent memory.

[0141] Step S503, the client establishes a session connection with the database system, and creates a main process for this client.

[0142] Step S504, obtain the memory threshold, and allocate a first buffer with the size of the memory threshold from the persistent memory.

[0143] Among them, after the client connects to the database system, the main process and the first buffer can be created synchronously, or the first buffer can be applied based on the main process. In this embodiment, the size of the first buffer is 1GB.

[0144] Step S505, obtain the sorting operation request of the client for the data to be processed, and determine the amount of the data to be processed.

[0145] Among them, the main process can parse the SQL query statement of the sorting operation request, generate a corresponding query plan, and determine the amount of the data to be processed.

[0146] Step S506: Determine whether the data volume of the data to be processed is less than the memory threshold. That is, determine whether the data volume of the data to be processed is less than 1 GB. If it is less than the memory threshold, proceed to step S507; otherwise, proceed to step S508.

[0147] Step S507: The main process writes the data to be processed into the first buffer and sorts the data to be processed to obtain the sorted result of the data to be processed.

[0148] Among them, the main process can perform quicksort on the data to be processed in the first buffer.

[0149] Step S508: Allocate a second buffer with the same size as the data volume of the data to be processed from the persistent memory and write the data to be processed into the second buffer.

[0150] Step S509: The main program generates a corresponding number of child processes according to the maximum number of processes and divides the data to be processed into corresponding data blocks.

[0151] As described above, if the maximum number of processes is 10, 10 child processes can be generated and the data to be processed in the second buffer is divided into 10 data blocks.

[0152] Step S510: Each child process sorts the corresponding data block to obtain the sorted result of each data block.

[0153] Among them, each child process can perform quicksort on the corresponding data block. Each child process has an independent memory space, which can ensure the isolation and security of data processing.

[0154] Step S511: The main process aggregates and sorts the sorted results of each child process to finally obtain the sorted result of the entire data to be processed.

[0155] Specifically, the main process can perform merge sort on the sorted results of each child process to obtain the sorted result of all data. And after the sorting is completed, each child process can be terminated.

[0156] Among them, the persistent memory is mapped to a shared memory area, which can be shared by multiple processes, enabling parallel sorting of data blocks and further accelerating the sorting process compared to traditional external sorting.

[0157] Step S512: Return the sorted result of the data to be processed to the client.

[0158] Among them, if the second buffer is applied, the second buffer can be released at this time.

[0159] Step S513: After the client disconnects the session connection, release the first buffer and terminate the main process of the client.

[0160] In this embodiment, the function interface of PostgreSQL can be used to implement the related operations of the in-memory sorting buffer. These operations include creating an in-memory sorting buffer, adding elements to the buffer, retrieving elements from the buffer, and releasing the buffer, etc.

[0161] The data processing method provided in this embodiment uses persistent memory to perform corresponding data operations on the data to be processed. The access speed of persistent memory is much faster than that of a disk, which can effectively reduce the disk read and write overhead involved in sorting operations in the database system, thereby improving the response speed of the entire system. Moreover, persistent memory reduces the need for temporary disk space required for data operations such as sorting, simplifies the management and maintenance of the database system, and reduces the demand for disk storage space. During the execution of data operations, the data to be processed is stored in persistent memory, and even in the event of a system crash or power failure, the data can be safely saved, with high security. Persistent memory can be easily added to existing database systems such as PostgreSQL, has high scalability, and a wide range of applications.

[0162] The embodiment of the present invention also provides a database system. The schematic diagram of the architecture of this database system can be seen in Figure 3 . As Figure 3 shown, the database system is equipped with a disk storage system and persistent memory. The disk storage system is used to store DB files, WAL files, etc.; the persistent cache can be mapped as a part of memory to store frequently accessed data and improve system performance.

[0163] Among them, the memory management and allocation logic of the database system can be reformed, including the management of memory pools, memory blocks, and memory pages, as well as the use of APIs (application programming interfaces) for interacting with persistent memory, so as to support the allocation of memory required for data operations in persistent memory.

[0164] For memory-intensive operations such as sorting operations, the related operations of the database system are optimized to be able to use a memory space of the memory threshold size, that is, the first buffer, in persistent memory, and to ensure the read and write security and data integrity at this time;

[0165] Regarding the data synchronization and durability issues of persistent memory, corresponding countermeasures need to be implemented for different operation types. For example, for write operations, specific APIs and functions need to be used to achieve data synchronization and persistence to persistent memory; for read operations, corresponding APIs and functions also need to be used to read corresponding data from persistent memory.

[0166] In addition, to ensure the correctness and stability of performing memory-intensive operations using persistent memory, the database system needs to be tested in advance. After the test passes, memory-intensive operations are performed on the database system configured with persistent memory. For the specific data processing method, refer to the above embodiments, which will not be elaborated here.

[0167] In this embodiment, a data processing device is further provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be elaborated again. As used hereinafter, the term "module" may be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0168] This embodiment provides a data processing device applied to a database system. The database system is configured with persistent memory, and the persistent memory includes a first buffer with a memory threshold size, as Figure 6 shown. The device includes:

[0169] A request processing module 601, configured to determine the data volume of the data to be processed in response to a request for performing a data operation on the data to be processed;

[0170] A first processing module 602, configured to perform a data operation on the data to be processed using the first buffer when the data volume of the data to be processed is less than the memory threshold;

[0171] A second processing module 603, configured to allocate a second buffer matching the data volume of the data to be processed from the persistent memory and perform a data operation on the data to be processed using the second buffer when the data volume of the data to be processed is greater than the memory threshold.

[0172] In some optional implementation manners, the second processing module 603 performing a data operation on the data to be processed using the second buffer includes:

[0173] Writing the data to be processed into the second buffer;

[0174] Dividing the data to be processed in the second buffer into multiple data blocks, and performing data operations on each data block in parallel to generate data operation results corresponding to each data block;

[0175] Performing a summary process on the data operation results corresponding to each data block to generate a data operation result of the data to be processed.

[0176] In some optional implementation manners, the second processing module 603 is further configured to:

[0177] When the data volume of the data to be processed is greater than the memory threshold, create N child processes;

[0178] The second processing module 603 divides the data to be processed in the second buffer into multiple data blocks, and performs data operations on each data block in parallel to generate data operation results corresponding to each data block, including:

[0179] Divide the data to be processed in the second buffer into N data blocks;

[0180] Use N child processes to perform data operations on the corresponding data blocks in parallel to obtain data operation results corresponding to each child process.

[0181] In some alternative embodiments, the second processing module 603 creates N child processes, including:

[0182] Obtain a pre-set maximum number of processes;

[0183] Create N child processes according to the maximum number of processes; N is less than or equal to the maximum number of processes.

[0184] In some alternative embodiments, after the second processing module 603 performs data operations on the data to be processed using the second buffer, it is further configured to:

[0185] Release the second buffer.

[0186] In some alternative embodiments, the device further includes an allocation module, configured to:

[0187] Create a corresponding main process for the client connecting to the database system;

[0188] Allocate a first buffer of the memory threshold size for the client from the persistent memory based on the main process.

[0189] In some alternative embodiments, the first processing module 602 is further configured to:

[0190] After the client disconnects from the database system, release the first buffer allocated for the client.

[0191] The further function descriptions of the above-mentioned various modules and units are the same as those in the corresponding above embodiments, and will not be repeated here.

[0192] The data processing device in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, including a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0193] An embodiment of the present invention further provides a computer device having the above-mentioned Figure 6 data processing device.

[0194] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of a computer device provided by an alternative embodiment of the present invention. As shown in Figure 7 , the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a set of blade servers, or a multi-processor system). Figure 7 In

[0195] , a single processor 10 is taken as an example.

[0196] The memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.

[0197] The memory 20 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 may optionally include a memory remotely provided with respect to the processor 10, and these remote memories can be connected to the computer device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0198] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk, or a solid-state drive; the memory 20 may further include a combination of the above types of memories.

[0199] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or communication networks.

[0200] The embodiments of the present invention also provide a computer-readable storage medium. The method according to the embodiments of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and to be stored in a local storage medium, so that the method described herein can be stored in such software processed on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium may also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiments is implemented.

[0201] A part of the present invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the present invention through the operations of the computer. Those skilled in the art should understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executes the instructions, or the computer compiles the instructions and then executes the corresponding compiled program, or the computer reads and executes the instructions, or the computer reads and installs the instructions and then executes the corresponding installed program. Herein, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible by the computer.

[0202] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations should all be covered within the protection scope of the present invention.

Claims

1. A data processing method, applied to a database system, characterized in that: The database system is configured with a persistent memory, the persistent memory includes a first buffer of a memory threshold size, and the method includes: In response to a request for performing a data operation on the data to be processed, determining a data volume of the data to be processed; When the amount of the data to be processed is less than the memory threshold, using the first buffer to perform data operations on the data to be processed; When the amount of the data to be processed is greater than the memory threshold, a second buffer matching the amount of the data to be processed is allocated from the persistent memory, and the second buffer is used to perform data operations on the data to be processed.

2. The method according to claim 1, characterized in that The using the second buffer to perform data operation on the data to be processed includes: Writing the data to be processed into the second buffer; Dividing the data to be processed in the second buffer into a plurality of data blocks, and performing data operations on each data block in parallel to generate data operation results corresponding to each data block; The data operation results corresponding to each of the data blocks are aggregated to generate the data operation results of the data to be processed.

3. The method according to claim 2, characterized in that The method further comprises: When the amount of the data to be processed is greater than the memory threshold, N child processes are created, where N is a positive integer; The step of dividing the data to be processed in the second buffer into a plurality of data blocks, and performing data operations on each data block in parallel to generate data operation results corresponding to each data block includes: Dividing the to-be-processed data in the second buffer into N data blocks, wherein the N data blocks correspond one-to-one to the N sub-processes; The N sub-processes are used to perform data operations on corresponding data blocks in parallel to obtain data operation results corresponding to each sub-process.

4. The method according to claim 3, characterized in that The step of creating N subprocesses includes: Get the preset maximum number of processes; Create N child processes according to the maximum number of processes; N is less than or equal to the maximum number of processes.

5. The method according to claim 1, characterized in that After performing data operation on the data to be processed by using the second buffer, the method further includes: Release the second buffer.

6. The method according to any one of claims 1 to 5, characterized in that The method further comprises: Creating a corresponding main process for a client connecting to the database system; A first buffer of the memory threshold size is allocated to the client from the persistent memory based on the main process.

7. The method according to claim 6, characterized in that The method further comprises: After the client is disconnected from the database system, the first buffer allocated to the client is released.

8. A data processing device, applied to a database system, characterized in that: The database system is configured with a persistent memory, the persistent memory includes a first buffer of a memory threshold size, and the device includes: A request processing module, configured to determine the amount of the data to be processed in response to a request for performing a data operation on the data to be processed; A first processing module, configured to perform data operations on the data to be processed by using the first buffer when the amount of the data to be processed is less than the memory threshold; The second processing module is used to allocate a second buffer matching the data volume of the data to be processed from the persistent memory when the data volume of the data to be processed is greater than the memory threshold, and use the second buffer to perform data operations on the data to be processed.

9. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the data processing method according to any one of claims 1 to 7 by executing the computer instructions.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the data processing method according to any one of claims 1 to 7.