An Asynchronous I / O Implementation Method, Device and Medium under NUMA Architecture
By allocating back-end processes and read and write queues on NUMA nodes, combined with batch read and write and merge request mechanisms, the performance loss problem caused by buffers accessing memory across nodes under NUMA architecture is solved, asynchronous I/O operations are realized, and database performance is improved.
Patent Information
- Application Number
- CN202510571283.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-05-06
AI Technical Summary
Under the NUMA architecture, the buffer read and write operations of PostgreSQL database frequently access memory across nodes, resulting in performance losses. The existing technology has failed to effectively optimize the memory access efficiency of buffers.
By allocating backend processes on the NUMA node, reading and writing requests are added to the read and write queue of the target NUMA node, and the read and write control process is batch managed by the read and write control process, combining the status flag bits and merge request mechanism of the I/O control structure to realize asynchronous I/O operation.
It reduces the cost of cross-node memory access, improves memory access efficiency, optimizes database performance, and reduces system overhead and latency.
Smart Images

Figure CN120086257B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital data processing, and particularly to an asynchronous I / O implementation method, device, and medium under a NUMA architecture. Background Art
[0002] In modern NUMA-type architecture operating systems, the physical hardware resources are divided into multiple independent NUMA nodes to improve the scalability and resource allocation efficiency of multi-core processors. Each NUMA node is essentially a self-contained computing unit, which integrates some CPU cores and their directly associated local memory. The core feature of this architecture lies in the asymmetry of memory access. That is to say, when a CPU core accesses the local memory of its affiliated NUMA node, low-latency and high-bandwidth data transmission can be achieved through a high-speed interconnect bus.
[0003] When the PostgreSQL database performs file read and write operations, it usually does not perform file read and write operations every time it accesses the file on the physical storage medium. The relational pages are saved in the buffer cache to balance the access time to the physical storage medium and memory. When data needs to be read from the physical storage medium, the data is stored in the buffer in memory; when the buffer in memory needs to be flushed to disk, the data in the buffer in memory is written to the physical storage medium.
[0004] The existing PostgreSQL database implements the method of reading and writing the buffer by calling the read and write functions of the operating system. For the PostgreSQL database running on a NUMA operating system, in the scenario of disk I / O that can be asynchronously executed, the processing of I / O may use a continuous memory area located in non-local memory as the buffer used before reading and writing. This buffer processing method increases the number of accesses to remote memory and causes performance loss due to cross-NUMA memory access. Summary of the Invention
[0005] To solve the above problems, the present invention proposes an asynchronous I / O implementation method under a NUMA architecture, including:
[0006] Based on the main process of the PostgreSQL database, multiple backend processes generated for different front-end query requests are allocated to different NUMA nodes; wherein, the backend processes are used to apply for reading and writing the buffer.
[0007] When the backend process generates a read / write request for the buffer, add the read / write request to the read / write queue of the target NUMA node corresponding to the buffer.
[0008] Call the read-write control process of the target NUMA node, and poll the read-write queue to add the pending read-write requests in the read-write queue as linked list elements to the I / O linked list in the I / O control structure corresponding to the read-write control process;
[0009] Poll the I / O control structure through the read-write control process to perform batch read-writes on the linked list elements included in the I / O linked list in a preset state, so as to implement asynchronous I / O for the PostgreSQL database.
[0010] In an implementation manner of the present invention, adding the pending read-write requests in the read-write queue as linked list elements to the I / O linked list in the I / O control structure corresponding to the read-write control process specifically includes:
[0011] Obtain the status flag bit in the I / O control structure corresponding to the read-write control process;
[0012] When the status flag bit is not in the specified state, add the pending read-write requests in the read-write queue as linked list elements to the I / O linked list in the I / O control structure corresponding to the read-write control process; wherein, the specified state is processing I / O, and the linked list element is an I / O structure.
[0013] In an implementation manner of the present invention, performing batch read-writes on the linked list elements included in the I / O linked list in a preset state specifically includes:
[0014] According to the status flag bit, when the I / O control structure is in a preset state, obtain the number of I / O linked list elements from the I / O control structure; wherein, the preset state is processing completed and harvesting completed;
[0015] Mark the status flag bit as the specified state, and perform batch read-writes on the linked list elements in the I / O linked list that meet the number of I / O linked list elements;
[0016] After completing the batch read-writes on the linked list elements, mark the status flag bit as the preset state.
[0017] In an implementation manner of the present invention, after adding the pending read-write requests in the read-write queue as linked list elements to the I / O linked list in the I / O control structure corresponding to the read-write control process, the method further includes:
[0018] Determine the logical address of the buffer required by the pending read-write request and the previous pending read-write request of the pending read-write request;
[0019] Determine whether the read / write request to be processed needs to be merged according to the logical address.
[0020] In an implementation manner of the present invention, determining whether the read / write request to be processed needs to be merged according to the logical address specifically includes:
[0021] Determine whether the logical addresses of the read / write request to be processed and the previous read / write request to be processed are consecutive according to whether the sum of the start address of the buffer corresponding to the previous read / write request to be processed and the previous read / write length corresponding to the previous read / write request to be processed is the same as the start address of the buffer corresponding to the read / write request to be processed;
[0022] If so, perform a merge request on the read / write request to be processed and the previous read / write request to be processed to accumulate the read / write length corresponding to the read / write request to be processed and the previous read / write length.
[0023] In an implementation manner of the present invention, adding the read / write request to the read / write queue of the target NUMA node corresponding to the buffer specifically includes:
[0024] Attempt to add the read / write request to the local read / write queue of the local NUMA node corresponding to the buffer to determine whether the local read / write queue is full;
[0025] If so, determine the access cost of the read / write request for each NUMA node;
[0026] According to the access cost, select the NUMA node with the minimum access cost from other NUMA nodes except the local NUMA node as the target NUMA node for adding the read / write request, and add the read / write request to the read / write queue of the target NUMA node.
[0027] In an implementation manner of the present invention, attempting to add the read / write request to the local read / write queue of the local NUMA node corresponding to the buffer to determine whether the local read / write queue is full specifically includes:
[0028] Obtain the free element pointer and the tail pointer in the local read / write queue;
[0029] Compare the positions corresponding to the free element pointer and the tail pointer, and determine that the local read / write queue is full when the next position of the free element pointer is the tail pointer.
[0030] In an implementation manner of the present invention, after batch reading and writing the linked list elements included in the I / O linked list that meet the preset state, the method further includes:
[0031] When the batch reading and writing is abnormal, determine the corresponding abnormal type;
[0032] If the abnormal type is a system call interruption, perform the batch reading and writing of the linked list elements again;
[0033] If the abnormal type is not a system call interruption, save the abnormal state corresponding to the batch reading and writing according to the reading and writing request type corresponding to the batch reading and writing.
[0034] An embodiment of the present invention provides an asynchronous I / O implementation device under a NUMA architecture. The device includes:
[0035] At least one processor;
[0036] And a memory communicatively connected to the at least one processor;
[0037] Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor, so that the at least one processor can execute an asynchronous I / O implementation method as described in any one of the above.
[0038] An embodiment of the present invention provides a non-volatile computer storage medium storing computer-executable instructions, and the computer-executable instructions are set as:
[0039] An asynchronous I / O implementation method under a NUMA architecture as described in any one of the above.
[0040] The asynchronous I / O implementation method proposed by the present invention can bring the following beneficial effects:
[0041] The backend process is no longer responsible for buffer reading and writing, but only for reading and writing requests. The asynchronous reading and writing requests in the buffer in the memory of the same NUMA node are handed over to the reading and writing control process running on the CPU core located on the same NUMA node for proxy processing. When the reading and writing control process accesses the buffer located in the local memory, compared with accessing the remote memory, the access cost is reduced. Combined with the batch processing of I / O, the result of reducing the performance loss of the database and optimizing the database performance can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The drawings described herein are used to provide a further understanding of the present invention, and constitute a part of the present invention. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0043] Figure 1 It is a schematic flowchart of an asynchronous I / O implementation method under a NUMA architecture provided by an embodiment of the present invention;
[0044] Figure 2 A schematic diagram of a NUMA architecture provided by an embodiment of the present invention;
[0045] Figure 3 A schematic diagram of a read-write queue data structure provided by an embodiment of the present invention;
[0046] Figure 4 A schematic diagram of a data structure provided by an embodiment of the present invention;
[0047] Figure 5 A schematic diagram of the structure of an asynchronous I / O implementation device under a NUMA architecture provided by an embodiment of the present invention. Detailed implementation manners
[0048] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with specific embodiments of the present invention and the corresponding drawings. Obviously, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0049] PostgreSQL (usually abbreviated as Postgres) is an open-source relational database management system, known for its powerful functions, high reliability and scalability. It supports the SQL language and provides rich extension capabilities, suitable for application scenarios of various scales. In the PostgreSQL database system running on a NUMA platform, when its buffer is applied for, initialized, allocated, read and written, it does not distinguish which local memory of which NUMA node it is located in, and the process responsible for reading and writing to this buffer is also scheduled by the operating system to select a CPU core according to the scheduling algorithm in the system. The CPU core for scheduling and this buffer are not necessarily located on the same NUMA node. This processing method of the buffer increases the number of accesses to remote memory and causes performance loss.
[0050] The following will detail the technical solutions provided by each embodiment of the present invention in conjunction with the drawings.
[0051] As Figure 1 shown, an asynchronous I / O implementation method under a NUMA architecture provided by an embodiment of the present invention includes:
[0052] S101: Based on the main process of the PostgreSQL database, allocate multiple backend processes generated for different front-end query requests to different NUMA nodes; wherein, the backend processes are used to apply for reading and writing to the buffer.
[0053] As Figure 2 shown in the schematic diagram of a NUMA architecture, when the PostgreSQL database starts, a main process Postgres is created, which is responsible for managing the life cycle of the entire database system, including listening for connection requests from clients, starting and managing backend processes, maintaining the global state of the database, etc. When a user initiates a query operation at the front end, such as querying data in the database, performing insert, update, or delete operations, these front-end query requests will be sent to the PostgreSQL database. When the main process receives a front-end query request, it forks a backend process for each request and runs these backend processes on different CPU cores according to the scheduling algorithm, that is, distributes them to different NUMA nodes. The difference between the embodiment of the present invention and the traditional I / O implementation method of NUMA is that the backend process only makes read / write requests and does not responsible for the actual read / write processing.
[0054] S102: When a read / write request for the buffer is generated by the backend process, add the read / write request to the read / write queue of the target NUMA node corresponding to the buffer.
[0055] In PostgreSQL, the buffer is used to cache data pages in the memory. When the backend process initiates a read / write request, the NUMA node will first check whether there is a corresponding data page in the buffer. If so, the read / write operation is directly performed; if not, the data page is read from the disk and loaded into the buffer. The role of the buffer is to reduce disk I / O operations and improve the performance of the database.
[0056] In the NUMA architecture, the memory access latency of different NUMA nodes is different. If the backend process randomly performs read / write operations on the buffer, it may frequently access memory across NUMA nodes, resulting in an increase in memory access latency and a decrease in system performance. To reduce this cross-node memory access and improve memory access efficiency, it is necessary to centralize the read / write requests of the backend process for the buffer to the read / write queue of the corresponding target NUMA node. The read / write queue stores the read / write requests for the target NUMA node in order. When multiple backend processes simultaneously initiate read / write requests for the buffer of the same NUMA node, these requests will be added to the read / write queue and then processed sequentially according to certain rules to avoid conflicts and ensure data consistency.
[0057] In one embodiment, under the NUMA architecture, in order to make full use of the local memory resources of each node and improve the memory access efficiency, the backend process usually gives priority to sending read / write requests to the read / write queue of the local NUMA node. However, the capacity of the local read / write queue is limited. When the queue is full, if requests continue to be added, it will cause request backlog and affect the read / write performance. Therefore, when adding read / write requests to the read / write queue of the target NUMA node, it is necessary to reasonably screen the read / write requests according to the access costs of different NUMA nodes to achieve dynamic balance in distributing read / write requests. Thus, when the local read / write queue is full, some read / write requests can be transferred to other nodes with lower access costs, avoiding blockage of the local queue due to overload, reducing the latency of critical requests, and optimizing resource utilization at the same time.
[0058] First, analyze the location of the local NUMA node of the memory where the written buffer is located. Then, try to add the read / write request to the local read / write queue of the local NUMA node corresponding to the buffer to determine whether the local read / write queue is full. If the queue is not full, the read / write request will be successfully added to the local read / write queue, and then the local NUMA node will process it according to the queue order. If the local queue is full and the request cannot be added to the local queue, at this time, it is necessary to find other NUMA nodes as the target NUMA node.
[0059] To select the optimal node, it is necessary to calculate the access cost of the read / write request for all other NUMA nodes except the local NUMA node. The calculation of the access cost usually considers multiple factors, such as the physical distance between nodes, the bandwidth of the communication link, the current load of the node, etc. Generally speaking, the farther the distance, the lower the bandwidth, and the higher the load, the greater the access cost. According to the calculated access costs of each NUMA node, select the NUMA node with the smallest access cost from other nodes except the local NUMA node. This node is the target NUMA node of the current read / write request. Finally, add the read / write request to the read / write queue of the target NUMA node, and the target NUMA node is responsible for processing this request. If no free space is found in the read / write queue of each NUMA node, spin waiting is required, and the read / write request queues are continuously polled in ascending order of the memory access costs of its own NUMA node and other NUMA nodes until there is free space in a certain read / write queue and then the read / write queue is inserted.
[0060] It should be noted that when judging whether the read / write queue is full, it is necessary to judge through the free element pointer and the tail pointer. As Figure 3Schematic diagram of a read-write queue data structure shown. The read-write queue is maintained by a read-write control process. The read-write queue is an idle queue. Each time the backend process performs a read or write operation, it needs to use an idle space in the read-write queue. When it needs to use an idle space, it locates according to the idle element pointer and fills the information into the queue member structure at the position of the idle element. Each queue member contains a buffer descriptor, physical file information, I / O operation type, whether it is idle, a pointer to the next element, etc., to ensure that read-write requests are processed in an orderly manner. The idle element pointer pointing to the idle element node represents the current position in the local read-write queue that can be used to store new read-write requests. It moves as requests are added and always indicates the next available idle node to be written to. The tail pointer pointing to the end of the request marks the last occupied request node in the local read-write queue and is used to define the boundary of the requests stored in the queue.
[0061] When the backend process attempts to add a read-write request to the local read-write queue, it first obtains two key pointers of the queue, namely the idle element pointer and the tail pointer. It makes a logical comparison of the positions corresponding to the idle element pointer and the tail pointer. In the Figure 3 queue structure shown, each node is connected by a pointer to the next element. When a read-write request is added to the read-write queue, it will first be added to the first idle position in the logical order according to the queue order. Then, after the read-write request is written, the idle element pointer is moved backward. If the next position of the idle element pointer is the tail pointer, it means that the idle element pointer currently points to the last idle node, and the next position of this node has reached the occupied node pointed to by the tail pointer. At this time, there is no extra idle node in the queue to store new requests, that is, the queue is full.
[0062] S103: Invoke the read-write control process of the target NUMA node to poll the read-write queue, and add the pending read-write requests in the read-write queue as linked list elements to the I / O linked list in the I / O control structure corresponding to the read-write control process.
[0063] There is a read-write control process on each NUMA node. The read-write control process has two responsibilities. One of the responsibilities is to maintain the read-write queue. The read-write control process of the target NUMA node will regularly poll the read-write queue. The read-write queue returns the current first pending read-write request to be processed to the read-write control process when polled by the read-write control process. If there are pending read-write requests on the queue and the I / O control structure is not in the state of I / O being processed, the pending request is added as a linked list element to the I / O linked list in the I / O control structure. The I / O linked list is a linked list structure in the I / O control structure used to store pending I / O requests. These requests exist in the form of I / O structures in the linked list, which is convenient for management and scheduling.
[0064] As Figure 4 shown in a schematic diagram of a data structure, the I / O control structure is responsible for coordinating and managing the core data structure for processing I / O requests of a local node. The I / O control structure stores an I / O linked list that has not been processed, a status flag bit, the number of elements in the I / O linked list, the number of elements in the I / O linked list that have been processed, and a buffer pointer. The linked list elements in the I / O linked list are I / O structures, and the I / O structure contains the status information of the current I / O, the buffer descriptor to be used, the physical file information, the I / O operation type, and a pointer to the next I / O structure. The status flag bit of the I / O control structure is used to indicate the usage status of the current I / O control structure. There are a total of 7 statuses, which are: idle (no requests), processing I / O, I / O processing exception, merging requests, waiting for harvesting after processing is complete, harvesting in progress after processing is complete, and harvesting completed after processing is complete. The buffer pointer is used to indicate the position of the buffer data when reading data during an I / O operation.
[0065] A large number of read and write requests are generated by the PostgreSQL database. To efficiently process these read and write requests, it is necessary to call the read and write control processes on each NUMA node for management. The I / O control structure is the key data structure for centrally managing these I / O operations. However, I / O operations need to be carried out sequentially, and only one request can be processed at the same time. Therefore, it is necessary to control the addition of requests through the status flag bit to avoid conflicts and errors caused by multiple requests being processed simultaneously.
[0066] In one embodiment, when the read and write control process needs to add a to-be-processed read and write request in the read and write queue to the I / O linked list of the I / O control structure, it will first obtain the status flag bit in the I / O control structure. The status flag bit can reflect the usage status of the current I / O control structure, thereby determining whether a new request can be added. After obtaining the status flag bit, it will be judged whether it is in a specified state. The specified state refers to processing I / O. If it is in the specified state, it means that the I / O control structure is processing an I / O request. At this time, a new request cannot be added to the I / O linked list, and it is necessary to wait for the request to be processed and the status flag bit to change before trying again. When the status flag bit is not in the specified state, it means that the I / O control structure is not currently processing an I / O request and can receive a new request. Then, the new read and write request can be directly added to the I / O linked list as a linked list element.
[0067] It should be noted that in I / O operations, frequent read and write requests will increase the system overhead. Especially on storage devices such as disks, each I / O operation has a certain time and resource cost. When the data accessed by multiple read and write requests is continuous or adjacent in the buffer, merging these requests into one request for processing can reduce unnecessary I / O operations and improve the read and write performance of the database. Therefore, after adding the read and write requests to be processed to the I / O linked list, it is necessary to determine whether request merging can be performed.
[0068] Specifically, after adding the read and write requests to be processed to the I / O linked list, determine the logical address of the buffer required by the read and write request to be processed. At the same time, the logical address of the buffer required by the previous read and write request to be processed is also determined. The logical address can be obtained by reading the information recorded in the I / O structure. The logical address can be information such as the start address and length of the buffer, which is used to determine the specific position of the data involved in the request in the buffer. Based on the two obtained logical addresses, determine whether these two requests can be merged. The basis for the determination is usually the continuity or adjacency of the logical addresses.
[0069] Specifically, obtain the start address of the buffer corresponding to the previous read and write request to be processed, and the previous read and write length corresponding to this request. Then, add the start address of the buffer corresponding to the previous read and write request to be processed to the previous read and write length to obtain the end address of the data involved in the previous request in the buffer. For example, if the start address of the buffer of the previous read and write request to be processed is start_addr_prev and the previous read and write length is length_prev, then the end address of the previous request is end_addr_prev = start_addr_prev + length_prev. Then, obtain the start address start_addr_current of the buffer corresponding to the current read and write request to be processed, and compare whether the end address end_addr_prev of the previous request is the same as the start address start_addr_current of the buffer corresponding to the current read and write request to be processed. If they are the same, it means that the data involved in these two requests is continuous in the buffer, that is, their logical addresses are continuous. When it is determined that the logical addresses of the read and write request to be processed and the previous read and write request to be processed are continuous, the operating system will merge these two requests.
[0070] The specific operation of merging is to accumulate the read / write length length_current corresponding to the read / write request to be processed with the previous read / write length length_prev to obtain a new read / write length new_length = length_prev + length_current. At the same time, update the relevant request information and treat these two requests as a whole for subsequent processing.
[0071] In addition, if the logical addresses of two read / write requests are not continuous but very close, they may also be merged according to specific strategies. For example, if the interval between two read / write requests is less than a preset time threshold, merging these two requests can reduce the number of I / O operations and improve efficiency.
[0072] S104: Poll the I / O control structure through the read / write control process to perform batch read / write operations on the linked list elements included in the I / O linked list in a preset state, so as to implement asynchronous I / O for the PostgreSQL database.
[0073] In addition to maintaining the read / write queue, the read / write control process is also responsible for periodically polling the I / O control structure, scheduling read / write operations by reading the information in the control structure, and performing batch read / write operations on the linked list elements included in the I / O linked list that meet the preset state. In the traditional synchronous I / O mode, the front end will block and wait for the operation to complete after initiating an I / O request, which will cause waste of CPU resources. Especially in high-concurrency scenarios, it will seriously affect the performance and response speed of the database. In the embodiment of the present invention, the I / O processing that can be performed asynchronously in the buffer area of the same NUMA node memory is handed over to the read / write control process running on the CPU core located on the same NUMA node for proxy processing. By polling the I / O control structure through the read / write control process and performing batch read / write operations, asynchronous I / O can be better implemented and the overhead of I / O operations can be reduced.
[0074] In one embodiment, the read / write control process polls the I / O control structure at a certain time interval. During the polling process, it determines whether the state of the I / O control structure meets the preset state according to the status flag bit. Here, the preset state refers to the state of processing completion and harvesting completion. When the I / O control structure is in the state of processing completion and harvesting completion and other states are not true, it indicates that the I / O control structure has completed the processing of all requests and resource recovery and is in a reusable state. At this time, the control system obtains the number of I / O linked list elements from the I / O control structure, and the number of I / O linked list elements is the number of read / write requests that need to be processed in the subsequent batch read / write operation.
[0075] To avoid concurrency issues during batch read and write operations, before performing batch read and write operations, the operating system marks the status flag as a specified status. The specified status indicates that the I / O control structure is performing batch read and write operations, and other processes or operations need to wait for this operation to complete. After marking the status flag, batch read and write operations are performed on the linked list elements in the I / O linked list that meet the number of I / O linked list elements. Specifically, the read and write requests represented by these linked list elements are combined into a group and sent to the corresponding buffer for read and write processing at one time. When the batch read and write operation is completed, the status flag is marked as the preset status again, which means that the I / O control structure has completed the current round of batch read and write tasks, and has cleaned up and harvested relevant resources, and can accept new read and write tasks again.
[0076] Through the above batch read and write mechanism based on the status flag, the efficiency of I / O operations can be effectively improved and system overhead can be reduced. At the same time, the use of the status flag ensures the order and correctness of I / O operations, avoids the occurrence of concurrency issues, and through marking the status flag, the entire scheduling process of the read and write control process forms a closed loop, thus achieving efficient query of the PostgreSQL database.
[0077] It should be noted that after the read and write control process performs batch read and write operations, abnormal situations may occur due to reasons such as hardware failures, insufficient system resources, and software errors. After the batch read and write operation, it is necessary to check and handle possible exceptions to ensure the reliability of the system and the integrity of the data.
[0078] Specifically, when the batch read and write operation is completed, first check whether there are any exceptions. If it is detected that there are exceptions in the batch read and write, it is necessary to further determine the corresponding exception type. If the error code EINTR is returned, that is, the system call is interrupted, regardless of whether the request being processed is a merged request at this time, the batch read and write of the linked list elements needs to be executed again. In the case of non-EINTR, that is, an exception that is not a system call interruption, it is necessary to save the exception status corresponding to the batch read and write according to the type of read and write request corresponding to the batch read and write. The merged read and write requests are first split and the I / O is executed again separately. If there are still exceptions, the exception statuses are saved separately. The transaction associated with the I / O will detect the completion status of the I / O before committing. If it is not completed, it will block and wait. If there is an exception, the transaction will roll back.
[0079] Such as Figure 2A schematic diagram of a NUMA architecture is shown. The main process Postgres of the PostgreSQL database will fork out backend processes for front-end query requests. The backend processes are allocated to different NUMA nodes, and each NUMA node is responsible for processing multiple I / O requests that can be executed asynchronously. When there are read / write requests from backend processes for the buffer, according to the NUMA node corresponding to the buffer and the load situation of the read / write queues in the NUMA node, the read / write requests are distributed to the corresponding target NUMA nodes. The target NUMA nodes here can be one or more. The read / write requests are preferentially distributed to the local NUMA node. When the local read / write queue is full, the read / write requests will be inserted into the read / write queues of the corresponding NUMA nodes in ascending order of access cost. The backend processes are not responsible for performing read / write operations. After the backend processes complete the distribution of read / write requests, the read / write control process obtains the read / write requests by polling the read / write queues and writes the read / write requests into the I / O control structure. The read / write control process schedules the read / write requests that meet the preset status by periodically reading the data in the I / O control structure and changing the status of the I / O control structure, realizes the batch read / write of the permanent storage medium, and obtains the corresponding return results.
[0080] The read / write method of the PostgreSQL buffer is not optimized for the NUMA architecture design. When its backend process performs read / write operations on the buffer, if the buffer is remote memory, it will generate more access costs compared to accessing local memory, thereby affecting the overall performance of the database. In the embodiment of the present invention, the backend process no longer performs buffer read / write, but hands over the read / write requests that can be asynchronously performed on the buffer in the memory of the same NUMA node to the read / write control process running on the CPU core located on the same NUMA node for proxy processing. When the read / write control process accesses the buffer located in local memory, compared to accessing remote memory, the access cost is reduced. Combined with the batch processing of I / O, the result of reducing the performance loss of the database and optimizing the database performance can be achieved.
[0081] The above is the method embodiment proposed by the present invention. Based on the same idea, some embodiments of the present invention also provide the corresponding device and non-volatile computer storage medium for the above method.
[0082] Figure 5 It is a schematic diagram of the structure of an asynchronous I / O implementation device under a NUMA architecture provided by an embodiment of the present invention. As Figure 5 shown, it includes:
[0083] At least one processor; and,
[0084] A memory communicatively connected to at least one processor; wherein,
[0085] The memory stores instructions that can be executed by at least one processor. The instructions are executed by at least one processor, enabling the at least one processor to execute an asynchronous I / O implementation method under a NUMA architecture as described in any one of the above.
[0086] An embodiment of the present invention provides a non-volatile computer storage medium storing computer-executable instructions, and the computer-executable instructions are set as:
[0087] An asynchronous I / O implementation method under a NUMA architecture as described in any one of the above.
[0088] Each embodiment in the present invention is described in a progressive manner. For parts that are the same or similar among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the relevant parts of the method embodiments for the relevant parts.
[0089] The device and medium provided by the embodiments of the present invention correspond one by one to the method. Therefore, the device and medium also have beneficial technical effects similar to those of the corresponding method. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the device and medium will not be elaborated here.
[0090] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0091] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0092] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0093] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0094] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0095] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash RAM. The memory is an example of computer-readable media.
[0096] Computer-readable media includes both permanent and non-permanent, removable and non-removable media and can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0097] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the said element.
[0098] The above are only embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the scope of the claims of the present invention.
Claims
1. An asynchronous I / O implementation method under the NUMA architecture, characterized in that, The method includes: Based on the main process of the PostgreSQL database, allocate multiple backend processes generated for different front-end query requests to different NUMA nodes; wherein, the backend process is used to apply for reading and writing to the buffer. When the backend process generates a read / write request for the buffer, add the read / write request to the read / write queue of the target NUMA node corresponding to the buffer. Call the read / write control process of the target NUMA node to poll the read / write queue, and add the pending read / write requests in the read / write queue as linked list elements to the I / O linked list in the I / O control structure corresponding to the read / write control process. Poll the I / O control structure through the read / write control process to perform batch reading and writing on the linked list elements included in the I / O linked list in a preset state, so as to implement asynchronous I / O for the PostgreSQL database.
2. The asynchronous I / O implementation method under a NUMA architecture according to claim 1, characterized in that, Adding the pending read / write requests in the read / write queue as linked list elements to the I / O linked list in the I / O control structure corresponding to the read / write control process specifically includes: Obtain the status flag bit in the I / O control structure corresponding to the read / write control process. When the status flag bit is not in the specified state, add the pending read / write requests in the read / write queue as linked list elements to the I / O linked list in the I / O control structure corresponding to the read / write control process; wherein, the specified state is processing I / O, and the linked list element is an I / O structure.
3. The asynchronous I / O implementation method under a NUMA architecture according to claim 2, wherein Performing batch reading and writing on the linked list elements included in the I / O linked list in a preset state specifically includes: According to the status flag bit, when the I / O control structure is in a preset state, obtain the number of I / O linked list elements from the I / O control structure; wherein, the preset state is processing completed and harvesting completed. Mark the status flag bit as the specified state, and perform batch reading and writing on the linked list elements in the I / O linked list that meet the number of I / O linked list elements. After completing the batch reading and writing of the linked list elements, mark the status flag bit as the preset state.
4. The asynchronous I / O implementation method under a NUMA architecture according to claim 1, wherein After adding the pending read / write requests in the read / write queue as linked list elements to the I / O linked list in the I / O control structure corresponding to the read / write control process, the method further includes: Determine the logical address of the buffer required by the pending read / write request and the previous pending read / write request of the pending read / write request. According to the logical address, determine whether the pending read / write request needs to be merged.
5. The asynchronous I / O implementation method under the NUMA architecture according to claim 4, wherein Determining whether the pending read / write request needs to be merged according to the logical address specifically includes: Determine whether the logical addresses of the pending read / write request and the previous pending read / write request are continuous according to whether the sum of the start address of the buffer corresponding to the previous pending read / write request and the previous read / write length corresponding to the previous pending read / write request is the same as the start address of the buffer corresponding to the pending read / write request. If so, merge the to-be-processed read / write request with the previous to-be-processed read / write request to accumulate the read / write length corresponding to the to-be-processed read / write request and the previous read / write length.
6. The asynchronous I / O implementation method under a NUMA architecture according to claim 1, wherein Adding the read / write request to the read / write queue of the target NUMA node corresponding to the buffer specifically includes: Attempt to add the read / write request to the local read / write queue of the local NUMA node corresponding to the buffer to determine whether the local read / write queue is full; If so, determine the access cost of the read / write request for each NUMA node; According to the access cost, select the NUMA node with the smallest access cost from the other NUMA nodes except the local NUMA node as the target NUMA node for adding the read / write request, and add the read / write request to the read / write queue of the target NUMA node.
7. The asynchronous I / O implementation method under the NUMA architecture according to claim 6, characterized in that Attempt to add the read / write request to the local read / write queue of the local NUMA node corresponding to the buffer to determine whether the local read / write queue is full, specifically including: Obtain the free element pointer and the tail pointer in the local read / write queue; Compare the positions corresponding to the free element pointer and the tail pointer, and determine that the local read / write queue is full when the next position of the free element pointer is the tail pointer.
8. The asynchronous I / O implementation method under the NUMA architecture according to claim 1, characterized in that After batch reading and writing the linked list elements that meet the preset status, the method further includes: When an exception occurs in the batch reading and writing, determine the corresponding exception type; If the exception type is a system call interruption, perform the batch reading and writing of the linked list elements again; If the exception type is not a system call interruption, save the exception status corresponding to the batch reading and writing according to the read / write request type corresponding to the batch reading and writing.
9. An asynchronous I / O implementation device under the NUMA architecture, characterized in that, The device includes: At least one processor; And a memory communicatively connected to the at least one processor; Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute an asynchronous I / O implementation method under a NUMA architecture according to any one of claims 1-8.
10. A non-volatile computer storage medium stores computer-executable instructions, characterized in that, The computer-executable instructions are set to: An asynchronous I / O implementation method under a NUMA architecture according to any one of claims 1-8.
Citation Information
Patent Citations
Instruction prefetch-based multi-core shared memory control equipment
CN102207916A
Method for preventing node controller from deadly embrace and node controller
CN102439571A