Data interaction processing method, processor and numa system architecture

By configuring a relay cache pool in the NUMA system architecture, the problem of high latency and low efficiency when accessing storage devices across nodes is solved, and more efficient data interaction is achieved. Interacting with the target storage device through the target cache, reducing the latency of cross-node access.

CN120216392APending Publication Date: 2025-06-27PHYTIUM TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510272753.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the NUMA system architecture, when accessing storage devices across nodes, data interaction delay is high and interaction efficiency is low.

Method used

A data interaction processing method is proposed. By configuring a redirected cache pool in the NUMA system architecture, cache is allocated from the memory of different NUMA nodes for cross-node data interaction. The specific steps include determining the NUMA node where the target storage device is located, allocating the target cache to the same belong to the memory of the node, and interacting with the target storage device through the target cache.

Benefits of technology

The relay cache pool reduces the latency of accessing storage devices across nodes, improves data interaction efficiency, and makes NUMA node programs more efficient when accessing storage devices across nodes, and the amount of data that can be transferred per unit time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216392A_ABST
    Figure CN120216392A_ABST
Patent Text Reader

Abstract

The invention provides a data interaction processing method, a processor and a numa system architecture, the method is applied to a first numa node of the numa system architecture, the numa system architecture comprises a plurality of numa nodes, a transit cache pool is configured in the numa architecture, the transit cache pool comprises caches distributed from memories of different numa nodes, and the cache nodes are configured in the transit cache pool. The method comprises the steps of determining a numa node where a target storage device corresponding to a data interaction request is located; the data interaction request is generated by a program running on a first numa node; when it is determined that the numa node where the target storage device is located is a second numa node, a target cache is allocated from a transfer cache pool, and the memory of the target cache and the memory of the second numa node belong to the same memory; and performing data interaction with the target storage device through the target cache. According to the method, the data interaction efficiency during cross-node access to the storage device in the numa system architecture can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a data interaction processing method, a processor, and a NUMA system architecture. Background Art

[0002] Non-Uniform Memory Access (NUMA) is a memory architecture design for multi-processor systems. The NUMA architecture includes at least two NUMA nodes, and each NUMA node includes a processor (or processor core) and memory. In addition, storage devices can be mounted under the NUMA node.

[0003] In the above NUMA system architecture, it usually occurs that a program running on one NUMA node interacts with the storage devices of other NUMA nodes, and the data interaction latency is relatively high and the interaction efficiency is low. Summary of the Invention

[0004] Based on the above technical problems, this application proposes a data interaction processing method, a processor, and a NUMA system architecture, which can improve the data interaction efficiency when accessing storage devices across nodes in the NUMA system architecture.

[0005] A first aspect of this application proposes a data interaction processing method, which is applied to a first NUMA node of a NUMA system architecture. The NUMA system architecture includes multiple NUMA nodes, and a transfer cache pool is configured in the NUMA architecture. The transfer cache pool includes caches allocated from the memories of different NUMA nodes. The method includes:

[0006] Determine the NUMA node where the target storage device corresponding to the data interaction request is located; the data interaction request is generated by a program running on the first NUMA node;

[0007] In the case where the determined NUMA node where the target storage device is located is the second NUMA node, allocate a target cache from the transfer cache pool, and the target cache belongs to the same memory as the memory of the second NUMA node;

[0008] Perform data interaction with the target storage device through the target cache.

[0009] In some implementation manners, performing data interaction with the target storage device through the target cache includes:

[0010] Modify the cache pointer of the data interaction request to point to the target cache, so that when the protocol controller of the target storage device receives the data interaction request, it performs the data interaction operation between the target cache and the target storage device according to the data interaction request.

[0011] In some implementations, when the data interaction request is a data sending request, the method further includes: storing the data to be sent in the target cache;

[0012] The protocol controller performing the data interaction operation between the target cache and the target storage device according to the data interaction request includes:

[0013] The protocol controller stores the data stored in the target cache to the target storage device according to the data interaction request;

[0014] Or,

[0015] When the data interaction request is a data reading request, the protocol controller performing the data interaction operation between the target cache and the target storage device according to the data interaction request includes:

[0016] The protocol controller reads the data requested to be read by the data interaction request from the target storage device and stores the read data in the target cache.

[0017] In some implementations, when it is determined that the numa node where the target storage device is located is the second numa node, the method further includes: determining the original cache for performing the data interaction operation corresponding to the data interaction request, and the original cache and the memory of the first numa node belong to the same memory;

[0018] When the data interaction request is a data sending request, storing the data to be sent in the target cache includes: copying the data stored in the original cache to the target cache;

[0019] Or,

[0020] When the data interaction request is a data reading request, the method further includes: copying the data stored in the target cache to the original cache.

[0021] In some implementations, when the data interaction request is a data reading request, before copying the data stored in the target cache to the original cache, the method further includes:

[0022] Detect the data interaction request from a pre-set completion queue; the completion queue is used to store completed data interaction requests;

[0023] When the data interaction request is detected from the completion queue, copy the data stored in the target cache to the original cache.

[0024] In some implementation manners, when it is determined that the numa node where the target storage device is located is the second numa node, the method further includes:

[0025] Configure a request queue and a completion queue in the memory of the second numa node, where the request queue is used to store uncompleted data interaction requests.

[0026] In some implementation manners, allocating a target cache from the transit cache pool includes:

[0027] Query a candidate cache from the transit cache pool, where the candidate cache is an idle cache that belongs to the same memory as the memory of the second numa node and can be used to execute the data interaction request;

[0028] When the candidate cache is queried from the transit cache pool, allocate a target cache from the candidate cache;

[0029] Or,

[0030] When the candidate cache is not queried from the transit cache pool, send a creation instruction to the second numa node, where the creation instruction is used to instruct the second numa node to allocate a specific size of memory of the second numa node to the transit cache pool as a newly created cache; the specific size of memory is not less than the candidate cache;

[0031] Allocate a target cache from the newly created cache.

[0032] In some implementation manners, the transit cache pool includes multiple cache queues, each cache queue includes multiple caches of the same size, and the cache sizes of different cache queues are different;

[0033] Querying a candidate cache from the transit cache pool includes:

[0034] Query a candidate cache from the target cache queue of the transit cache pool, where the cache in the target cache queue is the smallest cache that can meet the cache required by the data interaction request.

[0035] In some implementation manners, when the second numa node allocates a specific size of memory of the second numa node to the transit cache pool as a newly created cache, the method further includes:

[0036] When the data interaction request is completed, a recycling instruction is sent to the second NUMA node, and the recycling instruction is used to instruct the second NUMA node to recycle the memory of the specific size.

[0037] In some implementation manners, the data interaction request includes a data interaction request generated by a storage performance development toolkit program running on the first NUMA node;

[0038] and / or

[0039] The target storage device includes a storage device using the NVMe protocol mounted on the second NUMA node.

[0040] A second aspect of the present application proposes a processor, which is applied to the first NUMA node of a NUMA system architecture. The NUMA architecture includes multiple NUMA nodes. A transit cache pool is configured in the NUMA architecture, and the transit cache pool includes caches allocated from the memories of different NUMA nodes. The processor is configured to execute the above data interaction processing method.

[0041] A third aspect of the present application proposes a NUMA system architecture, and the NUMA system architecture is configured to implement the above data interaction processing method.

[0042] In the data interaction method proposed by the present application, a transit cache pool composed of the memories of different NUMA nodes is configured in the NUMA system architecture. When the first NUMA node determines that the target storage device corresponding to the data interaction request is located in the second NUMA node, the first NUMA node allocates a target cache from the transit cache pool. The target cache belongs to the same memory as the memory of the second NUMA node, and then data interaction is performed with the target storage device through the target cache. The above solution enables the NUMA node program to convert the cross-NUMA node access to the target storage device into interacting with the target storage device through the cache of the node where the target storage device is located when accessing the target storage device across NUMA nodes. In this way, the cache of the NUMA node where the target storage device is located can be used to reduce the latency of interacting with the target storage device and improve the interaction efficiency. Description of the Drawings

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0044] Figure 1A schematic structural diagram of a NUMA system architecture provided by an embodiment of the present application.

[0045] Figure 2 Another schematic structural diagram of a NUMA system architecture provided by an embodiment of the present application.

[0046] Figure 3 A schematic flowchart of a data interaction processing method provided by an embodiment of the present application.

[0047] Figure 4 A schematic diagram of the processing process for allocating a target cache from a transfer cache pool provided by an embodiment of the present application. Detailed implementation manners

[0048] The technical solution of the embodiment of the present application is applicable to the application scenario where a user program in a NUMA system architecture accesses a storage device across NUMA nodes. By adopting the technical solution of the embodiment of the present application, the access latency of the storage device in the above scenario can be reduced, and the data interaction efficiency between the user program and the storage device can be improved.

[0049] In a NUMA system architecture, processors or processor cores are divided into multiple NUMA nodes. Each node has its own independent memory space and bus system, and the various NUMA nodes are also interconnected through a bus.

[0050] In one implementation manner, a NUMA system architecture includes a processor, and the processor includes multiple processor cores. The multiple processor cores are divided into different groups, and each group includes at least one processor core. Each group serves as a NUMA node.

[0051] In another implementation manner, as Figure 1 shown, a NUMA system architecture includes multiple processors, and each processor includes multiple processor cores. The processor cores of each processor are divided into different NUMA nodes.

[0052] Regardless of the adopted manner, for each NUMA node, an independent memory controller is also configured. The processor cores in the same NUMA node can share the memory controller. At the same time, corresponding physical memory is configured for each NUMA node. Each NUMA node can access the memory space of its own configured local memory through its own corresponding memory controller, and at the same time, can also access the memory space of other NUMA nodes.

[0053] In a NUMA system architecture, each NUMA node corresponds to its own memory controller and memory. Any NUMA node can access its own corresponding memory or the memory of any other node in the architecture. According to the relationship between NUMA nodes, NUMA nodes can be divided into local nodes, neighbor nodes, and remote nodes. The speed at which a processor core accesses the memory of different types of nodes is different. The speed of accessing the local node is the fastest, and the speed of accessing the remote node is the slowest.

[0054] In addition, in each NUMA node of the NUMA system architecture, hardware storage devices such as non-volatile storage media like SSDs can be mounted to provide a more reliable data storage function for the NUMA system architecture. These storage devices can adopt any storage protocol to achieve higher performance and higher concurrency storage functions.

[0055] For example, NVME (Non-Volatile Memory Express) is a high-performance, high-concurrency storage protocol for non-volatile storage media (such as SSDs). NVMe uses the PCIe interface to directly connect to the storage device, which significantly improves the data transfer rate and I / O performance compared with traditional SATA and SAS interfaces. The NVMe protocol is designed simply and efficiently, supports a multi-queue architecture, greatly reduces latency, and enhances parallel processing capabilities. The low latency and high bandwidth characteristics of NVMe make it particularly suitable for high-performance computing, big data analysis, and enterprise-level storage solutions that require fast storage access.

[0056] Mounting an NVMe storage device on the NUMA node of the NUMA system architecture can form a NUMA system architecture with multiple NVMe storage devices as shown in Figure 2 wherein, on NUMA0, NUMA1, and NUMA2, NVMe storage devices are respectively connected through the PCIe interface.

[0057] SPDK (Storage Performance Development Kit) is a storage performance development toolkit program for high-performance storage applications. It aims to maximize storage performance by directly accessing and controlling storage hardware (such as NVMe SSDs) while minimizing CPU occupancy. SPDK can significantly improve the performance of the storage system and is particularly suitable for application scenarios with extremely high requirements for storage latency and throughput.

[0058] In Figure 2 the shown NUMA system architecture, applying SPDK can significantly improve the data interaction efficiency between the user program and the storage device under the NUMA node.

[0059] However, for Figure 2In the NUMA system architecture shown, the speed at which a user program running on a local NUMA node accesses the NVMe storage device mounted on the local NUMA node is higher than that of accessing the NVMe storage device mounted on other NUMA nodes. In the current NUMA system architecture, it often happens that a user program running on a certain NUMA node accesses the storage device of other NUMA nodes. This cross-node access to the storage device will bring a large interaction delay and affect the data reading and writing efficiency.

[0060] For example, the SPDK program running on the first NUMA node usually accesses the NVMe storage device under other NUMA nodes. This data interaction process requires a large time delay and it is difficult to exert the performance advantages of SPDK.

[0061] To address the above technical problems, a common solution at present is to specify the processor for the user program to run, that is, to control the user program to run on the processor or processor core of the NUMA node where the storage device it needs to access is located, so as to avoid the user program from accessing the storage device across NUMA nodes. However, this method requires the user to forcibly select a specific processor or processor core to run when running the program, which is sometimes unacceptable to the user. And the above method may also lead to unbalanced deployment of user programs in the NUMA system architecture, affecting the system performance.

[0062] To address the above technical problems, this application proposes a data interaction processing method applied to the NUMA system architecture, which can reduce the latency when a user program in the NUMA system architecture accesses a storage device across NUMA nodes and improve the efficiency of data interaction between the user program and the storage device across NUMA nodes.

[0063] Next, the technical solutions in the embodiments of this application will be described clearly and completely with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0064] The embodiment of this application proposes a data interaction method. This method is applied to the first NUMA node in the NUMA system architecture and can be specifically applied to the user program running on the first NUMA node. The above user program can be any application program, such as the SPDK program, etc.

[0065] The above NUMA system architecture includes multiple NUMA nodes, and at least one storage device is mounted in at least one NUMA node. This storage device can be a storage device of any type and implementing any storage protocol.

[0066] In this NUMA system architecture, a transfer cache pool is configured. The caches in this transfer cache pool are obtained by different NUMA nodes in the NUMA system architecture partitioning a part of their local memory into this memory transfer pool. That is, different NUMA nodes in the above NUMA system architecture partition different sizes of memory from their respective memories as caches, and these caches are aggregated into the above transfer cache pool. Any NUMA node in the NUMA system architecture can access and use the caches in the transfer cache pool, such as performing data interaction with the caches in the transfer cache pool, etc.

[0067] In some embodiments, the user program running on the above first NUMA node includes an application layer and a protocol layer. The application layer is used to interact with the user, perform basic initialization processing based on user operations, etc. The protocol layer then performs corresponding processing based on the tasks issued by the application layer or executes through the protocol process.

[0068] For example, SPDK can be divided into an SPDK application layer and an SPDK protocol layer.

[0069] Among them, the SPDK application layer is coded by the user layer, and its specific behaviors include: parameter parsing, device scanning, binding cores to create threads, queue initialization, and DMA memory initialization.

[0070] The interface of the SPDK application layer will finally call the library functions of the SPDK protocol layer to complete the processing flow of the storage protocol. For example, for an nvme storage device, it will complete the processing flow of the nvme protocol. For example: NVME requests are created, encapsulated and placed in the request queue, and after the hardware processing is completed, the processed requests are put back into the completion queue for user callback.

[0071] In the embodiments of the present application, the data interaction method proposed in the present application is executed through the protocol layer of the program. For example, the data interaction method is executed through the protocol layer of SPDK. By optimizing the protocol process of the protocol layer, the interaction efficiency when the user program interacts with the storage device across NUMA nodes is improved. In this way, the interaction efficiency between the user program and the storage device across NUMA nodes can be improved without the user's perception.

[0072] It should be noted that SPDK has a total of three NVME protocols: TCP, RDMA, and PCIE. Both the TCP and RDMA protocols are remote protocols. The protocol initiator sends requests to the protocol executor, and finally the protocol executed by the executor is still PCIE. Therefore, in the embodiments of the present application, only the optimization of the PCIE protocol of the SPDK protocol layer is concerned.

[0073] Based on the above NUMA system architecture, see Figure 3As shown in the figure, the data interaction processing method proposed in the embodiment of the present application includes the following processing steps executed by the user program of the first NUMA node:

[0074] S101. Determine the NUMA node where the target storage device corresponding to the data interaction request is located.

[0075] Among them, the data interaction request is generated by a program running on the first NUMA node. For example, when the SPDK program running on the first NUMA node reads data from the target storage device in response to a user operation, a data interaction request for the target storage device will be generated. The data interaction request can be to read data from the target storage device, or to write data to or send data to the target storage device.

[0076] The above-mentioned target storage device can be a storage device using the NVMe protocol.

[0077] The program running on the first NUMA node can parse and determine the target storage device corresponding to the destination address according to the destination address of the data interaction request, and determine the NUMA node where the target storage device is located according to the NUMA system architecture.

[0078] If the program running on the first NUMA node determines that the NUMA node where the target storage device is located is the first NUMA node, that is, the program of the first NUMA node issues a data interaction request for the storage device of this NUMA node. At this time, according to the conventional data interaction logic, data read and write and other interaction operations can be directly performed with the storage device of the first NUMA node.

[0079] S102. When it is determined that the NUMA node where the target storage device is located is the second NUMA node, allocate a target cache from the transfer cache pool.

[0080] When the program running on the first NUMA node determines that the NUMA node where the target storage device is located is the second NUMA node, if the program of the first NUMA node directly interacts with the target storage device, a situation of accessing the storage device across NUMA nodes will occur, and this situation will cause a relatively high data interaction delay.

[0081] In view of the above situation, in this embodiment, if the program of the first NUMA node determines through judgment that the target storage device is not local to the first NUMA node but in the second NUMA node, it does not directly perform data interaction with the target storage device of the second NUMA node, but first accesses the transfer cache pool and allocates a target cache from the transfer cache pool for performing data interaction operations with the target storage device.

[0082] Among them, the above-mentioned target cache and the memory of the second NUMA node belong to the same memory. That is, when the program of the first NUMA node selects a cache from the transfer cache pool, it specifically selects from the free caches from the memory of the second NUMA node.

[0083] Exemplarily, in the transfer cache pool, the attribute information of each cache carries information such as the NUMA node to which the cache belongs, the free status of the cache, and the size of the cache. When the program of the first NUMA node accesses the transfer cache pool, by parsing the attribute information of each cache in the transfer cache pool, it identifies a free cache that belongs to the same memory as the memory of the second NUMA node and has a size sufficient to execute the current data interaction request operation as the target cache.

[0084] Among them, the size of the cache required to execute the current data interaction request operation can be determined according to the size of the data to be read and written by the current data interaction request. For example, the size of the cache required to execute the current data interaction request operation is not less than the size of the data to be read and written by the current data interaction request.

[0085] S103. Perform data interaction with the target storage device through the target cache.

[0086] Specifically, after the program of the first NUMA node determines the target cache from the transfer cache pool, it performs data interaction corresponding to the above data interaction request with the target storage device through the target cache.

[0087] In some embodiments, the program of the first NUMA node performs data interaction with the target storage device through the above-mentioned target cache. Specifically, the target cache is used as a data transfer cache, and data transfer is performed through the target cache to perform data interaction with the target storage device.

[0088] For example, data is sent to the target storage device through data transfer via the target cache, or data is read from the target storage device through data transfer via the target storage device.

[0089] It can be understood that using the target cache as a data transfer cache to interact with the target storage device is actually a program on the first NUMA node interacting with the target cache, and then the target cache directly interacts with the target storage device on the second NUMA node. Since the latency of a program on the first NUMA node accessing the cache on the second NUMA node across nodes is much lower than that of a program on the first NUMA node directly accessing the storage device on the second NUMA node across nodes, therefore, compared with a program on the first NUMA node directly interacting with the target storage device on the second NUMA node, the embodiments of the present application enable the program on the first NUMA node to interact with the second NUMA node through the target cache transfer with higher efficiency and lower latency.

[0090] As can be seen from the above introduction, the data interaction method proposed in the embodiments of the present application configures a transfer cache pool composed of the memories of different NUMA nodes in the NUMA system architecture. When the first NUMA node determines that the target storage device corresponding to the data interaction request is located on the second NUMA node, the first NUMA node allocates a target cache from the transfer cache pool. The target cache and the memory of the second NUMA node belong to the same memory, and then data interaction is performed with the target storage device through the target cache. The above solution enables the NUMA node program to convert the cross-NUMA node access to the target storage device into interacting with the target storage device through the cache of the node where the target storage device is located when accessing the target storage device across NUMA nodes, so that the delay of interacting with the target storage device can be reduced by leveraging the cache of the NUMA node where the target storage device is located, and the interaction efficiency can be improved.

[0091] In some embodiments, the program on the first NUMA node interacts with the target storage device through the target cache, including:

[0092] The program on the first NUMA node modifies the cache pointer of the data interaction request to point to the target cache, so that when the protocol controller of the target storage device receives the data interaction request, it performs the data interaction operation between the target cache and the target storage device according to the data interaction request.

[0093] Specifically, taking the SPDK program as an example, when a user reads and writes data of a storage device through the SPDK program, a Buffer cache is passed in each time a data interaction request is issued to store the read and written data. Normally, the data that the user wants to write to the storage device is stored in this Buffer cache, and then the data is written from this Buffer cache to the storage device. Or, when the user wants to read data from the storage device, the data is first read from the storage device into this Buffer cache.

[0094] For the original data interaction request, the cache pointer therein points to the Buffer cache where the request originally comes in. When the protocol controller of the storage device executes the operation of this data interaction request, it will perform the data read and write operations of the storage device according to the Buffer cache passed in by this request.

[0095] In this embodiment, after the program of the first numa node determines the above target cache, it modifies the cache pointer of the data interaction request to point to this target cache, that is, uses the target cache to replace the original Buffer cache of the data interaction request. At this time, when the protocol controller of the target storage device receives this data interaction request, since the cache pointer of this data interaction request points to the target cache, the protocol controller will use the target cache pointed to by the cache pointer as the cache for data interaction operations with the target storage device. Therefore, the protocol controller of the target storage device performs the data interaction operation between the target cache and the target storage device according to this data interaction request, that is, according to the specific request of this data interaction request, conducts data interaction through the target cache and the target storage device. For example, the data in the target cache is stored in the target storage device, or the data in the target storage device is read into the target cache.

[0096] In some embodiments, when the above data interaction request is a data sending request, after the program of the first numa node allocates the target cache from the transfer cache pool, it first stores the data to be sent in this target cache.

[0097] Based on the above operations, when the program of the first numa node modifies the cache pointer of the data interaction request to point to this target cache, the protocol controller of the target storage device stores the data stored in the target cache to the target storage device according to this data interaction request.

[0098] For example, when the data interaction request specifies the storage location of the data in the target storage device, the protocol controller of the target storage device stores the data stored in the target cache at the above storage location of the target storage device.

[0099] In other embodiments, when the above data interaction request is a data reading request, when the program of the first numa node modifies the cache pointer of the data interaction request to point to the above target cache, the protocol controller of the target storage device reads the data requested to be read by this data interaction request from the target storage device according to this data interaction request, and stores the read data in this target cache.

[0100] In some other embodiments, when the program of the first NUMA node determines that the NUMA node where the target storage device corresponding to the data interaction request is located is the second NUMA node, it first determines the original cache for executing the data interaction operation corresponding to the data interaction request, that is, determines the Buffer cache into which the data interaction request is incoming, and saves the address of the original cache.

[0101] Since the data interaction request is a data interaction request initiated by the program of the first NUMA node, this original cache and the memory of the first NUMA node belong to the same memory.

[0102] On this basis, in the case where the above data interaction request is a data sending request, during the data interaction process, when the program of the first NUMA node writes the data to be sent into the target cache, specifically, it copies the data to be sent stored in the above original cache to the above target cache.

[0103] In the case where the above data interaction request is a data reading request, during the data interaction process, when the protocol controller of the target storage device reads the data requested to be read by the data interaction request from the target storage device and stores the read data into the target cache, the program of the first NUMA node copies the data stored in the target cache to the above original cache, and then, when the program of the first NUMA node executes the subsequent data interaction protocol process, it can feed back the data in the original cache to the user layer.

[0104] In practical applications, the above scheme of sending data to the target storage device through the target cache can be applied only when the program of the first NUMA node sends data to the storage device of the second NUMA node. For example, when the program of the first NUMA node sends data to the target storage device of the second NUMA node, the above scheme of sending data to the target storage device through the target cache is applied, while when the program of the first NUMA node reads data from the target storage device of the second NUMA node, the cross-node direct reading method is still adopted.

[0105] Alternatively, the above scheme of reading data from the target storage device through the target cache can be applied only when the program of the first NUMA node reads data from the storage device of the second NUMA node. For example, when the program of the first NUMA node sends data to the target storage device of the second NUMA node, the program of the first NUMA node directly sends data across nodes to the target storage device of the second NUMA node, while when the program of the first NUMA node reads data from the target storage device of the second NUMA node, the scheme of reading data from the target storage device through the target cache is adopted.

[0106] Alternatively, when the program on the first NUMA node sends data to the target storage device on the second NUMA node and the program on the first NUMA node reads data from the target storage device on the second NUMA node, the above-described solution for data interaction through the target cache and the target storage device is applied. That is, when the program on the first NUMA node sends data to the target storage device on the second NUMA node, as introduced in the above embodiments, it is sent through the target cache as a data transfer cache. When the program on the first NUMA node reads data from the target storage device on the second NUMA node, as introduced in the above embodiments, it is read through the target cache as a data transfer cache.

[0107] In some other embodiments, in order to detect the status of data interaction requests in real time to facilitate the program on the first NUMA node to perform cached data copying, a request queue and a completion queue are set up in the data interaction protocol process to implement the status monitoring of data interaction requests.

[0108] Among them, the request queue is used to store uncompleted data interaction requests. When the user program issues a data interaction request, the data interaction request will be first stored in the request queue.

[0109] The completion queue is used to store completed data interaction requests. When the protocol controller of the storage device finishes performing the data interaction operation corresponding to the data interaction request, the data interaction request will be transferred from the request queue to the completion queue.

[0110] Based on the above settings of the request queue and the completion queue, the program on the first NUMA node can determine whether the data interaction request is completed by detecting the data interaction request from the completion queue.

[0111] For example, in the case where the data interaction request is a data read request, after the program on the first NUMA node modifies the cache pointer of the data interaction request and issues the modified data interaction request to the protocol controller of the target storage device, the program on the first NUMA node periodically detects the data interaction request from the completion queue.

[0112] In the case where the data interaction request is detected from the completion queue, the program on the first NUMA node can determine that the data interaction operation corresponding to the data interaction request has been completed. At this time, the program on the first NUMA node can copy the data stored in the target cache to the original cache.

[0113] The above request queue and completion queue can be set to run in the memory of the first NUMA node or can be set to run in the memory of the second NUMA node. The embodiments of the present application do not make any limitations.

[0114] Next, taking the example that the SPDK program on the first NUMA node reads data from the NVMe storage device on the second NUMA node, the complete processing process of the data interaction method proposed in the embodiments of the present application will be introduced by way of example:

[0115] First, when the upper-layer user program reads data from the NVMe storage device on the second NUMA node through the SPDK program, the SPDK program first initializes based on the user's data read request. During this initialization process, a request queue and a completion queue will be created in the memory of the first NUMA node or in the memory of the second NUMA node, and the data read request will be stored in the request queue. Preferably, the request queue and the completion queue are created in the memory of the second NUMA node to more efficiently sense the completion of the request.

[0116] Then, when the SPDK program determines that the NVMe device to be accessed is the NVMe storage device on the second NUMA node, it determines the address and size of the original cache of the data read request, and then selects, from the transfer cache pool, an idle cache that belongs to the same memory as the memory of the second NUMA node and is not less than the above original cache as the target cache. SPDK modifies the cache pointer of the data read request to point to the target cache, and at the same time saves the above original cache.

[0117] The SPDK program continues the submission process of the read request operation using the above target cache according to the NVMe protocol. Finally, the target cache will be converted into a physical address at the NVMe underlying protocol layer and written into the register corresponding to the NVMe protocol.

[0118] Next, the NVMe controller performs DMA data transfer according to the target cache address that has been written into the register in the request process, that is, reads the data required to be read by the data read request from the NVMe storage device on the second NUMA node and stores it in the target cache. After the relevant data transfer is completed, the data read request will be transferred from the request queue to the completion queue.

[0119] The SPDK program polls and detects the completion queue to detect whether the data read request exists in it. When the data read request is detected from the completion queue, it can be determined that the data has been read from the target NVMe storage device to the target cache. At this time, the SPDK program copies the data in the target cache to the original cache. The upper-layer user program can access the original cache to obtain the data read from the NVMe storage device on the second NUMA node.

[0120] In some embodiments, when the program on the first NUMA node allocates the target cache from the transfer cache pool, the allocation is performed according to the processing steps A1 to A4 as shown in Figure 4 :

[0121] A1. Query candidate caches from the transfer cache pool.

[0122] Among them, the candidate cache is an idle cache that belongs to the same memory as the memory of the second NUMA node and can be used to execute the data interaction request.

[0123] Specifically, when a program of the first NUMA node selects a target cache from the transfer cache pool, it should ensure that the selected target cache meets the following conditions: 1. Belongs to the same memory as the memory of the second NUMA node; 2. The cache size can be used for this data interaction request, that is, it is sufficient to store the interaction data during the execution of this data interaction request; 3. The cache is idle.

[0124] Based on the above conditions, the program of the first NUMA node first queries the caches in the transfer cache pool that meet the above conditions simultaneously as candidate caches.

[0125] In some embodiments, the caches in the transfer cache pool can be divided into an available queue and an unavailable queue. Among them, the available queue is used to store idle caches, and the unavailable queue is used to store caches that are being used. The transfer cache pool maintains the storage of each cache in the available queue and the unavailable queue. When a certain cache is idle, the transfer cache pool stores it in the available queue. When a cache in the available queue is used, this cache is transferred from the available queue to the unavailable queue. At the same time, when a cache in the unavailable queue is released after being used up, this cache is transferred from the unavailable queue to the available queue.

[0126] Based on the above settings of the available queue and the unavailable queue, when the program of the first NUMA node queries candidate caches from the transfer cache pool, it can directly query from the available queue to improve the query efficiency.

[0127] In the case where the candidate cache is queried from the transfer cache pool, execute step A2. Allocate a target cache from the candidate cache.

[0128] Specifically, if the program of the first NUMA node queries a candidate cache that meets the above conditions simultaneously from the transfer cache pool, the program of the first NUMA node determines any one of the queried candidate caches as the target cache.

[0129] Or,

[0130] In the case where the candidate cache is not queried from the transfer cache pool, execute step A3. Send a creation instruction to the second NUMA node, and the creation instruction is used to instruct the second NUMA node to allocate a specific size of memory of the second NUMA node to the transfer cache pool as a newly created cache; the specific size of memory is not less than the candidate cache.

[0131] Specifically, if the program of the first NUMA node fails to find a candidate cache in the transfer cache pool, it indicates that there is currently no cache in the transfer cache pool that can execute the above data interaction request. At this time, the program of the first NUMA node sends a creation instruction to the second NUMA node through the first NUMA node. For example, the program sends a creation instruction to the processor or processor core of the second NUMA node through the processor or processor core of the first NUMA node, instructing the second NUMA node to allocate a memory of a specific size in the second NUMA node to the transfer cache pool as a newly created cache.

[0132] Among them, the above-mentioned memory of a specific size is not less than the above-mentioned candidate cache, that is, this memory of a specific size should be able to be used for this data interaction request, that is, it is sufficient to store the interaction data during the execution of this data interaction request.

[0133] The above creation instruction can instruct the second NUMA node to allocate a new cache to the transfer cache pool for executing the current data interaction request. This way of dynamically allocating caches enables the caches in the transfer cache pool to meet the cross-node data interaction requirements in real time.

[0134] Since during the operation of the NUMA system architecture, it is impossible to predict what data interaction operations the user will perform, nor can it predict the size of the interaction data required by the user's data interaction request and the data interaction frequency. Therefore, it is impossible to ensure that the caches in the transfer cache pool can meet the interaction requirements in real time. In this case, the embodiment of the present application can, through the above-mentioned way of dynamically allocating caches, timely instruct the target NUMA node to allocate a new cache that can meet the interaction requirements to the transfer cache pool when there is no available cache, so that the transfer cache pool can meet the data interaction requirements in real time.

[0135] A4. Allocate a target cache from the newly created cache.

[0136] Specifically, when there is only one newly created cache, this newly created cache can be directly used as the target cache; when there are multiple newly created caches, one of them can be selected as the target cache.

[0137] In some embodiments, when a program in the first NUMA node sends a creation instruction to the second NUMA node, instructing the second NUMA node to allocate a memory of a specific size to the transfer cache pool as a newly created cache, and using the newly created cache as the target cache to complete the above data interaction request, the program in the first NUMA node sends a recycling instruction to the second NUMA node, instructing the second NUMA node to recycle the memory of the above specific size. This can, when the data interaction request is completed, enable the second NUMA node to promptly recycle the temporarily allocated memory from the transfer cache pool, avoid occupying the memory of the second NUMA node for a long time and affecting the performance of the second NUMA node, and maintain the performance stability of the entire NUMA system architecture.

[0138] In other embodiments, the caches in the transfer cache pool are divided into cache queues of different sizes. Among them, each cache queue is used to store caches of a certain size, and the cache sizes of different cache queues are different. For example, a 2K cache queue and a 2M cache queue are set in the transfer cache pool. Multiple 2K caches are stored in the 2K cache queue, and multiple 2M caches are stored in the 2M cache queue.

[0139] Since the sizes of the data to be interacted in different data interaction requests are different, the cache sizes required for different data interaction requests are different. Moreover, in the data interaction protocol of the NUMA system architecture, the interaction requests for larger data are usually split into multiple interaction requests for smaller data. For example, in SPDK, the data interaction requests exceeding 2M are segmented, and a single request is split into multiple sub-requests, and the corresponding data size of each sub-request does not exceed 2M. Therefore, in the NUMA system architecture, the cache sizes required for data interaction requests are usually several fixed cache sizes.

[0140] Based on the above characteristics, the embodiments of the present application set multiple cache queues in the transfer cache pool, which are respectively used to store caches of different sizes, so as to meet the cache requirements of different data interaction requests.

[0141] Based on the above multiple cache queues, when a program in the first NUMA node queries for candidate caches from the transfer cache pool, it can first screen out the target cache queue from the transfer cache pool. The caches in the target cache queue are the smallest caches that can meet the cache required for the data interaction demand. For example, assuming that the cache required for the data interaction request does not exceed 2K, the 2K cache queue is determined as the target cache queue; assuming that the cache required for the data interaction request is greater than 2K but does not exceed 2M, the 2M cache queue is determined as the target cache queue.

[0142] Then, query for candidate caches from the target cache queue.

[0143] In some embodiments, an available queue and an unavailable queue introduced in the above embodiments are provided in the above cache queue, and the working modes of the available queue and the unavailable queue are the same as those introduced in the above embodiments. When a program of a first numa node determines a target cache queue, when querying candidate caches from the target cache queue, it directly queries in the available queue of the target cache queue, thereby improving the query efficiency.

[0144] The above cache selection method can select the smallest cache that can meet the data interaction requirements from the transfer cache pool, thereby avoiding cache waste, improving the cache read / write performance, and improving the data interaction efficiency.

[0145] In some other embodiments, in the case where multiple cache queues of different sizes are provided in the transfer cache pool, in individual cases, there may also be a situation where the cache required for a data interaction request in the numa system architecture is relatively large, and there is no target cache in the transfer cache pool that can meet the cache required for the data interaction request. At this time, it is possible to allocate a cache that meets the requirements of the current data interaction request to the transfer cache pool in real time according to the processing of step A3 in the above embodiments, and, when the current data interaction request is completed, release the cache back to its numa node where it is located. This solution can avoid setting up special cache queues for individual special cases, but instead adopts a dynamic allocation and recycling method, which can improve the utilization efficiency of the numa node memory and maintain the performance stability of the numa system architecture. Moreover, the caches in each cache queue in the above transfer cache pool are recycled, improving the system performance.

[0146] The data interaction processing method proposed in the above embodiments of the present application converts the process of a numa node program accessing a storage device across nodes into a process of a numa node program accessing memory across nodes, that is, through the transfer memory to perform interactions with the local node of the target storage device, and then the numa node program interacts with the transfer memory. The data interaction processing method proposed in the embodiments of the present application has higher data interaction efficiency and can transfer a larger amount of data per unit time compared with the numa node program directly accessing the storage device across nodes.

[0147] The inventor of the present application conducted a data interaction efficiency test comparison between the data interaction processing method proposed in the above embodiments of the present application and the process of a conventional numa node program accessing a storage device across nodes, and the comparison results are shown in Table 1.

[0148] Table 1

[0149]

[0150] Table 1 above shows the method of accessing a storage device across NUMA nodes, as well as the amount of data that can be transmitted per unit time during the interaction of different lengths of data between a NUMA node program and a cross-node storage device through the data interaction method proposed in this application.

[0151] As can be seen from Table 1, in the scenario of transmitting 1K of data, since the data length is relatively small, even when accessing the storage device across NUMA nodes, the bottleneck of accessing the storage device across NUMA nodes is not reached. Therefore, the data interaction efficiency improvement of the data interaction method proposed in the embodiments of this application compared to the data interaction of accessing the storage device across NUMA nodes is not obvious.

[0152] As the length of the transmitted data increases, compared with the method of accessing the storage device across NUMA nodes, the amount of data transmitted per unit time by the data interaction method proposed in the embodiments of this application increases significantly, and its data interaction efficiency is significantly improved.

[0153] It can be proved by the above experiments that the data interaction method proposed in the embodiments of this application has an obvious effect on improving the data interaction efficiency of a NUMA node program accessing a storage device across nodes, and can improve the data interaction performance of the NUMA system architecture.

[0154] Corresponding to the above data interaction processing method, an embodiment of this application also proposes a processor, which can be applied to the first NUMA node in a NUMA system architecture, and a program can run on this processor, such as running an operating system program, an application program, etc.

[0155] The above NUMA system architecture includes multiple NUMA nodes, and a transfer cache pool is configured in the NUMA architecture. The transfer cache pool includes caches allocated from the memories of different NUMA nodes. For the specific structure introduction of this NUMA system architecture, reference can be made to the introduction of the NUMA system architecture described in the above embodiments.

[0156] The above processor is configured to execute the data interaction processing method described in any of the above embodiments. For the specific processing process when the processor executes the above data interaction processing method, reference can be made to the corresponding processing process in the data interaction processing method described in the above embodiments, and details will not be repeated here.

[0157] In an embodiment of the present application, a processor is a circuit with the ability to process signals. In one implementation, the processor can be a circuit with the ability to read and execute instructions, such as a CPU, microprocessor, GPU, or DSP, etc.; in another implementation, the processor can implement certain functions through the logical relationship of a hardware circuit, and the logical relationship of the hardware circuit is fixed or can be reconfigured. For example, the processor is a hardware circuit implemented by an ASIC or PLD, such as an FPGA, etc. In a reconfigurable hardware circuit, the process of the processor loading a configuration document to implement the configuration of the hardware circuit can be understood as the process of the processor loading instructions to implement the data interaction processing method described in the above embodiments. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as a type of ASIC, such as an NPU, TPU, DPU, etc.

[0158] In another embodiment, a NUMA system architecture is further provided. The NUMA system architecture is configured to implement the data interaction processing method described in any of the above embodiments. For example, the NUMA system architecture includes multiple NUMA nodes, and a transit cache pool is configured in the NUMA architecture. The transit cache pool includes caches allocated from the memories of different NUMA nodes. The first NUMA node in this NUMA system architecture executes the processing process of the data interaction processing method described in any of the above embodiments.

[0159] The processor and the NUMA system architecture provided in this embodiment belong to the same inventive concept as the data interaction processing method provided in the above embodiments of the present application, can execute the data interaction processing method provided in any of the above embodiments of the present application, and have the corresponding functional modules and beneficial effects for executing this method. For technical details not described in detail in this embodiment, reference can be made to the specific processing content of the data interaction processing method provided in the above embodiments of the present application, and details will not be elaborated here.

[0160] Other embodiments of the present application also propose a computer device, which includes the above-mentioned processor or includes the above-mentioned NUMA system architecture. The computer device can specifically be a computer, server, intelligent terminal, handheld terminal, wearable device, etc.

[0161] In addition to the above methods and devices, an embodiment of the present application can also be a computer program product, which includes computer program instructions. When the computer program instructions are run by a processor, the processor is caused to execute the steps in the data interaction processing method described in any of the above embodiments of this specification.

[0162] The computer program product can be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present application. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The programming code can be executed entirely on the user computing device, partially on the user device, executed as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0163] In addition, an embodiment of the present application can also be a storage medium on which a computer program is stored, and the computer program is executed by a processor to perform the steps in the data interaction processing method described in any of the above embodiments of this specification.

[0164] For the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, some steps can be in other sequences or performed simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0165] It should be noted that the embodiments in this specification are all described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0166] The steps in the methods of the embodiments of the present application can be adjusted, combined, and deleted according to actual needs. The technical features recorded in each embodiment can be replaced or combined.

[0167] The modules and sub-modules in the devices and terminals in the embodiments of the present application can be combined, divided, and deleted according to actual needs.

[0168] In several embodiments provided by the present application, it should be understood that the disclosed terminals, devices, and methods can be implemented in other ways. For example, the terminal embodiments described above are merely illustrative. For example, the division of modules or sub-modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple sub-modules or modules can be combined or integrated into another module, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be in electrical, mechanical, or other forms.

[0169] The modules or sub-modules described as separate components may or may not be physically separated. The components as modules or sub-modules may or may not be physical modules or sub-modules, that is, they can be located in one place, or can be distributed to multiple network modules or sub-modules. Some or all of the modules or sub-modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0170] In addition, each functional module or sub-module in various embodiments of the present application can be integrated in a processing module, or each module or sub-module can exist physically alone, or two or more modules or sub-modules can be integrated in one module. The above-mentioned integrated modules or sub-modules can be implemented in the form of hardware or in the form of software functional modules or sub-modules.

[0171] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0172] The steps of the method or algorithm described in combination with the embodiments disclosed in this article can be directly implemented by hardware, software units executed by a processor, or a combination of both. The software units can be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0173] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.

[0174] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A data interaction processing method, characterized in that: A first NUMA node applied to a NUMA system architecture, the NUMA system architecture comprising a plurality of NUMA nodes, a transit cache pool configured in the NUMA architecture, the transit cache pool comprising cache allocated from memory of different NUMA nodes, the method comprising: Determine the numa node where the target storage device corresponding to the data interaction request is located; the data interaction request is generated by a program running on the first numa node; In the case where it is determined that the numa node where the target storage device is located is a second numa node, a target cache is allocated from the transit cache pool, and the target cache and the memory of the second numa node belong to the same memory; Data is exchanged with the target storage device through the target cache.

2. The method according to claim 1, characterized in that The method of performing data interaction with the target storage device through the target cache includes: The cache pointer of the data interaction request is modified to point to the target cache, so that when the protocol controller of the target storage device receives the data interaction request, it performs the data interaction operation between the target cache and the target storage device according to the data interaction request.

3. The method according to claim 2, characterized in that In the case where the data interaction request is a data sending request, the method further comprises: storing the data to be sent in the target cache; The protocol controller performs a data interaction operation between the target cache and the target storage device according to the data interaction request, including: The protocol controller stores the data stored in the target cache to the target storage device according to the data exchange request; or, In the case where the data interaction request is a data read request, the protocol controller performs a data interaction operation between the target cache and the target storage device according to the data interaction request, including: The protocol controller reads the data requested by the data interaction request from the target storage device, and stores the read data into the target cache.

4. The method according to claim 3, characterized in that In the case where it is determined that the numa node where the target storage device is located is a second numa node, the method further includes: determining an original cache for executing the data interaction operation corresponding to the data interaction request, the original cache and the memory of the first numa node belong to the same memory; In the case where the data interaction request is a data sending request, storing the data to be sent in the target cache includes: copying the data stored in the original cache to the target cache; or, In the case where the data interaction request is a data read request, the method further includes: copying the data stored in the target cache to the original cache.

5. The method according to claim 4, characterized in that In the case where the data interaction request is a data read request, before copying the data stored in the target cache to the original cache, the method further includes: Detecting the data interaction request from a preset completion queue; the completion queue is used to store completed data interaction requests; When the data exchange request is detected from the completion queue, the data stored in the target cache is copied to the original cache.

6. The method according to claim 5, characterized in that In the case where it is determined that the numa node where the target storage device is located is a second numa node, the method further includes: In the memory of the second numa node, a request queue and a completion queue are configured, and the request queue is used to store unfinished data interaction requests.

7. The method according to any one of claims 1 to 6, characterized in that: Allocating a target cache from the transit cache pool includes: Querying a candidate cache from the transit cache pool, where the candidate cache is an idle cache that belongs to the same memory as the memory of the second NUMA node and can be used to execute the data interaction request; When the candidate cache is found from the transit cache pool, a target cache is allocated from the candidate cache; or, In the case where the candidate cache is not queried from the transit cache pool, a creation instruction is sent to the second NUMA node, wherein the creation instruction is used to instruct the second NUMA node to allocate a specific size memory of the second NUMA node to the transit cache pool as a newly created cache; the specific size memory is not less than the candidate cache; Assign a target cache from the newly created cache.

8. The method according to claim 7, characterized in that The transfer buffer pool includes a plurality of buffer queues, each buffer queue includes a plurality of buffers of the same size, and the buffer sizes of different buffer queues are different; Querying a candidate cache from the transit cache pool includes: A candidate cache is queried from a target cache queue of the transit cache pool, wherein the cache in the target cache queue is a minimum cache that can satisfy the cache required by the data interaction request.

9. The method according to claim 8, characterized in that In the case where the second NUMA node allocates a specific size memory of the second NUMA node to the transit cache pool as a newly created cache, the method further includes: When the data interaction request is completed, a reclaim instruction is sent to the second NUMA node, where the reclaim instruction is used to instruct the second NUMA node to reclaim the memory of the specific size.

10. The method according to any one of claims 1 to 6, characterized in that: The data interaction request includes a data interaction request generated by a storage performance development toolkit program running on the first numa node; and / or, The target storage device includes a storage device using the NVME protocol mounted on the second NUMA node.

11. A processor, characterized in that: A first NUMA node applied to a NUMA system architecture, the NUMA architecture comprising a plurality of NUMA nodes, a transit cache pool configured in the NUMA architecture, the transit cache pool comprising caches allocated from memories of different NUMA nodes, the processor being configured to execute the data interaction processing method as described in any one of claims 1 to 10.

12. A numa system architecture, characterized in that: The NUMA system architecture is configured to implement the data interaction processing method as described in any one of claims 1 to 10.