Data processing system and method, device and nonvolatile readable storage medium
By binding exclusive physical network cards, hard disks and memory in the disk array, and using RDMA technology, the low performance and inefficiency problems caused by span die access in the NUMA architecture are solved, and efficient data access is achieved.
Patent Information
- Application Number
- PCT/CN2024/122512
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-17
- Filing Date
- 2024-09-29
- Publication Date
- 2025-07-24
AI Technical Summary
When processors of NUMA architecture access across dies, the access paths are long, resulting in high consumption of computer resources and low access performance and efficiency.
In the disk array, each CPU core group is bound to an exclusive physical network card, physical hard disk and physical memory, and establishes a communication link that is not probed by other hosts and CPU core groups through network devices. It uses RDMA technology to achieve exclusive access and avoid cross-die access.
Shorten access paths, save computer resources, and improve access performance and efficiency.
Smart Images

Figure CN2024122512_24072025_PF_FP_ABST
Abstract
Description
Data processing system, method, device and non-volatile readable storage medium
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to the Chinese patent application filed with the China Patent Office on January 17, 2024, with application number 202410066357.3, and application name “A Data Processing System, Method, Device and Medium”, all contents of which are incorporated by reference into this application. Technical Field
[0003] The present application relates to the field of computer technology, and in particular to a data processing system, method, device, and non-volatile readable storage medium. Background Art
[0004] Currently, CPUs with a NUMA (Non-Uniform Memory Access) architecture often access data across multiple dies, where a single die contains a subset of CPU cores. Cross-die accesses result in longer access paths and consume more computer resources, limiting performance and efficiency. By partitioning the CPU cores within a processor, N CPU cores can be grouped into a single die.
[0005] Therefore, how to solve the low access performance and low access efficiency caused by cross-die access is a problem that needs to be solved by those skilled in the art.
[0006] Summary of the Invention
[0007] In view of this, the purpose of this application is to provide a data processing system, method, device, and non-volatile readable storage medium to solve the low access performance and low access efficiency caused by cross-die access. The optional solutions are as follows:
[0008] In a first aspect, the present application provides a data processing system, comprising: a disk array, a network device, and at least one host;
[0009] The disk array includes: multiple CPU core groups, each CPU core group is bound to a dedicated physical network card, physical hard disk and physical memory that are not detected by other CPU core groups;
[0010] The network device is configured to: enable the disk array and the at least one host to communicate;
[0011] Any host is configured to: detect the physical network card, physical hard disk and physical memory bound to each CPU core group in the disk array through the network device; establish a communication link with the physical network card exclusive to a single CPU core group among the multiple CPU core groups that is not detected by other hosts and other CPU core groups, and access the current CPU core group, as well as the physical hard disk and physical memory bound to the current CPU core group through the communication link.
[0012] Optionally, each of the CPU core groups and the physical network card, the physical hard disk, and the physical memory bound thereto are located in the same physical area.
[0013] Optionally, the physical locations of the different CPU core groups and the bound physical network cards, the physical hard disks, and the physical memories are in the same circuit area of the processor integrated circuit of the disk array.
[0014] Optionally, the disk array is configured to: divide all CPU cores in the disk array into different CPU core groups according to a load balancing strategy to obtain the multiple CPU core groups;
[0015] Correspondingly, the disk array is also configured to bind at least one physical memory instance, at least one physical network card instance, and at least one hard disk instance to each CPU core group, so that each CPU core group is bound to an exclusive physical network card, physical hard disk, and physical memory that are not detected by other CPU core groups.
[0016] Optionally, the memory instance is the control process of the physical memory, the network card instance is the control process of the physical network card, and the hard disk instance is the control process of the physical hard disk;
[0017] Optionally, the network device is a network switch;
[0018] Accordingly, the network switch is configured to enable the disk array and the at least one host to communicate using an Ethernet protocol.
[0019] Optionally, any host is specifically configured to connect to the remote direct data access (RDMA) network port of the physical network card dedicated to a single CPU core group among the multiple CPU core groups to establish an RDMA communication link that is not detected by other hosts and other CPU core groups, and access the current CPU core group, as well as the physical hard disk and physical memory bound to the current CPU core group through the RDMA communication link.
[0020] Optionally, any host is further configured to create an access queue and a memory area for each communication link in itself.
[0021] Optionally, the physical hard disk is a non-volatile memory host controller interface specification NVMe hard disk and / or a serial attached SCSI interface protocol SAS hard disk;
[0022] Accordingly, any host is specifically configured to access the physical hard disk bound to the current CPU core group through the NVMe protocol and / or SAS protocol.
[0023] Optionally, the disk array further includes: a main controller;
[0024] Correspondingly, the main controller is configured to: receive an access request sent by any host, and forward the access request to a corresponding CPU core group according to a destination port of the access request.
[0025] Optionally, any CPU core group is specifically set as follows: if the received access request is a read request, a memory area is applied for in the physical memory bound to the current CPU core group, the target data to be read by the read request is read from the physical hard disk bound to the current CPU core group to the applied memory area, and the target data in the memory area is sent to the host corresponding to the read request through the physical network card bound to the current CPU core group.
[0026] Optionally, any CPU core group is further configured to release memory resources occupied by the read request.
[0027] Optionally, any CPU core group is also configured to: while reading the target data to be read by the read request from the physical hard disk bound to the current CPU core group to the applied memory area, read associated data whose correlation with the target data is greater than a preset threshold from the physical hard disk to the applied memory area; after sending the target data in the memory area to the host corresponding to the read request through the physical network card bound to the current CPU core group, retain the target data and the associated data in the applied memory area.
[0028] Optionally, any CPU core group is specifically configured to calculate the correlation between data key information of other data in the physical hard disk bound to the current CPU core group and data key information of the target data; data key information includes: data storage address, hash value and / or metadata.
[0029] Optionally, any CPU core group is specifically set to: if the received access request is a write request, then apply for a memory area in the physical memory bound to the current CPU core group, write the data to be written by the write request into the applied memory area, write the data in the applied memory area into the physical hard disk bound to the current CPU core group, and return a write request completion message to the host corresponding to the write request through the physical network card bound to the current CPU core group.
[0030] In a second aspect, the present application provides a data processing method, which is applied to any host that communicates with a disk array through a network device, wherein the disk array includes: multiple CPU core groups, each CPU core group is bound to a dedicated physical network card, a physical hard disk, and a physical memory that is not detected by other CPU core groups;
[0031] The method includes:
[0032] Detecting the physical network card, physical hard disk, and physical memory bound to each CPU core group in the disk array through the network device;
[0033] Establishing a communication link with a physical network card dedicated to a single CPU core group among the multiple CPU core groups, which is not detected by other hosts and other CPU core groups;
[0034] The current CPU core group, and the physical hard disk and physical memory bound to the current CPU core group are accessed through the communication link.
[0035] Optionally, the host and the disk array discover each other through the Ethernet protocol, and the host and the RDMA network port on the physical network card dedicated to each CPU core group in the disk array establish a connection, and create a queue and data memory for interactive data, wherein there are multiple RDMA network ports on the physical network card dedicated to one CPU core group, one RDMA network port allows connection to one host, and the same host allows connection to the RDMA network ports on multiple physical network cards dedicated to the CPU core group.
[0036] Optionally, there is an exclusive communication link between the host and each of the CPU core groups, and the communication link includes multiple paths. The host discovers the physical hard disk on the CPU core group through the communication link of each of the CPU core groups, and manages and accesses data on the physical hard disk.
[0037] Optionally, the disk array divides all its CPU cores into different CPU core groups according to the load balancing strategy to obtain the multiple CPU core groups; binds at least one physical memory instance, at least one physical network card instance and at least one hard disk instance to each CPU core group, so that each CPU core group is bound to an exclusive physical network card, physical hard disk and physical memory that are not detected by other CPU core groups.
[0038] Optionally, the network device is a network switch; the network switch uses an Ethernet protocol to enable the disk array and the at least one host to communicate.
[0039] Optionally, any host is connected to the RDMA network port of the physical network card dedicated to a single CPU core group among the multiple CPU core groups to establish an RDMA communication link that is not detected by other hosts and other CPU core groups, and accesses the current CPU core group, as well as the physical hard disk and physical memory bound to the current CPU core group through the RDMA communication link.
[0040] Optionally, any host creates an access queue and memory area in itself for each communication link.
[0041] Optionally, any host accesses the physical hard disk bound to the current CPU core group through the NVMe protocol and / or SAS protocol.
[0042] Optionally, if the access request received by any CPU core group is a read request, a memory area is requested in the physical memory bound to the current CPU core group, target data to be read by the read request is read from the physical hard disk bound to the current CPU core group to the requested memory area, and the target data in the memory area is sent to the host corresponding to the read request via the physical network card bound to the current CPU core group. Memory resources occupied by the read request are released.
[0043] Optionally, any CPU core group reads the target data to be read by the read request from the physical hard disk bound to the current CPU core group to the applied memory area, and at the same time reads associated data whose correlation with the target data is greater than a preset threshold from the physical hard disk to the applied memory area; after sending the target data in the memory area to the host corresponding to the read request through the physical network card bound to the current CPU core group, the target data and the associated data in the applied memory area are retained.
[0044] Optionally, any CPU core group calculates the correlation between data key information of other data in the physical hard disk bound to the current CPU core group and data key information of the target data; data key information includes: data storage address, hash value and / or metadata.
[0045] Optionally, if the access request received by any CPU core group is a write request, a memory area is applied for in the physical memory bound to the current CPU core group, the data to be written by the write request is written into the applied memory area, the data in the applied memory area is written into the physical hard disk bound to the current CPU core group, and a write request completion message is returned to the host corresponding to the write request through the physical network card bound to the current CPU core group.
[0046] In a third aspect, the present application provides an electronic device, comprising:
[0047] a memory arranged to store a computer program;
[0048] The processor is configured to execute the computer program to implement the aforementioned disclosed data processing method.
[0049] In a fourth aspect, the present application provides a non-volatile readable storage medium, which is configured to store a computer program, wherein the computer program implements the aforementioned disclosed data processing method when executed by a processor.
[0050] It can be seen from the above scheme that the present application provides a data processing system, including: a disk array, a network device and at least one host; the disk array includes: multiple CPU core groups, each CPU core group is bound to an exclusive physical network card, physical hard disk and physical memory that are not detected by other CPU core groups; the network device is configured to: enable the disk array and the at least one host to communicate; any host is configured to: detect the physical network card, physical hard disk and physical memory bound to each CPU core group in the disk array through the network device; establish a communication link with the physical network card exclusive to a single CPU core group among the multiple CPU core groups that is not detected by other hosts and other CPU core groups, and access the current CPU core group, and the physical hard disk and physical memory bound to the current CPU core group through the communication link.
[0051] It can be seen that the technical effect of the present application is: each CPU core group of the disk array is bound to an exclusive physical network card, physical hard disk and physical memory that are not detected by other CPU core groups, thereby realizing the isolation and division of physical network card resources, physical hard disk resources and physical memory resources in the disk array; any host detects the physical network card, physical hard disk and physical memory bound to each CPU core group in the disk array through a network device; the host also establishes an exclusive communication link with the physical network card exclusive to a single CPU core group in multiple CPU core groups, which is not detected by other hosts and other CPU core groups. Then, a CPU core group (i.e., a die) and a host have a dedicated communication link, so that the host can access the CPU core group and the bound physical hard disk and physical memory without cross-die access. The access path is shortened, computer resources are saved, and access performance and access efficiency are improved.
[0052] Correspondingly, the data processing method, device and non-volatile readable storage medium provided by this application also have the above-mentioned technical effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.
[0054] FIG1 is a schematic diagram of a data processing system disclosed in this application;
[0055] FIG2 is a schematic diagram of another data processing system disclosed in this application;
[0056] FIG3 is a flow chart of a data processing method disclosed in this application;
[0057] FIG4 is a schematic diagram of an electronic device disclosed in this application;
[0058] FIG5 is a structural diagram of a server provided by this application;
[0059] FIG6 is a diagram of a terminal structure provided by this application. DETAILED DESCRIPTION
[0060] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0061] Currently, processors with a NUMA architecture often access data across multiple dies, where a single die contains a subset of CPU cores. Cross-die accesses result in longer access paths and consume more computer resources, limiting access performance and efficiency. By partitioning the different CPU cores within a processor, N CPU cores can be grouped into a single die. To address this issue, this application provides a data processing solution that avoids cross-die accesses and addresses the poor performance and efficiency associated with cross-die accesses.
[0062] As shown in FIG1 , an embodiment of the present application discloses a data processing system comprising: a disk array, a network device, and at least one host. The disk array comprises: multiple CPU core groups, each CPU core group being bound to a dedicated physical network card, a physical hard disk, and a physical memory that are not detected by other CPU core groups. The network device is configured to enable the disk array to communicate with at least one host. Any host is configured to: detect the physical network card, the physical hard disk, and the physical memory that are bound to each CPU core group in the disk array through the network device; establish a communication link that is not detected by other hosts and other CPU core groups with the physical network card that is exclusive to a single CPU core group among the multiple CPU core groups, and access the current CPU core group, as well as the physical hard disk and physical memory that are bound to the current CPU core group through the communication link. A CPU core group is a die that includes multiple CPU cores. The communication link that is not detected by other hosts and other CPU core groups means that other hosts and other CPU core groups cannot perceive the existence of the communication link.
[0063] Each CPU core group, along with its associated physical network card, physical hard drive, and physical memory, resides in the same physical area. This means that the division of CPU cores is based on the physical layout of the devices, not just a logical division. More importantly, different CPU core groups and their associated physical network cards, physical hard drives, and physical memory are physically close together, within the same circuit area of the disk array's processor integrated circuit. This indicates that the physical layout of different CPU core groups is determined by circuit design.
[0064] In this embodiment, the disk array is configured to divide all CPU cores in the disk array into different CPU core groups according to the load balancing strategy, thereby obtaining a plurality of CPU core groups; accordingly, the disk array is further configured to bind at least one physical memory instance (i.e., the memory control process), at least one physical network card instance (i.e., the network card control process), and at least one hard disk instance (i.e., the hard disk control process) to each CPU core group, so that each CPU core group is bound to a dedicated physical network card, physical hard disk, and physical memory that is not detected by other CPU core groups. The load balancing strategy can evenly divide all CPU cores in the disk array into different CPU core groups. In other words, any CPU core group cannot detect the existence of its unbound physical network card, physical hard disk, and physical memory.
[0065] It should be noted that the host and disk array can discover each other via the Ethernet protocol. In one example, the network device is a network switch; accordingly, the network switch is configured to enable communication between the disk array and at least one host using the Ethernet protocol. Ethernet protocols such as RoCE (RDMA over Converged Ethernet) are examples.
[0066] It should be noted that any host can connect to the RDMA (Remote Direct Memory Access) network port of the physical network card dedicated to any CPU core group through a network device, thereby establishing a dedicated RDMA communication link. In one example, any host is specifically configured to connect to the RDMA network port of the physical network card dedicated to a single CPU core group among multiple CPU core groups to establish an RDMA communication link that is not detected by other hosts and other CPU core groups, and access the current CPU core group, as well as the physical hard disk and physical memory bound to the current CPU core group through the RDMA communication link. Accordingly, any host is also configured to create an access queue and a memory area for each communication link in itself.
[0067] In one example, the physical hard disk is an NVMe (Non-Volatile Memory express, non-volatile memory host controller interface specification) hard disk and / or a SAS (Serial Attached SCSI, serial attached SCSI interface protocol) hard disk; accordingly, any host is specifically configured to access the physical hard disk bound to the current CPU core group through the NVMe protocol and / or SAS protocol.
[0068] Furthermore, the disk array further includes: a main controller; accordingly, the main controller is configured to: receive an access request sent by any host, and forward the access request to a corresponding CPU core group according to a destination port of the access request.
[0069] In one example, any CPU core group is specifically configured to: if the received access request is a read request, request a memory area in the physical memory bound to the current CPU core group, read the target data of the read request from the physical hard disk bound to the current CPU core group into the requested memory area, and send the target data in the memory area to the host corresponding to the read request via the physical network card bound to the current CPU core group. Accordingly, any CPU core group is also configured to release the memory resources occupied by the read request.
[0070] It should be noted that any CPU core group is also configured to: read the target data to be read by the read request from the physical hard disk bound to the current CPU core group to the requested memory area, and at the same time read the associated data whose correlation with the target data is greater than a preset threshold from the physical hard disk to the requested memory area; after sending the target data in the memory area to the host corresponding to the read request through the physical network card bound to the current CPU core group, retain the target data and associated data in the requested memory area. Among them, any CPU core group also calculates the correlation between the data key information of other data in the physical hard disk bound to the current CPU core group and the data key information of the target data; the data key information includes: data storage address, hash value and / or metadata.
[0071] In one example, any CPU core group is specifically set as follows: if the received access request is a write request, a memory area is applied for in the physical memory bound to the current CPU core group, the data to be written by the write request is written into the applied memory area, the data in the applied memory area is written into the physical hard disk bound to the current CPU core group, and a write request completion message is returned to the host corresponding to the write request through the physical network card bound to the current CPU core group.
[0072] It can be seen that in this embodiment, each CPU core group in the disk array is bound to an exclusive physical network card, physical hard disk, and physical memory that are not detected by other CPU core groups, thereby achieving the isolation and division of physical network card resources, physical hard disk resources, and physical memory resources in the disk array; any host can detect the physical network card, physical hard disk, and physical memory bound to each CPU core group in the disk array through a network device; the host also establishes an exclusive communication link with the physical network card exclusive to a single CPU core group in multiple CPU core groups, which is not detected by other hosts and other CPU core groups. Then, a CPU core group (i.e., a die) and a host have a dedicated communication link, allowing the host to access the CPU core group and the bound physical hard disk and physical memory without cross-die access, thereby shortening the access path, saving computer resources, and improving access performance and efficiency.
[0073] Referring to Figure 2, the data processing system provided in this application includes n hosts, a network switch, and a disk array. The disk array includes four dies: die0, die1, die2, and die3, corresponding to four cores. Each die is bound to a dedicated 100G network card, NVMe disk, and memory. In this system, Ethernet and Remote Data Memory (RDMA) technologies are used to establish communication between the disk array and the hosts. RDMA enables hosts to access the disk array's storage space. Host access to the disk array's storage space follows the principle that the network card on the current die can only access the disks on the current die.
[0074] Each CPU core group, along with its associated physical network card, physical hard drive, and physical memory, resides in the same physical area. This means that the division of CPU cores is based on the physical layout of the devices, not just a logical division. More importantly, different CPU core groups and their associated physical network cards, physical hard drives, and physical memory are physically close together, within the same circuit area of the disk array's processor integrated circuit. This indicates that the physical layout of different CPU core groups is determined by circuit design.
[0075] The host can be a server, which is the initiator of data access. The host establishes a communication connection with the disk array through the Ethernet protocol. The host also creates a hardware queue set to receive data, and then sends and receives data through the RDMA data link.
[0076] In some embodiments, the host and the disk array discover each other through the Ethernet protocol. The host and the disk array establish a connection with the RDMA network port on each die-specific network card, and create a queue and data memory for interactive data. Among them, there are multiple RDMA network ports on a die-specific network card, and one RDMA network port can be connected to one host. The same host can be connected to the RDMA network ports on multiple die-specific network cards. In this case, there is an exclusive communication link between the host and each die; the same host can be connected to multiple RDMA network ports on a single die-specific network card. In this case, there is an exclusive communication link between the host and the die, but the communication link includes multiple paths. The host discovers the hard disk on each die through the communication link of the die, and manages the hard disk and accesses data.
[0077] The host records the information of the hard drives found on each communication link and sends read and write requests through the communication link. The corresponding die in the disk array receives the read and write requests, processes the data using RDMA technology, and returns a response to the host. For read requests, the host sends a read request to the corresponding die through the communication link. The die then requests memory within itself, reads the data into the requested memory using a hard drive access protocol such as NVMe / SAS, then sends the data to the host through the die's network port using RDMA. The read response is then returned to the host via the Ethernet connection. Write requests are similar and will not be described here.
[0078] This embodiment divides the CPU core, memory, and network card by NUMA die, and realizes the division of memory, network card, and hard disk resources by die through instance binding. The data access process provided can avoid data access across NUMA die. It can be seen that by grouping the resources of PCIe (Peripheral Component Interconnect express, a high-speed serial computer expansion bus standard) network card devices, memory resources, and CPU cores according to NUMA die in the disk array, cross-die access is avoided by connecting to each die in a way that only the disk on the current die is seen, and ultimately high performance is achieved. At the same time, RDMA technology is used to realize the separation of storage and computing between the disk array and the host and the free expansion of the disk array storage space, which can provide higher performance access capabilities.
[0079] The following introduces a data processing method provided in an embodiment of the present application. The data processing method described below can be referenced with other embodiments described in this document.
[0080] As shown in Figure 3, an embodiment of the present application discloses a data processing method, which is applied to any host, which communicates with a disk array through a network device. The disk array includes: multiple CPU core groups, each CPU core group is bound to an exclusive physical network card, physical hard disk and physical memory that are not detected by other CPU core groups.
[0081] The method provided in the embodiment of the present application includes:
[0082] S301 , detecting the physical network card, physical hard disk, and physical memory bound to each CPU core group in the disk array through a network device.
[0083] S302 : Establish a communication link with a physical network card dedicated to a single CPU core group among the multiple CPU core groups, which is not detected by other hosts and other CPU core groups.
[0084] S303: Access the current CPU core group, and the physical hard disk and physical memory bound to the current CPU core group through the communication link.
[0085] Each CPU core group, along with its associated physical network card, physical hard drive, and physical memory, resides in the same physical area. This means that the division of CPU cores is based on the physical layout of the devices, not just a logical division. More importantly, different CPU core groups and their associated physical network cards, physical hard drives, and physical memory are physically close together, within the same circuit area of the disk array's processor integrated circuit. This indicates that the physical layout of different CPU core groups is determined by circuit design.
[0086] In one example, a disk array divides all of its CPU cores into different CPU core groups according to a load balancing strategy, thereby obtaining multiple CPU core groups. At least one physical memory instance, at least one physical network card instance, and at least one hard disk instance are bound to each CPU core group, so that each CPU core group is bound to an exclusive physical network card, physical hard disk, and physical memory that are not detected by other CPU core groups.
[0087] In one example, the network device is a network switch; the network switch uses an Ethernet protocol to enable the disk array and at least one host to communicate.
[0088] In one example, any host connects to the RDMA (Remote Direct Memory Access) network port of the physical network card dedicated to a single CPU core group among multiple CPU core groups to establish an RDMA communication link that is not detected by other hosts and other CPU core groups, and accesses the current CPU core group, as well as the physical hard disk and physical memory bound to the current CPU core group through the RDMA communication link.
[0089] In one example, any host creates an access queue and memory area within itself for each communication link.
[0090] In one example, any host accesses the physical hard disk bound to the current CPU core group through the NVMe protocol and / or the SAS protocol.
[0091] In one example, if the access request received by any CPU core group is a read request, a memory area is requested in the physical memory bound to the current CPU core group. The target data to be read from the physical hard disk bound to the current CPU core group is read into the requested memory area. The target data in the memory area is then sent to the host corresponding to the read request via the physical network card bound to the current CPU core group. Memory resources occupied by the read request are released.
[0092] In one example, any CPU core group reads the target data to be read in the read request from the physical hard disk bound to the current CPU core group to the applied memory area, and at the same time reads associated data whose correlation with the target data is greater than a preset threshold from the physical hard disk to the applied memory area; after sending the target data in the memory area to the host corresponding to the read request through the physical network card bound to the current CPU core group, the target data and associated data in the applied memory area are retained.
[0093] In one example, any CPU core group calculates the correlation between data key information of other data in the physical hard disk bound to the current CPU core group and data key information of the target data; the data key information includes: data storage address, hash value and / or metadata.
[0094] In one example, if the access request received by any CPU core group is a write request, a memory area is applied for in the physical memory bound to the current CPU core group, the data to be written by the write request is written into the applied memory area, the data in the applied memory area is written into the physical hard disk bound to the current CPU core group, and a write request completion message is returned to the host corresponding to the write request through the physical network card bound to the current CPU core group.
[0095] Among them, regarding the optional working processes of each module and unit in this embodiment, reference can be made to the corresponding contents disclosed in the aforementioned embodiments, and no further details will be given here.
[0096] It can be seen that in this embodiment, each CPU core group in the disk array is bound to an exclusive physical network card, physical hard disk, and physical memory that are not detected by other CPU core groups, thereby realizing the isolation and division of physical network card resources, physical hard disk resources, and physical memory resources in the disk array; any host establishes an exclusive communication link that is not detected by other hosts and other CPU core groups through a network device with the physical network card exclusive to a single CPU core group in multiple CPU core groups. In this way, a CPU core group (i.e., a die) and a host communicate via a dedicated communication link, allowing the host to access the CPU core group and the bound physical hard disk and physical memory without cross-die access. This shortens the access path, saves computer resources, and improves access performance and efficiency.
[0097] An electronic device provided in an embodiment of the present application is introduced below. The electronic device described below can be referenced with other embodiments described herein.
[0098] As shown in FIG4 , an embodiment of the present application discloses an electronic device, including:
[0099] a memory arranged to store a computer program;
[0100] The processor is configured to execute a computer program to implement the method disclosed in any of the above embodiments.
[0101] In this embodiment, when the processor executes the computer program stored in the memory, it can specifically implement the following steps: detect the physical network card, physical hard disk and physical memory bound to each CPU core group in the disk array through the network device; establish a communication link with the physical network card exclusive to a single CPU core group among multiple CPU core groups that is not detected by other hosts and other CPU core groups; access the current CPU core group, as well as the physical hard disk and physical memory bound to the current CPU core group through the communication link.
[0102] In this embodiment, when the processor executes the computer program stored in the memory, it can specifically implement the following steps: divide all the CPU cores in itself into different CPU core groups according to the load balancing strategy to obtain multiple CPU core groups; bind at least one physical memory memory instance, at least one physical network card network card instance and at least one hard disk hard disk instance to each CPU core group, so that each CPU core group is bound to an exclusive physical network card, physical hard disk and physical memory that cannot be detected by other CPU core groups.
[0103] In this embodiment, when the processor executes the computer program stored in the memory, it can specifically implement the following steps: connect to the RDMA (Remote Direct Memory Access) network port of the physical network card dedicated to a single CPU core group among multiple CPU core groups to establish an RDMA communication link that is not detected by other hosts and other CPU core groups, and access the current CPU core group, as well as the physical hard disk and physical memory bound to the current CPU core group through the RDMA communication link.
[0104] In this embodiment, when the processor executes the computer program stored in the memory, the following steps may be specifically implemented: creating an access queue and a memory area for each communication link in the processor.
[0105] In this embodiment, when the processor executes the computer program stored in the memory, the following steps can be specifically implemented: accessing the physical hard disk bound to the current CPU core group through the NVMe protocol and / or SAS protocol.
[0106] In this embodiment, when the processor executes a computer program stored in the memory, the following steps may be specifically implemented: if the received access request is a read request, a memory area is requested in the physical memory bound to the current CPU core group, target data to be read by the read request is read from the physical hard disk bound to the current CPU core group into the requested memory area, and the target data in the memory area is sent to the host corresponding to the read request via the physical network card bound to the current CPU core group. Memory resources occupied by the read request are released.
[0107] In this embodiment, when the processor executes the computer program stored in the memory, the following steps can be specifically implemented: while reading the target data to be read by the read request from the physical hard disk bound to the current CPU core group to the applied memory area, the associated data whose correlation with the target data is greater than a preset threshold is read from the physical hard disk to the applied memory area; after sending the target data in the memory area to the host corresponding to the read request through the physical network card bound to the current CPU core group, the target data and associated data in the applied memory area are retained.
[0108] In this embodiment, when the processor executes the computer program stored in the memory, it can specifically implement the following steps: calculate the correlation between the data key information of other data in the physical hard disk bound to the current CPU core group and the data key information of the target data; the data key information includes: data storage address, hash value and / or metadata.
[0109] In this embodiment, when the processor executes a computer program stored in the memory, the following steps can be specifically implemented: if the received access request is a write request, a memory area is applied for in the physical memory bound to the current CPU core group, the data to be written by the write request is written into the applied memory area, the data in the applied memory area is written into the physical hard disk bound to the current CPU core group, and a write request completion message is returned to the host corresponding to the write request through the physical network card bound to the current CPU core group.
[0110] Furthermore, an embodiment of the present application also provides an electronic device. The electronic device can be either a server as shown in FIG5 or a terminal as shown in FIG6. FIG5 and FIG6 are both structural diagrams of electronic devices according to an exemplary embodiment, and the contents of the diagrams are not to be construed as limiting the scope of use of the present application.
[0111] Figure 5 is a schematic diagram of the structure of a server provided in an embodiment of the present application. The server may include: at least one processor, at least one memory, a power supply, a communication interface, an input / output interface, and a communication bus. The memory is configured to store a computer program, which is loaded and executed by the processor to implement the relevant steps of the data processing disclosed in any of the aforementioned embodiments.
[0112] In this embodiment, the power supply is configured to provide operating voltage for each hardware device on the server; the communication interface can create a data transmission channel between the server and external devices, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of this application and is not limited here; the input and output interface is configured to obtain external input data or output data to the outside world, and its optional interface type can be selected according to specific application needs and is not limited here.
[0113] In addition, the memory as a carrier for resource storage can be a read-only memory, random access memory, disk or CD, etc. The resources stored thereon include operating system, computer programs and data, etc. The storage method can be temporary storage or permanent storage.
[0114] The operating system is used to manage and control the hardware devices and computer programs on the server, enabling the processor to operate and process data in the memory. It can be Windows Server, NetWare, Unix, Linux, etc. In addition to computer programs capable of performing the data processing methods disclosed in any of the aforementioned embodiments, computer programs can also include computer programs capable of performing other specific tasks. Data can include data such as application update information and other data such as application developer information.
[0115] FIG6 is a schematic structural diagram of a terminal provided in an embodiment of the present application. The terminal may include but is not limited to a smart phone, a tablet computer, a laptop computer, or a desktop computer.
[0116] Generally, the terminal in this embodiment includes: a processor and a memory.
[0117] The processor may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor may also include a main processor and a coprocessor. The main processor is a processor configured to process data in an awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor configured to process data in a standby state. In some embodiments, the processor may be integrated with a GPU (Graphics Processing Unit), which is configured to be responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor may also include an AI (Artificial Intelligence) processor, which is configured to process computing operations related to machine learning.
[0118] The memory may include one or more computer non-volatile readable storage media, which may be non-transitory. The memory may also include high-speed random access memory, and non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In this embodiment, the memory is at least configured to store the following computer program, wherein, after the computer program is loaded and executed by the processor, it can implement the relevant steps in the data processing method performed by the terminal side disclosed in any of the aforementioned embodiments. In addition, the resources stored in the memory may also include an operating system and data, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system may include Windows, Unix, Linux, etc. The data may include but is not limited to update information of the application.
[0119] In some embodiments, the terminal may further include a display screen, an input and output interface, a communication interface, a sensor, a power supply, and a communication bus.
[0120] Those skilled in the art will appreciate that the structure shown in FIG6 does not limit the terminal and may include more or fewer components than shown in the figure.
[0121] A non-volatile readable storage medium provided in an embodiment of the present application is introduced below. The non-volatile readable storage medium described below can be referenced with other embodiments described herein.
[0122] A non-volatile readable storage medium configured to store a computer program, wherein the computer program, when executed by a processor, implements the data processing method disclosed in the aforementioned embodiments. The non-volatile computer readable storage medium, serving as a carrier for resource storage, may be a read-only memory, random access memory, a magnetic disk, or an optical disk, and the resources stored thereon may include an operating system, a computer program, and data, and the storage method may be either transient or permanent.
[0123] In this embodiment, the computer program executed by the processor can specifically implement the following steps: detecting the physical network card, physical hard disk and physical memory bound to each CPU core group in the disk array through the network device; establishing a communication link with the physical network card exclusive to a single CPU core group among multiple CPU core groups that is not detected by other hosts and other CPU core groups; accessing the current CPU core group, as well as the physical hard disk and physical memory bound to the current CPU core group through the communication link.
[0124] In this embodiment, the computer program executed by the processor can specifically implement the following steps: divide all CPU cores in the processor into different CPU core groups according to the load balancing strategy to obtain multiple CPU core groups; bind at least one physical memory memory instance, at least one physical network card network card instance and at least one hard disk hard disk instance to each CPU core group, so that each CPU core group is bound to an exclusive physical network card, physical hard disk and physical memory that is not detected by other CPU core groups.
[0125] In this embodiment, the computer program executed by the processor can specifically implement the following steps: connect to the RDMA (Remote Direct Memory Access) network port of the physical network card dedicated to a single CPU core group among multiple CPU core groups to establish an RDMA communication link that is not detected by other hosts and other CPU core groups, and access the current CPU core group, as well as the physical hard disk and physical memory bound to the current CPU core group through the RDMA communication link.
[0126] In this embodiment, the computer program executed by the processor may specifically implement the following steps: creating an access queue and a memory area for each communication link in the processor.
[0127] In this embodiment, the computer program executed by the processor can specifically implement the following steps: access the physical hard disk bound to the current CPU core group through the NVMe protocol and / or SAS protocol.
[0128] In this embodiment, the computer program executed by the processor can specifically implement the following steps: if the received access request is a read request, a memory area is requested in the physical memory bound to the current CPU core group, target data to be read by the read request is read from the physical hard disk bound to the current CPU core group into the requested memory area, and the target data in the memory area is sent to the host corresponding to the read request via the physical network card bound to the current CPU core group. Memory resources occupied by the read request are released.
[0129] In this embodiment, the computer program executed by the processor can specifically implement the following steps: while reading the target data to be read by the read request from the physical hard disk bound to the current CPU core group to the applied memory area, read the associated data whose correlation with the target data is greater than a preset threshold from the physical hard disk to the applied memory area; after sending the target data in the memory area to the host corresponding to the read request through the physical network card bound to the current CPU core group, retain the target data and associated data in the applied memory area.
[0130] In this embodiment, the computer program executed by the processor can specifically implement the following steps: calculate the correlation between the data key information of other data in the physical hard disk bound to the current CPU core group and the data key information of the target data; the data key information includes: data storage address, hash value and / or metadata.
[0131] In this embodiment, the computer program executed by the processor can specifically implement the following steps: if the received access request is a write request, a memory area is applied for in the physical memory bound to the current CPU core group, the data to be written by the write request is written into the applied memory area, the data in the applied memory area is written into the physical hard disk bound to the current CPU core group, and a write request completion message is returned to the host corresponding to the write request through the physical network card bound to the current CPU core group.
[0132] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0133] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM (Compact Disc Read Only Memory), or any other form of non-volatile readable storage medium known in the art.
[0134] Optional examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core idea of the present application. At the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A data processing system, characterized in that, Comprising: a disk array, a network device, and at least one host; The disk array includes: multiple CPU core groups, and each CPU core group is bound with an exclusive physical network card, physical hard disk, and physical memory that cannot be probed by other CPU core groups; The network device is configured to: enable communication between the disk array and the at least one host; Any host is configured to: probe the physical network card, physical hard disk, and physical memory bound to each CPU core group in the disk array through the network device; establish a communication link that cannot be probed by other hosts and other CPU core groups with the physical network card exclusive to a single CPU core group among the multiple CPU core groups, and access the current CPU core group, as well as the physical hard disk and physical memory bound to the current CPU core group, through the communication link.
2. The system according to claim 1, wherein Each of the CPU core groups and the physical network card, physical hard disk, and physical memory bound to it are located in the same physical area.
3. The system according to claim 1, wherein The physical locations of different CPU core groups and the physical network cards, physical hard disks, and physical memories bound to them are in the same circuit area of the processor integrated circuit of the disk array.
4. The system according to claim 1, characterized in that, The disk array is configured to: divide all CPU cores in itself into different CPU core groups according to a load balancing strategy to obtain the multiple CPU core groups; Correspondingly, the disk array is further configured to: bind a memory instance of at least one physical memory, a network card instance of at least one physical network card, and a hard disk instance of at least one physical hard disk to each CPU core group, so that each CPU core group is bound with an exclusive physical network card, physical hard disk, and physical memory that cannot be probed by other CPU core groups.
5. The system according to claim 4, characterized in that, The memory instance is the control process of the physical memory, the network card instance is the control process of the physical network card, and the hard disk instance is the control process of the physical hard disk.
6. The system according to claim 1, wherein The network device is a network switch; Correspondingly, the network switch is configured to: enable communication between the disk array and the at least one host using the Ethernet protocol.
7. The system according to claim 1, characterized in that Any host is specifically configured to: connect to the remote direct memory access (RDMA) network port of the physical network card exclusive to a single CPU core group among the multiple CPU core groups to establish an RDMA communication link that cannot be probed by other hosts and other CPU core groups, and access the current CPU core group, as well as the physical hard disk and physical memory bound to the current CPU core group, through the RDMA communication link.
8. The system according to claim 1, wherein Any host is further configured to: create an access queue and a memory area for each communication link in itself.
9. The system according to claim 1, characterized in that, The physical hard disk is a non-volatile memory host controller interface specification (NVMe) hard disk and / or a serial attached SCSI (SAS) interface protocol hard disk; Correspondingly, any host is specifically configured to: access the physical hard disk bound to the current CPU core group through the NVMe protocol and / or the SAS protocol.
10. The system according to any one of claims 1 to 9, characterized in that, The disk array further includes: a main controller; Correspondingly, the main controller is configured to: receive an access request sent by any host and forward the access request to the corresponding CPU core group according to the destination port of the access request.
11. The system according to any one of claims 1 to 9, characterized in that, Any CPU core group is specifically configured as follows: If the received access request is a read request, an area in the physical memory bound to the current CPU core group is allocated, the target data to be read by the read request is read from the physical hard disk bound to the current CPU core group into the allocated memory area, and the target data in the memory area is sent to the host corresponding to the read request through the physical network card bound to the current CPU core group.
12. The system according to claim 11, wherein Any CPU core group is also configured to release the memory resources occupied by the read request.
13. The system according to claim 11, wherein Any CPU core group is also configured to: while reading the target data to be read by the read request from the physical hard disk bound to the current CPU core group into the allocated memory area, read the associated data whose association degree with the target data is greater than a preset threshold from the physical hard disk into the allocated memory area; after sending the target data in the memory area to the host corresponding to the read request through the physical network card bound to the current CPU core group, retain the target data and the associated data in the allocated memory area.
14. The system according to claim 13, wherein Any CPU core group is specifically configured to calculate the association degree between the data key information of other data in the physical hard disk bound to the current CPU core group and the data key information of the target data. The data key information includes: data storage address, hash value, and / or metadata.
15. The system according to any one of claims 1 to 9, characterized in that, Any CPU core group is specifically configured as follows: If the received access request is a write request, an area in the physical memory bound to the current CPU core group is allocated, the data to be written by the write request is written into the allocated memory area, the data in the allocated memory area is written into the physical hard disk bound to the current CPU core group, and a write request completion message is returned to the host corresponding to the write request through the physical network card bound to the current CPU core group.
16. A data processing method, characterized in that, Applied to any host, the host communicates with a disk array through a network device. The disk array includes: multiple CPU core groups, and each CPU core group is bound to a dedicated physical network card, physical hard disk, and physical memory that cannot be probed by other CPU core groups. The method includes: Probing the physical network card, physical hard disk, and physical memory bound to each CPU core group in the disk array through the network device. Establishing a communication link that cannot be probed by other hosts and other CPU core groups with the physical network card dedicated to a single CPU core group among the multiple CPU core groups. Accessing the current CPU core group, as well as the physical hard disk and physical memory bound to the current CPU core group, through the communication link.
17. The method according to claim 16, characterized in that, The host and the disk array discover each other through the Ethernet protocol. A connection is established between the host and the RDMA network interface on the physical network card dedicated to each CPU core group in the disk array, and a queue and data memory for data interaction are created. There are multiple RDMA network interfaces on the physical network card dedicated to one CPU core group, and one RDMA network interface allows connection to one host, and the same host allows connection to the RDMA network interfaces on the physical network cards dedicated to multiple CPU core groups.
18. The method according to claim 17, wherein There is an exclusive communication link between the host and each of the CPU core groups. The communication link includes multiple paths. The host discovers the physical hard disks on each CPU core group through the communication link of each CPU core group, and manages and accesses data from the physical hard disks.
19. An electronic device, characterized in that, Comprising: A memory configured to store a computer program; A processor configured to execute the computer program to implement the method as claimed in claim 16.
20. A non-volatile readable storage medium, characterized in that, Configured to save a computer program, wherein the computer program, when executed by the processor, implements the method as claimed in claim 16.
Citation Information
Patent Citations
Computing system and data transmission method
CN116841946A
Network card mixed nucleophilic hardware binding method and device and storage medium
CN116866167A
Solid state disk configuration management method and device, computer equipment and storage medium
CN117311646A
Data processing system, method, equipment and medium
CN117591450A
Storage System Multiprocessing and Mutual Exclusion in a Non-Preemptive Tasking Environment
US20170090999A1
Cited By
Hard disk sequence identification method, electronic equipment and server
CN120723849A