Data processing method and device, electronic equipment and storage medium
By mapping the independent memory areas of the local disk and the second target device to a unified address space, and using data to accelerate hardware devices and handling hardware devices for data processing and transmission, the performance bottleneck caused by CPU dependence in the prior art is solved, and efficient data processing and transmission is achieved.
Patent Information
- Application Number
- CN202510138883.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-30
AI Technical Summary
Existing CPU-based data processing solutions have performance bottlenecks in high-throughput big data processing scenarios, resulting in increased data transmission delay and reduced processing efficiency.
By mapping the independent memory areas of the local disk and the second target device into a unified address space, data acceleration hardware devices and handling hardware devices are used to accelerate data processing and transmission, the CPU is avoided frequently scheduled memory access and the multiple transfers of data between different devices or memory are reduced.
The data handling process that reduces CPU participation is realized, eliminates the data transmission bottleneck during memory access based on the central processor, optimizes the memory usage efficiency, and improves the performance of data processing.
Smart Images

Figure CN120066410A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computers, and specifically to a data processing method, apparatus, electronic device, and storage medium. Background Art
[0002] The existing technical solutions mainly build a framework based on the CPU to perform operations such as compression, decompression, encryption, decryption, and data transfer. However, since the processing process depends on the CPU, this solution leads to too many data transfers in the CPU memory, resulting in a performance bottleneck. Frequent data transfers not only increase the burden of memory access but also increase data transmission latency, affecting the overall processing efficiency. In addition, the computing power and memory bandwidth of the CPU are relatively limited, unable to fully utilize the advantages of large-scale data processing, thereby restricting the maximum throughput of the system.
[0003] Therefore, in the high-throughput big data processing scenario, this CPU-based solution has significant performance bottlenecks and cannot meet the requirements of large-scale data transmission and processing. Summary of the Invention
[0004] Embodiments of this application provide a data processing method, apparatus, electronic device, and storage medium, which can improve the performance of data processing.
[0005] Embodiments of this application provide a data processing method, including:
[0006] Mapping the local disk and the independent memory areas of the second target device into a unified address space, and the unified address space is used as a shared memory area accessed by multiple hardware devices;
[0007] Using a data acceleration hardware device to accelerate the data to be processed in the first target device to obtain accelerated data, and writing the accelerated data into the local disk in the unified address space;
[0008] Performing data processing on the accelerated data in the local disk to obtain processed data;
[0009] Using a transfer hardware device to transfer the processed data in the local disk in the unified address space to the second target device in the unified address space.
[0010] Embodiments of this application also provide a data processing apparatus, including:
[0011] A mapping unit, configured to map the local disk and the independent memory areas of the second target device into a unified address space, and the unified address space is used as a shared memory area accessed by multiple hardware devices;
[0012] An acceleration unit for accelerating the data to be processed in the first target device using a data acceleration hardware device to obtain accelerated data and writing the accelerated data into a local disk in a unified address space;
[0013] A processing unit for performing data processing on the accelerated data in the local disk to obtain processed data;
[0014] A transfer unit for transferring the processed data in the local disk in the unified address space to the second target device in the unified address space using a transfer hardware device.
[0015] In some embodiments, the local disk is a node disk of a computing node in a computing device cluster, and the processing unit is used for:
[0016] Performing data redistribution processing among multiple node disks of the computing device cluster so that the accelerated data in the node disks is redistributed as processed data.
[0017] In some embodiments, the data to be processed includes multiple compressed data segments, and the acceleration unit is used for:
[0018] Using a data acceleration hardware device to decompress each compressed data segment to obtain multiple accelerated data segments;
[0019] For each data segment, writing the accelerated data segment into a local disk in a unified address space.
[0020] In some embodiments, the processing unit includes:
[0021] A grouping subunit for locally grouping the accelerated data segments in the local disk to obtain grouped data;
[0022] A sending subunit for separately sending each grouped data in the local disk to other local disks.
[0023] In some embodiments, the sending subunit is used for:
[0024] Using a data acceleration hardware device to compress each grouped data in the local disk in the unified address space to obtain compressed grouped data;
[0025] Using a transfer hardware device to separately send each compressed grouped data in the local disk to other local disks so that a data acceleration hardware device can decompress each compressed grouped data in other local disks in the unified address space to obtain processed data.
[0026] In some embodiments, a handling hardware device is used to move the processed data in the local disk in the unified address space to a second target device in the unified address space, including:
[0027] Using a data acceleration hardware device to perform compression processing on the processed data in the local disk in the unified address space to obtain compressed data;
[0028] Using a handling hardware device to move the compressed data in the local disk in the unified address space to a second target device in the unified address space.
[0029] In some embodiments, using a data acceleration hardware device to perform compression processing on each packet of data in the local disk in the unified address space to obtain compressed packet data, including:
[0030] Obtaining the packet data in the local disk in the unified address space;
[0031] Performing block processing on the packet data to obtain a plurality of packet data blocks;
[0032] Using the compression parameters configured by the data acceleration hardware device to perform compression processing on the plurality of packet data blocks simultaneously to obtain a plurality of compressed packet data blocks;
[0033] Merging the plurality of compressed packet data blocks to obtain compressed packet data.
[0034] In some embodiments, using a data acceleration hardware device to perform decompression processing on each compressed packet data in other local disks in the unified address space to obtain processed data, including:
[0035] Obtaining the compressed packet data in the local disk in the unified address space;
[0036] Performing block processing on the compressed packet data to obtain a plurality of compressed packet data blocks;
[0037] Using the decompression parameters configured by the data acceleration hardware device to perform decompression processing on the plurality of compressed packet data blocks simultaneously to obtain a plurality of decompressed data blocks;
[0038] Merging the plurality of decompressed data blocks to obtain processed data.
[0039] In some embodiments, the independent memory area of the local disk includes a controller memory buffer, and the independent memory area of the second target device includes an internal buffer.
[0040] In some embodiments, the independent memory area of the local disk includes a controller memory buffer, and the mapping unit is used for:
[0041] Obtaining the physical address of the controller memory buffer in the local disk;
[0042] Register the physical address of the controller memory buffer into the address mapping table of the system kernel to obtain the physical address upon completion of local disk registration;
[0043] Map the physical address upon completion of local disk registration into the unified address space.
[0044] In some embodiments, the independent memory area of the second target device includes an internal buffer, and the mapping unit is configured to:
[0045] Obtain the physical address of the internal buffer in the second target device;
[0046] Map the physical address of the internal buffer in the second target device into the unified address space;
[0047] Configure the access controller of the second target device such that the second target device accesses the internal buffer through the unified address space.
[0048] An embodiment of the present application further provides an electronic device, including a memory storing multiple instructions; the processor loads the instructions from the memory to execute the steps in any one of the data processing methods provided by the embodiments of the present application.
[0049] An embodiment of the present application further provides a computer-readable storage medium storing multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps in any one of the data processing methods provided by the embodiments of the present application.
[0050] Embodiments of the present application can map the local disk and the independent memory area of the second target device into the unified address space, and the unified address space serves as a shared memory area accessed by multiple hardware devices; use a data acceleration hardware device to accelerate the data to be processed in the first target device to obtain accelerated data, and write the accelerated data into the local disk in the unified address space; perform data processing on the accelerated data in the local disk to obtain processed data; use a transfer hardware device to transfer the processed data in the local disk in the unified address space to the second target device in the unified address space.
[0051] The unified address space proposed in the embodiments of the present application supports the collaboration of multiple hardware devices, enabling data sharing and access between different storage devices. Each hardware device can directly access and process data, avoiding multiple transfers of data between different devices or memories and no longer relying on the central processing unit to frequently schedule memory accesses. This method eliminates the data transfer bottleneck during memory access based on the central processing unit and optimizes the memory usage efficiency. The acceleration operations such as data compression and decompression of the central processing unit can be transferred to the data acceleration hardware device, and at the same time, the data transfer process of the central processing unit can be transferred to the transfer hardware device, avoiding the participation of the central processing unit and reducing the performance bottleneck in traditional data transfer. Thus, the embodiments of the present application improve the performance of data processing.
[0052] The embodiments of the present application can be applied to various environments that require high throughput, large-scale data processing, and high requirements for security and computing efficiency. For example, data compression and decompression in big data computing hardware devices, CRC check (Cyclic Redundancy Check) in the financial field, encryption and decryption operations in cloud computing, real-time video stream and large-scale image processing, data encryption and processing of Internet of Things devices, data synchronization and backup in large-scale distributed databases, and many other scenarios.
[0053] For example, the embodiments of the present solution can be applied to shuffle operations in large-scale distributed architectures, verification of financial transaction records, large-scale data encryption and decryption operations in cloud computing environments, etc. The present solution can be applied to multiple scenarios with high throughput and high security requirements. The execution processes of multiple hardware devices (such as data acceleration hardware devices and transfer hardware devices) can be executed in parallel, no longer relying on the traditional synchronization waiting mechanism and no longer relying on the traditional central processing unit computing, so that operations such as encryption and decryption, decompression, shuffle operation, and CRC check no longer become performance bottlenecks, thereby improving the throughput of the entire data processing flow, reducing latency, and improving system efficiency. Brief Description of the Drawings
[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0055] Figure 1a It is a schematic diagram of the shuffle scenario of the data processing method provided by the embodiments of the present application;
[0056] Figure 1b It is a flowchart of the data processing method provided by the embodiments of the present application;
[0057] Figure 1c It is a schematic diagram of the transfer scenario of the data processing method provided by the embodiment of the present application;
[0058] Figure 1d It is a schematic diagram of the shuffle scenario of the data processing method provided by the embodiment of the present application;
[0059] Figure 2a It is a schematic diagram of the traditional data transfer scenario;
[0060] Figure 2b It is a schematic diagram of the data transfer scenario provided by the embodiment of the present application;
[0061] Figure 2c It is a schematic diagram of the process of applying the data processing method provided by the embodiment of the present application in the shuffle scenario;
[0062] Figure 3 It is a schematic diagram of the structure of the data processing device provided by the embodiment of the present application;
[0063] Figure 4 It is a schematic diagram of the structure of the electronic device provided by the embodiment of the present application. Detailed implementation manners
[0064] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0065] The embodiment of the present application provides a data processing method, device, electronic device and storage medium.
[0066] Among them, the data processing device can be specifically integrated in an electronic device, and the electronic device can be a device such as a terminal or a server. Among them, the terminal can be a device such as a mobile phone, a tablet computer, a smart Bluetooth device, a laptop computer, or a personal computer (PC); the server can be a single server or a server cluster composed of multiple servers.
[0067] In some embodiments, the data processing device can also be integrated in multiple electronic devices. For example, the data processing device can be integrated in multiple servers, and the data processing method of the present application is implemented by multiple servers.
[0068] In some embodiments, the terminal can also be used as a server to implement some or all of the functions of the server.
[0069] For example, referring to Figure 1a , an embodiment of the present application provides a data processing system, which includes a computing device cluster, multiple first target devices, and multiple second target devices. The computing device cluster may include multiple computing device nodes. The system may be initialized in advance, that is, the local disk and the independent memory area of the second target device are mapped into a unified address space, and the unified address space is used as a shared memory area accessed by multiple hardware devices. Then, a data acceleration hardware device is used to accelerate the data to be processed in the first target device to obtain accelerated data, and the accelerated data is written into the local disk in the unified address space. A reallocation (shuffle) process is performed on the accelerated data in the local disk to obtain processed data. A transfer hardware device is used to transfer the processed data in the local disk in the unified address space to the second target device in the unified address space.
[0070] The following will be described in detail respectively. It should be noted that the serial numbers of the following embodiments do not limit the preferred order of the embodiments.
[0071] In this embodiment, a data processing method is provided. As Figure 1b shown, the specific process of the data processing method may be as follows:
[0072] 110. Map the local disk and the independent memory area of the second target device into a unified address space, and the unified address space is used as a shared memory area accessed by multiple hardware devices.
[0073] The local disk refers to a disk storage device installed inside or directly connected to a computing node in the computing device cluster, and is used to store the operating system, application programs, and user data. Among them, the local disk may be a solid-state drive (SSD, Solid State Drive), a hard disk drive (HDD, Hard Disk Drive), a hybrid hard disk (SSHD, Solid State Hybrid Drive), an optical disc (CD, DVD, Blu-ray), a network attached storage (NAS, Network Attached Storage), and so on.
[0074] The independent memory area refers to a memory area with independent management and access permissions.
[0075] Hardware devices refer to hardware components specifically designed to accelerate specific computing tasks, data transfer, or storage operations, which are used to improve the overall performance of the system. For example, hardware devices can be DPU (Data Processing Unit), SNIC (Storage Network Interface Card), GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), SSD (Solid-State Drive), RAM (Random Access Memory), data acceleration hardware devices, transfer hardware devices, and so on.
[0076] Among them, the first target device and the second target device can be the same device or different devices. The first target device can be any software and hardware device or system, and the second target device can be any hardware device.
[0077] For example, in some embodiments, the first target device can be a software system, such as HDFS (Hadoop Distributed File System); the second target device can be a hardware device, such as DPU.
[0078] Hardware devices refer to hardware components specifically designed to accelerate specific computing tasks, data transfer, or storage operations, which are usually used to relieve the burden on the CPU and improve the overall performance of the system.
[0079] When a process needs to transfer data from one place to another, the operating system will perform multiple copy operations. For example, when data is received from a disk or network, the data is first copied to the buffer in the operating system kernel; then, it is copied from the kernel buffer to the buffer in the user space (used by the application); finally, the application may need to send the data to a network or storage device. For example, if the traditional solution is used to transfer the data in the local disk to the second target device, the central processing unit (CPU) needs to read the data to be transferred into the DRAM cache through its memory DRAM (Dynamic Random Access Memory), such as DDR (Double Data Rate), SDRAM (Synchronous DRAM), etc., and then transfer it from the DRAM cache to the second target device.
[0080] Zero-Copy is an optimization technique aimed at reducing redundant copy operations when data is transferred between memory and storage, thereby improving performance. In the embodiments of this application, in order to implement zero-copy, a unified address space is proposed, which supports the cooperation of multiple hardware devices, maximizes data processing efficiency and throughput, and reduces the burden on the central processing unit.
[0081] The embodiments of this application propose a unified address space, which is used as a shared memory area accessed by multiple hardware devices. Among them, the unified address space can be a Unified Memory Architecture (UMA) or a Unified Virtual Address Space (UVAS).
[0082] Among them, UMA is a memory architecture, which means that multiple hardware devices share a unified physical memory pool and do not require complex memory allocation and management. UVAS is a system architecture in which all hardware components share the same virtual address space, and hardware devices can directly access the data in the same virtual address space without address translation.
[0083] For example, referring to Figure 1c , in some embodiments, the independent memory areas of the local disk and the second target device are mapped to the unified address space. Then, when moving data from the local disk to the second target device, the CPU does not need to read the data through its DDR, but directly moves the data in the local disk to the independent memory area of the second target device through the channels of the moving hardware device (such as IOAT).
[0084] In some embodiments, the independent memory areas of all hardware devices can be mapped to the unified address space so that the hardware devices can access these hardware devices.
[0085] In some embodiments, the independent memory area of the local disk includes a controller memory buffer. The disk controller memory buffer is a dedicated memory located in the storage device controller, which is used to cache the data to be written or read. For example, the independent memory area of an SSD includes a CMB (Controller Memory Buffer).
[0086] In some embodiments, the independent memory area of the second target device includes an Internal Buffer. For example, the independent memory areas of DPU, GPU, etc. include an internal buffer, which is used to optimize data transmission, store temporary data, or accelerate the calculation process.
[0087] Thus, in some embodiments, the independent memory region of the local disk includes a controller memory buffer. Mapping the independent memory region of the local disk into the unified address space includes:
[0088] Obtain the physical address of the controller memory buffer in the local disk;
[0089] Register the physical address of the controller memory buffer into the address mapping table of the system kernel to obtain the physical address after the local disk registration is completed;
[0090] Map the physical address after the local disk registration is completed into the unified address space.
[0091] The controller memory buffer is an independent memory region inside the disk controller, and its physical address is hardware-specific and usually cannot be directly accessed. Among them, its physical address can be obtained through a hardware interface (such as PCIe). Registering this address into the address mapping table of the kernel can ensure that the system can manage and map this memory region so as to map the physical address into the unified address space.
[0092] Thus, in some embodiments, the independent memory region of the second target device includes an internal buffer. Mapping the independent memory region of the second target device into the unified address space includes:
[0093] Obtain the physical address of the internal buffer in the second target device;
[0094] Map the physical address of the internal buffer in the second target device into the unified address space;
[0095] Configure the access controller of the second target device so that the second target device accesses the internal buffer through the unified address space.
[0096] The internal buffer in the second target device is an independent memory region inside the second target device, and its physical address is hardware-specific and usually cannot be directly accessed. Among them, its physical address can be obtained through a hardware interface (such as PCIe).
[0097] In some embodiments, to ensure that the second target device can access its internal buffer through the unified address space, the system needs to configure the access controller of the device and set appropriate permissions and access rules. This usually involves an access control mechanism (such as permission setting) to ensure that only authorized devices or operations can access these memory regions.
[0098] In step 110, by mapping to a unified address space, all hardware devices (including CPUs, GPUs, storage devices, etc.) can directly access and share the same virtual memory. This simplifies data exchange and communication between hardware components, eliminating the need to handle complex address translations or transfer data through intermediaries (such as DMA). This mapping method avoids frequent memory copies or direct I / O operations, enhancing the performance of hardware devices. Especially in scenarios with frequent data transfers, there is no need to transfer data through the main memory, thus reducing latency and bandwidth bottlenecks.
[0099] For example, in some embodiments, all device address spaces on a server can be mapped to a unified memory address space, enabling data copies between heterogeneous memories to be treated as ordinary memory address spaces visible to the CPU.
[0100] 120. Use a data acceleration hardware device to accelerate the processing of the data to be processed in the first target device, obtaining accelerated data, and write the accelerated data into the local disk in the unified address space.
[0101] A data acceleration hardware device refers to a hardware device specifically designed to accelerate data processing, transmission, or computing tasks. Through specific hardware acceleration technologies, such as encryption / decryption, compression / decompression, etc., it significantly improves the processing speed and reduces the computational burden on the CPU. For example, a data acceleration hardware device can be QAT (Quick Assist Technology), GPU, ASIC (Application-Specific Integrated Circuit, an integrated circuit designed for specific tasks), TPM (a dedicated encryption chip, Trusted Platform Module), HSM (a dedicated hardware device for encryption / decryption and key management, Hardware Security Module), Intel SGX (a hardware-enhanced encryption technology, SoftwareGuard Extensions), and so on.
[0102] Among them, the acceleration processing can include processing solutions for accelerating data processing, transmission, or computing tasks such as encryption / decryption, compression / decompression. Using a data acceleration hardware device for encryption / decryption, compression / decompression, etc. to accelerate data processing can significantly improve the processing speed and reduce latency when dealing with large-scale data.
[0103] Accelerated data is the data after hardware acceleration, referring to the result after optimized processing. For example, after the original data undergoes encryption, compression, or other computing tasks, a more efficient processing result is obtained. This accelerated data can be decrypted data, compressed data, etc.
[0104] In some embodiments, a data acceleration hardware device can significantly accelerate data processing tasks through technologies such as parallel processing, large-scale computing optimization, and hardware acceleration algorithms. This process enables tasks that originally required long CPU execution times to be completed in a shorter time, thereby improving the overall processing efficiency. For example, in some embodiments, step 120 can use QAT to accelerate the data to be processed in the SSD, obtain the accelerated data, and write the accelerated data into the CMB of the SSD mapped in the UMA. Among them, by utilizing the CMB characteristics of the NVMe SSD, the processed data can be directly written to the disk through the memory buffer of the controller, without passing through the traditional transmission path from DRAM to SSD, but directly mapped to the local disk (such as NVMe SSD) through the UMA. This method can avoid the traditional data transmission path and directly achieve efficient data processing and storage with the help of hardware acceleration. By directly mapping the data to the UMA address space, efficient storage and access of the accelerated data are realized.
[0105] In some embodiments, accelerating the data to be processed in the first target device by using a data acceleration hardware device, obtaining the accelerated data, and writing the accelerated data into the local disk in the unified address space includes:
[0106] Obtain a file path, where the file path corresponds to the data to be processed in the first target device, and the data to be processed includes multiple compressed data segments;
[0107] According to the file path, use the data acceleration hardware device to decompress each compressed data segment to obtain multiple accelerated data segments;
[0108] For each data segment, write the accelerated data segment into the local disk in the unified address space.
[0109] The file path refers to the storage location of the data to be processed. For example, the file path of HDFS can be obtained.
[0110] Among them, the data to be processed includes multiple compressed data segments. For example, files used for large-scale data analysis or processing are usually split into multiple data blocks. In some embodiments, these data blocks may be stored in a compressed format (such as.gz,.zip,.tar.gz, etc.) to reduce storage space and accelerate data transmission. Therefore, it is necessary to use a data acceleration hardware device to decompress each compressed data segment.
[0111] Therefore, in some embodiments, step 120 can utilize QAT hardware acceleration to accelerate the compression and decompression operations, and then write the accelerated data to the local disk efficiently through the CMB feature of UMA and NVMe SSD, achieving fast data processing and storage. Step 120 can effectively reduce the bottleneck of data transmission and improve the processing capacity and storage performance of the entire system.
[0112] 130. Perform data processing on the accelerated data in the local disk to obtain processed data.
[0113] For example, the accelerated data refers to the original data obtained after decompressing the compressed data block in HDFS, that is, the accelerated data.
[0114] Among them, the specific steps of data processing depend on the data type and the requirements of the task, and can include various data processing schemes such as data redistribution, data cleaning, data conversion, data analysis, and data integration.
[0115] For example, the data processing can be data redistribution (Shuffle). Shuffle refers to reordering and distributing data to different nodes or partitions in a distributed computing framework (such as Spark, Mapreduce, etc.). Shuffle in big data scenarios requires high network overhead, which can lead to problems such as increased disk I / O, calculation latency, uneven load, and excessive CPU pressure.
[0116] To solve this problem, in some embodiments, the local disk is the node disk of a computing node in a computing device cluster, and step 130 includes performing data redistribution processing among multiple node disks of the computing device cluster, so that the accelerated data in the node disks is redistributed into processed data.
[0117] Among them, in some embodiments, step 130 includes:
[0118] Perform local grouping processing on the accelerated data segments in the local disk to obtain grouped data;
[0119] Send each grouped data in the local disk to other local disks respectively.
[0120] For example, referring to Figure 1d , after QAT decompresses the compressed data block [9, 2, 1] in HDFS and loads the decompressed data segments into the NVMe SSD of a computing node, the computing node can perform local grouping processing on it according to the partition key corresponding to the accelerated data segments to obtain group [2, 1] and group [9], and then send group [2, 1] and group [9] to the local disks of other computing nodes.
[0121] In some embodiments, in order to improve the transmission speed, a data acceleration hardware device may be used to separately send each packet of data in the local disk to other local disks.
[0122] In some embodiments, in order to reduce bandwidth consumption and improve the transmission speed, when sending packet data, a data acceleration hardware device may be used to compress the packet data, and before other local disks receive the compressed packet data, a data acceleration hardware device may be used to decompress the compressed packet data.
[0123] Therefore, in some embodiments, separately sending each packet of data in the local disk to other local disks includes:
[0124] Using a data acceleration hardware device to perform compression processing on each packet of data in the local disk in the unified address space to obtain compressed packet data;
[0125] Using a transfer hardware device to separately send each compressed packet of data in the local disk to other local disks, so that a data acceleration hardware device can perform decompression processing on each compressed packet of data in other local disks in the unified address space to obtain processed data.
[0126] In order to further optimize and accelerate computationally intensive tasks such as data compression, decompression, encryption, and decryption, the data acceleration hardware device can process data in parallel in blocks. Each data block can be independently processed by different processing units and can process multiple data blocks in parallel, significantly improving the processing speed. And the block operation helps manage the memory consumption of large-scale data and avoid memory overload. Therefore, it can efficiently process large-scale data.
[0127] For example, in some embodiments, using a data acceleration hardware device to perform compression processing on each packet of data in the local disk in the unified address space to obtain compressed packet data includes:
[0128] Obtaining the packet data in the local disk in the unified address space;
[0129] Performing block processing on the packet data to obtain multiple packet data blocks;
[0130] Using the compression parameters configured by the data acceleration hardware device to simultaneously perform compression processing on multiple packet data blocks to obtain multiple compressed packet data blocks;
[0131] Merging the multiple compressed packet data blocks to obtain compressed packet data.
[0132] In some embodiments, a data acceleration hardware device is used to decompress each compressed packet data in other local disks in the unified address space to obtain processed data, including:
[0133] Obtain the compressed packet data in the local disk in the unified address space;
[0134] Perform block processing on the compressed packet data to obtain multiple compressed packet data blocks;
[0135] Using the decompression parameters configured by the data acceleration hardware device, simultaneously decompress multiple compressed packet data blocks to obtain multiple decompressed data blocks;
[0136] Merge multiple decompressed data blocks to obtain processed data.
[0137] Among them, the compression parameters may include the block size of each block, compression level, dictionary size of the compression algorithm, window size, compression method, parallelism, etc.; the decompression parameters may include block size, decompression level, dictionary size of the decompression algorithm, window size, decompression method, parallelism, etc.
[0138] 140. Use a transfer hardware device to transfer the processed data in the local disk in the unified address space to the second target device in the unified address space.
[0139] The transfer hardware device is a hardware device specifically designed to accelerate or directly execute data transfer and transfer tasks. Its main function is to efficiently transfer data from one storage location to another without the intervention of the CPU, or reduce the burden on the CPU, thereby improving the overall performance of the system.
[0140] Among them, the transfer hardware device may include IOAT (I / O Acceleration Technology), NVMe over Fabrics, etc.
[0141] For example, IOAT can be used to transfer the processed data in the SSD in UMA to the DPU in UMA. Among them, by using IOAT, the processed data in UMA can be directly transferred from the SSD to the DPU. Through the direct memory access (DMA) and the acceleration ability of IOAT, the traditional CPU copy operation is avoided, and the data transfer efficiency is improved. Through hardware acceleration devices such as IOAT, data transfer no longer depends on the CPU, thereby releasing CPU resources for higher-level computing tasks; data flows directly in UMA, without the need for traditional memory copy and I / O operations, reducing latency and increasing throughput.
[0142] In some embodiments, to further improve the handling efficiency, the data acceleration hardware device and the handling hardware device can be jointly operated to perform data pressurization processing and handling in the unified address space (UMA). Therefore, by using the handling hardware device, the processed data in the local disk in the unified address space is handled into the second target device in the unified address space, including:
[0143] Using the data acceleration hardware device to pressurize the processed data in the local disk in the unified address space to obtain compressed data;
[0144] Using the handling hardware device to handle the compressed data in the local disk in the unified address space into the second target device in the unified address space.
[0145] As can be seen from the above, the embodiments of the present application can map the independent memory areas of the local disk and the second target device into the unified address space, and the unified address space is used as a shared memory area accessed by multiple hardware devices; using the data acceleration hardware device to accelerate the data to be processed in the first target device to obtain accelerated data, and writing the accelerated data into the local disk in the unified address space; performing data processing on the accelerated data in the local disk to obtain processed data; using the handling hardware device to handle the processed data in the local disk in the unified address space into the second target device in the unified address space.
[0146] Thus, this solution collaborates the data acceleration hardware device and the handling hardware device to perform efficient data compression and handling in the unified address space. By reducing the participation of the CPU, it optimizes data processing and transmission, and can significantly improve performance in scenarios of big data, high throughput, and low latency, thereby enhancing the performance of data processing.
[0147] According to the method described in the above embodiments, further detailed description will be made below.
[0148] Figure 2a For the traditional data handling solution, where the handling path of data between the NIC (Network Interface Card) and the NVMe SSD is NIC → DRAM → NVMe SSD, that is, data usually reaches the host memory (DRAM) in the CPU from the NIC through the PCIe (Peripheral Component Interconnect) bus; after the data is stored in the DRAM, the CPU accesses the data in the DRAM through the memory bus DDR, and the CPU issues a write command to the NVMe SSD through the PCIe channel, and the data is transmitted from the host memory to the NVMe SSD through the PCIe.
[0149] Refer to Figure 2b, in this embodiment, the transfer path is NIC → NVMe SSD. Through the cooperation between hardware devices supported by the unified address space, DRAM is bypassed to reduce latency and improve data transfer efficiency, and data transfer between different IO devices is reduced to achieve zero copy.
[0150] As Figure 2c shown, this embodiment will take Spark Shuffle as an example. For the Shuffle operation of the Shuffle Manager component in the Executor 1 (Executor JVM#1) within the Worker node, to achieve the transfer of data between the SSD and the SNIC or DPU, this example will detail the method of this application embodiment.
[0151] The specific process of a data processing method is as follows:
[0152] ① Shuffle Writer stage.
[0153] The data to be processed is stored as an object (obj) in the ByteBuffer (byte buffer). In the Shuffle operation, the data to be processed is serialized to Off-heap (off-heap memory) instead of Heap (heap memory) to improve performance and reduce computational pressure;
[0154] ② Storage stage.
[0155] The traditional solution writes shuffle data to the Spark local directory (Spark local.dir), while this solution uses the CMB of the NVMe SSD as a cache to directly store the compressed shuffle data; and since the CMB belongs to the high-speed storage area of the NVMe SSD internal controller, compared with ordinary SSD read and write, it has faster speed, lower latency, and can reduce the dependence on local disk IO.
[0156] In this solution, the QAT engine accelerates the data to be processed in the Kernel Space to obtain the accelerated data, and directly transfers the data to the unified address space (Unified Virtual Address Space), and directly writes it to the CMB address space of the NVMe SSD.
[0157] Among them, the QAT engine can offload computing tasks such as encryption, decryption, compression, and decompression that were originally executed by the CPU to this hardware to improve performance and reduce CPU load. In the embodiment of this application, using its direct read and write characteristics and combining with the CMB characteristics of the NVMe SSD, the data can be automatically written directly to the disk without other assistance.
[0158] Among them, the QAT engine can be used in various scenarios such as big data acceleration, WAN acceleration, HTTP compression, file system compression, database compression, etc. Its compression and decompression algorithms can include the DEFLATE algorithm. Among them, the DEFLATE algorithm first uses the LZ77 compression algorithm, and then through Huffman coding, uses the gzip or zlib header.
[0159] In addition, the QAT engine also has other features, such as the ability to configure the engine to perform compression or decompression, support for stateful (de)compression, support for storing specific functions, and so on.
[0160] When different processing platforms use QAT technology for data compression and decompression, they all have excellent performance effects. For example, for different processing platforms, such as the Coleto Creek platform and the Lewisburg platform, their performance is shown in the following table:
[0161] Performance Coleto Creek Lewisburg Compression 24 Gbps 100 Gbps Decompression 24 Gbps 160 Gbps Compression + Decompression 24 Gbps 100 Gbps
[0162] According to this table, when the QAT engine performs data compression and decompression operations on different hardware platforms, the Lewisburg platform has obvious advantages over the Coleto Creek in compression and decompression tasks.
[0163] ③ Shuffle Reader stage.
[0164] When reading data, use IOAT to transfer the processed data of CMB in the local disk (NVMe SSD) in the unified address space to the internal buffer of the second target device (SNIC / DPU) in the unified address space. The data can also be returned to the user space through the drivers for the Shuffle Reader to use.
[0165] The embodiments of the present application can be applied in a big data computing engine. For example, in order to reduce network bandwidth and disk usage, data is compressed, and at the same time, in application scenarios with high security requirements such as finance, CRC verification is performed on the data. In the embodiments of the present application, software encryption and decryption or CRC through CPU calculation usually wait synchronously, and only continue to execute after the encryption, decryption / CRC calculation is completed. By implementing these features through a dedicated hardware device on the CPU, software execution can be fully parallelized. That is, after the software initiates the encryption, decryption / CRC calculation, it can continue to execute other processes that do not depend on the calculation results of this part. While the hardware performs parallel calculations, the software does not stop to wait synchronously for the calculation results of this part. When the results are returned, the subsequent processes are continued through a callback mechanism such as message notification, fully streamlining the data processing process and maximizing the system throughput.
[0166] The embodiments of this application can be applied to the scenario of accelerating Spark shuffle calculations. To accelerate the Spark shuffle calculations, there are usually a large number of SSD disk writes and remote data reads. At the same time, to improve the calculation efficiency, the transmitted data is compressed to reduce the limitation of the data network bandwidth. The unified address space proposed in the embodiments of this application supports the collaborative and parallel work of multiple hardware devices. The software compression and decompression parts on the CPU can be done on the QAT; to reduce the number of memory copies, the transfer part can be done on the IOAT; at the same time, by directly reading and writing commands to the CMB, the process of the controller reading the command queue after a Host write command is avoided. Based on the above three hardware device characteristics, the embodiments of this application hand over the processes of software compression, decompression, data transfer, network card data sending and storage of Spark Shuffle to the above-mentioned IOAT, QAT and CMB to collaborate, reducing the large amount of data transfer processes involving the CPU in the middle.
[0167] The embodiments of this application can reduce the number of memory copies by introducing QAT and IOAT in the big data cluster, and through the proprietary compression and decompression of QAT, the native Spark performance can be improved by 30%.
[0168] The experimental scenario proposed in the embodiments of this application is based on the Intel SPR processor and the Spark computing engine. Through the SPR processor and the acceleration of the QAT engine for the compression and decompression of Spark Shuffle data, the end-to-end performance is improved by 5% - 20% compared with the software solution. Then, through the CMB feature, it is directly written to the NVMe SSD, and the overall performance is improved by 10% - 25%.
[0169] As can be seen from the above, the unified address space proposed in the embodiments of this application supports the collaborative and parallel work of multiple hardware devices, reduces the number of memory copies, and can improve the performance of data processing.
[0170] It can be understood that in the specific implementation of this application, it involves relevant data such as the data to be processed and files. When the following embodiments of this application are applied to specific products or technologies, permission or consent needs to be obtained, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions.
[0171] To better implement the above method, the embodiments of this application also provide a data processing device. This data processing device can be specifically integrated in an electronic device, and this electronic device can be devices such as a terminal, a server, etc. Among them, the terminal can be devices such as a mobile phone, a tablet computer, a smart Bluetooth device, a laptop computer, a personal computer, etc.; the server can be a single server or a server cluster composed of multiple servers.
[0172] For example, in this embodiment, taking the data processing device specifically integrated in the server cluster as an example, the method of the embodiment of the present application will be described in detail.
[0173] For example, as Figure 3 shown, the data processing device may include a mapping unit 310, an acceleration unit 320, a processing unit 330, and a transfer unit 340, as follows:
[0174] (1) Mapping unit 310.
[0175] The mapping unit 310 is used to map the local disk and the independent memory area of the second target device into a unified address space, and the unified address space is used as a shared memory area accessed by multiple hardware devices.
[0176] In some embodiments, the independent memory area of the local disk includes a controller memory buffer, and the mapping unit 310 is used for:
[0177] Obtain the physical address of the controller memory buffer in the local disk;
[0178] Register the physical address of the controller memory buffer into the address mapping table of the system kernel to obtain the registered physical address of the local disk;
[0179] Map the registered physical address of the local disk into the unified address space.
[0180] In some embodiments, the independent memory area of the second target device includes an internal buffer, and the mapping unit 310 is used for:
[0181] Obtain the physical address of the internal buffer in the second target device;
[0182] Map the physical address of the internal buffer in the second target device into the unified address space;
[0183] Configure the access controller of the second target device so that the second target device accesses the internal buffer through the unified address space.
[0184] (2) Acceleration unit 320.
[0185] The acceleration unit 320 is used to perform acceleration processing on the data to be processed in the first target device by using a data acceleration hardware device to obtain acceleration data, and write the acceleration data into the local disk in the unified address space.
[0186] In some embodiments, the acceleration unit 320 is used for:
[0187] Obtain a file path, which corresponds to the data to be processed in the first target device, and the data to be processed includes multiple compressed data segments;
[0188] According to the file path, use the data acceleration hardware device to decompress each compressed data segment to obtain multiple accelerated data segments;
[0189] For each data segment, write the accelerated data segment into the local disk in the unified address space.
[0190] (3) Processing unit 330.
[0191] The processing unit 330 is used to perform data processing on the accelerated data in the local disk to obtain processed data.
[0192] In some embodiments, the local disk is the node disk of a computing node in a computing device cluster, and the processing unit 330 is used to perform data redistribution processing among multiple node disks of the computing device cluster, so that the accelerated data in the node disks is redistributed into processed data.
[0193] In some embodiments, the processing unit 330 includes:
[0194] A grouping subunit, which includes performing local grouping processing on the accelerated data segments in the local disk to obtain grouped data;
[0195] A sending subunit, which includes sending each grouped data in the local disk to other local disks respectively.
[0196] In some embodiments, the sending subunit is used for:
[0197] Use the data acceleration hardware device to compress each grouped data in the local disk in the unified address space to obtain compressed grouped data;
[0198] Use the transfer hardware device to send each compressed grouped data in the local disk to other local disks respectively, so as to use the data acceleration hardware device to decompress each compressed grouped data in other local disks in the unified address space to obtain processed data.
[0199] In some embodiments, using the transfer hardware device to transfer the processed data in the local disk in the unified address space to the second target device in the unified address space includes:
[0200] Use the data acceleration hardware device to pressurize the processed data in the local disk in the unified address space to obtain compressed data;
[0201] Use a transfer hardware device to transfer the compressed data in the local disk in the unified address space to the second target device in the unified address space.
[0202] In some embodiments, use a data acceleration hardware device to compress each packet data in the local disk in the unified address space to obtain compressed packet data, including:
[0203] Obtain the packet data in the local disk in the unified address space;
[0204] Perform block processing on the packet data to obtain multiple packet data blocks;
[0205] Use the compression parameters configured by the data acceleration hardware device to simultaneously compress multiple packet data blocks to obtain multiple compressed packet data blocks;
[0206] Merge multiple compressed packet data blocks to obtain compressed packet data.
[0207] In some embodiments, use a data acceleration hardware device to decompress each compressed packet data in other local disks in the unified address space to obtain processed data, including:
[0208] Obtain the compressed packet data in the local disk in the unified address space;
[0209] Perform block processing on the compressed packet data to obtain multiple compressed packet data blocks;
[0210] Use the decompression parameters configured by the data acceleration hardware device to simultaneously decompress multiple compressed packet data blocks to obtain multiple decompressed data blocks;
[0211] Merge multiple decompressed data blocks to obtain processed data.
[0212] In some embodiments, the independent memory area of the local disk includes a controller memory buffer, and the independent memory area of the second target device includes an internal buffer.
[0213] (IV) Transfer unit 340.
[0214] The transfer unit 340 is used to use a transfer hardware device to transfer the processed data in the local disk in the unified address space to the second target device in the unified address space.
[0215] In specific implementation, each of the above units can be implemented as an independent entity, or can be combined arbitrarily and implemented as the same or several entities. For the specific implementation of each of the above units, reference can be made to the foregoing method embodiments, which will not be elaborated herein.
[0216] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the functions of that module or unit.
[0217] As can be seen from the above, in the data processing device of this embodiment, the mapping unit maps the local disk and the independent memory area of the second target device into a unified address space, and the unified address space is used as a shared memory area accessed by multiple hardware devices; the acceleration unit uses data acceleration hardware devices to accelerate the data to be processed in the first target device to obtain accelerated data, and writes the accelerated data into the local disk in the unified address space; the processing unit performs data processing on the accelerated data in the local disk to obtain processed data; the transfer unit uses transfer hardware devices to transfer the processed data in the local disk in the unified address space to the second target device in the unified address space. Thus, the embodiments of the present application can improve the performance of data processing.
[0218] The embodiments of the present application also provide an electronic device, which can be a device such as a terminal, a server, etc. Among them, the terminal can be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop computer, a personal computer, etc.; the server can be a single server or a server cluster composed of multiple servers, etc.
[0219] In some embodiments, the data processing device can also be integrated in multiple electronic devices. For example, the data processing device can be integrated in multiple servers, and multiple servers are used to implement the data processing method of the present application.
[0220] In this embodiment, the electronic device of this embodiment will be described in detail by taking the example that the electronic device is a server. For example, as Figure 4 shown, it shows a schematic structural diagram of the electronic device involved in the embodiments of the present application. Specifically:
[0221] The electronic device may include a processor 410 with one or more processing cores, a memory 420 with one or more computer-readable storage media, a power supply 430, an input module 440, a communication module 450 and other components. Those skilled in the art can understand that Figure 4 the structural diagram of the electronic device shown in does not constitute a limitation on the electronic device, and may include more or fewer components than shown, or combine some components, or have different component arrangements. Among them:
[0222] The processor 410 is the control center of the electronic device, connecting various parts of the entire electronic device through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 420, and by invoking the data stored in the memory 420, it executes various functions of the electronic device and processes data, thereby performing an overall detection of the electronic device. In some embodiments, the processor 410 may include one or more processing cores; in some embodiments, the processor 410 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 410 either.
[0223] The memory 420 can be used to store software programs and modules. The processor 410 executes various functional applications and data processing by running the software programs and modules stored in the memory 420. The memory 420 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, image playback function, etc.); the data storage area can store the data created according to the use of the electronic device. In addition, the memory 420 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory 420 may also include a memory controller to provide the processor 410 with access to the memory 420.
[0224] The electronic device further includes a power supply 430 that powers each component. In some embodiments, the power supply 430 may be logically connected to the processor 410 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 430 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0225] The electronic device may further include an input module 440, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.
[0226] The electronic device may further include a communication module 450. In some embodiments, the communication module 450 may include a wireless module. The electronic device can perform short-range wireless transmission through the wireless module of the communication module 450, thereby providing users with wireless broadband Internet access. For example, the communication module 450 can be used to help users send and receive emails, browse the web, and access streaming media, etc.
[0227] Although not shown, the electronic device may further include a display unit and the like, which will not be elaborated here. Specifically, in this embodiment, the processor 410 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 420 according to the following instructions, and the processor 410 will run the application programs stored in the memory 420 to implement various functions as follows:
[0228] Map the local disk and the independent memory area of the second target device to a unified address space, and the unified address space is used as a shared memory area accessed by multiple hardware devices;
[0229] Use a data acceleration hardware device to accelerate the data to be processed in the first target device to obtain accelerated data, and write the accelerated data into the local disk in the unified address space;
[0230] Perform data processing on the accelerated data in the local disk to obtain processed data;
[0231] Use a transfer hardware device to transfer the processed data in the local disk in the unified address space to the second target device in the unified address space.
[0232] For the specific implementation of each of the above operations, reference may be made to the previous embodiments, which will not be elaborated here.
[0233] As can be seen from the above, the embodiments of the present application can improve data processing performance.
[0234] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions or by controlling relevant hardware through instructions. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0235] For this reason, the embodiments of the present application provide a computer-readable storage medium, in which multiple instructions are stored, and these instructions can be loaded by a processor to execute the steps in any one of the data processing methods provided by the embodiments of the present application. For example, the instructions can execute the following steps:
[0236] Map the local disk and the independent memory area of the second target device to a unified address space, and the unified address space is used as a shared memory area accessed by multiple hardware devices;
[0237] Use a data acceleration hardware device to accelerate the data to be processed in the first target device to obtain accelerated data, and write the accelerated data into the local disk in the unified address space;
[0238] Perform data processing on the accelerated data in the local disk to obtain processed data;
[0239] Using a handling hardware device, the processed data in the local disk within the unified address space is moved into a second target device within the unified address space.
[0240] Among them, the storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk, optical disc, etc.
[0241] According to one aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the methods provided in various alternative implementations in the data processing aspect or the data encryption / decryption, compression / decompression, and handling aspects provided in the above embodiments.
[0242] Since the instructions stored in the storage medium can execute the steps in any of the data processing methods provided in the embodiments of the present application, the beneficial effects that can be achieved by any of the data processing methods provided in the embodiments of the present application can be realized. For details, see the previous embodiments and will not be repeated here.
[0243] The above has introduced in detail a data processing method, apparatus, electronic device, and computer-readable storage medium provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A data processing method, characterized in that: include: Mapping the local disk and the independent memory area of the second target device into a unified address space, wherein the unified address space is used as a shared memory area accessed by multiple hardware devices; Using a data acceleration hardware device to accelerate the data to be processed in the first target device to obtain accelerated data, and writing the accelerated data into the local disk in the unified address space; Performing data processing on the accelerated data in the local disk to obtain processed data; The processed data in the local disk in the unified address space is transported to the second target device in the unified address space by using a transport hardware device.
2. The data processing method according to claim 1, characterized in that: The local disk is a node disk of a computing node in a computing device cluster, and performing data processing on the accelerated data in the local disk to obtain processed data includes: Data redistribution processing is performed between the plurality of node disks of the computing device cluster, so that the accelerated data in the node disks are redistributed into processed data.
3. The data processing method according to claim 2, characterized in that: The data to be processed includes a plurality of compressed data segments, and the data acceleration hardware device is used to accelerate the data to be processed in the first target device to obtain accelerated data, and the accelerated data is written into the local disk in the unified address space, including: Decompressing each of the compressed data segments using a data acceleration hardware device to obtain a plurality of accelerated data segments; For each of the data segments, the accelerated data segment is written into a local disk in the unified address space.
4. The data processing method according to claim 3, characterized in that: The performing data processing on the accelerated data in the local disk to obtain processed data includes: Performing local grouping processing on the accelerated data segments in the local disk to obtain grouped data; Each group of data in the local disk is sent to other local disks respectively.
5. The data processing method according to claim 4, characterized in that: The step of sending each packet data in the local disk to other local disks respectively includes: Using a data acceleration hardware device to compress each packet data in the local disk in the unified address space to obtain compressed packet data; The transport hardware device is used to send each compressed grouped data in the local disk to other local disks respectively, so that the data acceleration hardware device is used to decompress each compressed grouped data in other local disks in the unified address space to obtain processed data.
6. The data processing method according to claim 5, characterized in that: The method of using a transport hardware device to transport the processed data in the local disk in the unified address space to a second target device in the unified address space includes: Using a data acceleration hardware device to compress the processed data in the local disk in the unified address space to obtain compressed data; The compressed data in the local disk in the unified address space is transferred to the second target device in the unified address space by using a transfer hardware device.
7. The data processing method according to claim 5, characterized in that: The method of using the data acceleration hardware device to compress each packet data in the local disk in the unified address space to obtain compressed packet data includes: Acquire the packet data in the local disk in the unified address space; Processing the packet data in blocks to obtain a plurality of packet data blocks; Using the compression parameters configured by the data acceleration hardware device, the plurality of grouped data blocks are compressed simultaneously to obtain a plurality of compressed grouped data blocks; The multiple compressed packet data blocks are merged to obtain compressed packet data.
8. The data processing method according to claim 7, characterized in that: The data acceleration hardware device is used to decompress each compressed packet data in other local disks in the unified address space to obtain processed data, including: Acquire the compressed packet data in the local disk in the unified address space; Processing the compressed packet data in blocks to obtain a plurality of compressed packet data blocks; Decompression parameters configured by the data acceleration hardware device are used to simultaneously decompress the plurality of compressed packet data blocks to obtain a plurality of decompressed data blocks; The multiple decompressed data blocks are merged to obtain processed data.
9. The data processing method according to claim 1, characterized in that: The independent memory area of the local disk includes a controller memory buffer, and the independent memory area of the second target device includes an internal buffer.
10. The data processing method according to claim 1, characterized in that: The independent memory area of the local disk includes a controller memory buffer, and mapping the independent memory area of the local disk to a unified address space includes: Get the physical address of the controller memory buffer in the local disk; Registering the physical address of the controller memory buffer into the address mapping table of the system kernel to obtain the physical address of the local disk registration; The physical address of the local disk registration is mapped into the unified address space.
11. The data processing method according to claim 1, characterized in that: The independent memory area of the second target device includes an internal buffer, and mapping the independent memory area of the second target device to the unified address space includes: Obtaining the physical address of the internal buffer in the second target device; Mapping the physical address of the internal buffer in the second target device into a unified address space; An access controller of the second target device is configured so that the second target device accesses the internal buffer through the unified address space.
12. A data processing device, characterized in that: include: A mapping unit, used for mapping the local disk and the independent memory area of the second target device into a unified address space, wherein the unified address space is used as a shared memory area accessed by multiple hardware devices; An acceleration unit, configured to accelerate the data to be processed in the first target device by using a data acceleration hardware device to obtain accelerated data, and write the accelerated data into a local disk in the unified address space; A processing unit, configured to perform data processing on the acceleration data in the local disk to obtain processed data; The transport unit is used to transport the processed data in the local disk in the unified address space to the second target device in the unified address space by using a transport hardware device.
13. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a plurality of instructions; the processor loads instructions from the memory to execute the steps in the data processing method according to any one of claims 1 to 11.
14. A computer program product or a computer program, wherein the computer program product or the computer program comprises computer instructions, wherein the computer instructions are stored in a computer-readable storage medium, wherein a processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the steps in the data processing method according to any one of claims 1 to 11.
15. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the data processing method according to any one of claims 1 to 11.