Data processing method, electronic equipment, storage medium and product

By introducing a transparent file system and a four-level pipeline software stack into the data processing method, and using the FPGA hardware platform to process data blocks in parallel, the problem of low data encryption, decryption and compression and decompression efficiency in the prior art is solved, efficient data transmission and transparent operation are achieved, and system performance and user experience are improved.

CN120234822AActive Publication Date: 2025-07-01INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510725299.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-07-01
Estimated Expiration
2045-05-30

AI Technical Summary

Technical Problem

In the prior art, data encryption, decryption and compression decompression methods are inefficient and cannot realize transparent operations, resulting in poor user experience, and hardware implementation cannot realize efficient data transmission and transparent operations.

Method used

By introducing a transparent file system into the data processing method, designing a four-level pipeline software stack, using the FPGA hardware platform to perform parallel processing of data blocks, and using point-to-point data transmission method to realize efficient data transmission, transparent compression encryption and transparent decryption and decompression of software and hardware collaboration.

Benefits of technology

It realizes efficient data transmission and transparent operations, improves the efficiency and security of data processing, reduces CPU resource consumption, and improves system performance and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234822A_ABST
    Figure CN120234822A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method, electronic equipment, a storage medium and a product, and relates to the field of data processing.The method comprises the steps that in response to a first operation instruction, a to-be-processed file is divided, and multiple data blocks corresponding to the to-be-processed file are obtained; a first processing operation is executed on a first data block in the multiple data blocks, a second processing operation is executed on a second data block in the multiple data blocks at the same time, the first data block is read before the second data block is read, and the first processing operation is an operation executed before the second processing operation according to the execution sequence of the multiple processing operations; the plurality of processing operations include at least two of a read process, a compression process, an encryption process, and a write process, or the plurality of processing operations include at least two of a read process, a decryption process, a decompression process, and a write process. And software and hardware collaborative parallel data processing can be realized, and the requirements of high-efficiency data transmission, high-efficiency transparent compression encryption and high-efficiency transparent decryption decompression can be met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing, and in particular, to a data processing method, an electronic device, a storage medium, and a product. Background Art

[0002] With the advent of the big data era, the security requirements for data have increased day by day, and a large amount of data has also grown accordingly. In order to protect the security of data, data encryption technology has become indispensable. In order to improve the storage efficiency of data, the data compression and decompression technology located in the data center has attracted more and more attention in the industry.

[0003] Currently, related technologies can use software methods to encrypt and decrypt data and perform software compression and decompression. However, using software methods for encryption, decryption, compression, and decompression requires a large amount of software computing resources, resulting in a decrease in the efficiency of other programs. In addition, some hardware implementation methods for encryption, decryption, compression, and decompression in related technologies are relatively cumbersome to operate and cannot achieve transparent operation, resulting in a poor user experience. Summary of the Invention

[0004] The present application provides a data processing method, an electronic device, a storage medium, and a product, so as to at least solve the problems in related technologies that the data encryption, decryption, compression, and decompression methods have low efficiency, and cannot achieve transparent compression and transparent encryption of files, nor can they achieve transparent decryption and transparent decompression of files.

[0005] The present application provides a data processing method, including: in response to a first operation instruction, dividing a file to be processed to obtain a plurality of data blocks corresponding to the file to be processed; performing a first processing operation on a first data block among the plurality of data blocks, and simultaneously performing a second processing operation on a second data block among the plurality of data blocks, where the first data block is read before the second data block, and the first processing operation is an operation that is executed before the second processing operation according to the execution order of a plurality of processing operations, and the plurality of processing operations include at least two of a reading processing, a compression processing, an encryption processing, and a writing processing, or the plurality of processing operations include at least two of a reading processing, a decryption processing, a decompression processing, and a writing processing.

[0006] The present application also provides a data processing device, including: a first processing unit, configured to divide a file to be processed in response to a first operation instruction, so as to obtain a plurality of data blocks corresponding to the file to be processed; a second processing unit, configured to perform a first processing operation on a first data block among the plurality of data blocks, and simultaneously perform a second processing operation on a second data block among the plurality of data blocks, where the first data block is read before the second data block is read, and the first processing operation is an operation that is executed before the second processing operation according to the execution order of a plurality of processing operations, and the plurality of processing operations include at least two of a reading processing, a compression processing, an encryption processing, and a writing processing, or the plurality of processing operations include at least two of a reading processing, a decryption processing, a decompression processing, and a writing processing.

[0007] The present application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any one of the above data processing methods when executing the computer program.

[0008] The present application also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any one of the above data processing methods are implemented.

[0009] The present application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of any one of the above data processing methods are implemented.

[0010] Through the present application, in response to a first operation instruction, a file to be processed is divided to obtain a plurality of data blocks corresponding to the file to be processed; a first processing operation is performed on a first data block among the plurality of data blocks, and a second processing operation is simultaneously performed on a second data block among the plurality of data blocks, where the first data block is read before the second data block is read, and the first processing operation is an operation that is executed before the second processing operation according to the execution order of a plurality of processing operations, and the plurality of processing operations include at least two of a reading processing, a compression processing, an encryption processing, and a writing processing, or the plurality of processing operations include at least two of a reading processing, a decryption processing, a decompression processing, and a writing processing. It is possible to implement software-hardware collaborative parallel data processing, and simultaneously meet the requirements of high-efficiency data transmission, high-efficiency transparent compression and encryption, and high-efficiency transparent decryption and decompression. Description of the Drawings

[0011] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0012] Figure 1Schematic flowchart of a data processing method provided by an embodiment of the present application; Figure 2 Schematic flowchart of another data processing method provided by an embodiment of the present application; Figure 3 Schematic flowchart of another data processing method provided by an embodiment of the present application; Figure 4 Schematic diagram of the topology structure of a PCIe bus protocol provided by an embodiment of the present application; Figure 5 Schematic diagram of a data transmission scheme of a traditional DMA provided by an embodiment of the present application; Figure 6 Schematic diagram of a point-to-point DMA data transmission scheme provided by an embodiment of the present application; Figure 7 Schematic diagram of realizing point-to-point data stream transmission between a disk and an FPGA provided by an embodiment of the present application; Figure 8 Schematic diagram of the single-step execution steps when a data encryption / decryption algorithm and a query algorithm are offloaded to FPGA hardware provided by an embodiment of the present application; Figure 9 Design scheme of a compression and encryption pipeline software stack provided by an embodiment of the present application; Figure 10 Serial system call method provided by an embodiment of the present application; Figure 11 Design of a four-stage pipeline software stack for data decryption and decompression provided by an embodiment of the present application; Figure 12 Schematic diagram of the composition structure of a compression and encryption, decryption and decompression heterogeneous computing data transmission scheme provided by an embodiment of the present application; Figure 13 Schematic flowchart of the specific implementation process of transparent compression and encryption of data provided by an embodiment of the present application; Figure 14 Schematic flowchart of the specific implementation process of transparent decryption and decompression of data provided by an embodiment of the present application; Figure 15 Schematic diagram of the structure of a data processing device provided by an embodiment of the present application. Detailed implementation manners

[0013] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0014] It should be noted that in the description of this application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects and not to describe a particular order or sequence.

[0015] With the advent of the big data era, the security requirements for data have increased day by day, and a huge amount of data has also grown significantly. To protect the security of data, data encryption technology has become indispensable. To improve the storage efficiency of data, data compression and decompression technology located in the data center has also attracted increasing attention in the industry.

[0016] For data security and data storage technology, traditional technologies can use software methods to encrypt and decrypt data and perform software compression and decompression. However, this traditional software encryption / decryption and software compression / decompression method has many disadvantages in the computer system of the data center. Whether it is software encryption / decryption or software compression / decompression, it belongs to the category of software acceleration technology and ultimately requires the integrated central processing unit (CPU) to execute. This will consume a large amount of CPU computing resources in the server and affect other software programs in the system. In addition, the efficiency of software acceleration technology is generally low and the acceleration effect is poor.

[0017] Traditional data compression / decompression and encryption / decryption are usually implemented at the software level, but this method has the following problems: 1) Slow processing speed: The data compression / decompression and data encryption implemented by software rely on the processing power of the CPU. For large-scale data processing, the performance bottleneck of the CPU will lead to a slower processing speed. 2) High CPU resource consumption: Software implementation requires a large amount of CPU and memory resources. Especially in a multi-tasking environment, resource competition may lead to a decline in performance.

[0018] To solve the above problems, hardware acceleration technology has been introduced into the data compression / decompression and encryption / decryption systems. Existing hardware solutions do not achieve efficient data transmission, do not achieve transparent operation, do not achieve a pipelined software stack design, cannot efficiently parallelly schedule software and hardware computing IPs, and cannot achieve the effect of efficient software and hardware cooperation.

[0019] Although the above solutions solve some problems of data processing to a certain extent, they cannot meet the requirements of high-speed data transmission, high security, and high storage efficiency at the same time. Therefore, how to achieve efficient data transmission while implementing transparent compression / decompression and transparent encryption / decryption calculations at the hardware level has become an urgent problem to be solved.

[0020] To solve the problems of slow data processing speed, high resource consumption, and lack of transparent operations in related solutions, the embodiments of the present application provide a data processing method. By introducing a transparent file system and designing two four-stage pipeline software stacks inside it, efficient data transmission, efficient data transparent compression and encryption, and efficient data transparent decryption and decompression are achieved, aiming to solve a series of problems in related technologies such as low data storage efficiency, high CPU resource consumption, low security, low transmission efficiency, poor flexibility and scalability, and provide an efficient, secure, and flexible data transmission, compression / decompression, encryption / decryption processing solution. To enable those skilled in the art of the present technology to better understand the solution of the present application, the following further details the present application in conjunction with the accompanying drawings and specific embodiments.

[0021] Figure 1 It is a schematic flowchart of a data processing method provided by an embodiment of the present disclosure.

[0022] As Figure 1 shown, this method can be executed by a host, and this method includes the following steps: Step 101, in response to a first operation instruction, divide the file to be processed to obtain a plurality of data blocks corresponding to the file to be processed.

[0023] In some embodiments, the file to be processed can be, for example, a video file, a document file, etc. The host of this solution can include an operating system and a file system. In this embodiment, the operating system is taken as Linux for illustration. It can be used to perform multiple operation processes on the file to be processed. The multiple processing operations include at least two of reading processing, compression processing, encryption processing, and writing processing, or the multiple processing operations include at least two of reading processing, decryption processing, decompression processing, and writing processing.

[0024] In some embodiments, the method further includes: when at least two of the multiple processing operations include read processing, compression processing, encryption processing, and write processing, determining that the first operation instruction is a data write operation instruction; when at least two of the multiple processing operations include read processing, decryption processing, decompression processing, and write processing, determining that the first operation instruction is a data read operation instruction. For example, when a file needs to be stored (written) or transmitted, read processing, compression processing, encryption processing, and write processing can be performed on the file to be processed. When the file to be processed needs to be read or transmitted, read processing, decryption processing, decompression processing, and write processing can be performed on the file to be processed, that is, when the file to be processed needs to be transmitted.

[0025] In some embodiments, the host can include a user layer and a kernel layer. The user layer and the kernel layer are two different layers in the host-side operating system. A user can trigger the first operation instruction in the user layer of the host. Optionally, after the user initiates a data write operation, the data write operation enters the kernel layer through a system call. The kernel layer can divide the file to be processed in response to the data write operation to obtain multiple data blocks corresponding to the file to be processed. Similarly, after the user initiates a data read operation, the data read operation enters the kernel layer through a system call. The kernel layer can divide the file to be processed in response to the data read operation to obtain multiple data blocks corresponding to the file to be processed.

[0026] Optionally, when the first operation instruction is a data write operation instruction, when the user triggers the first operation instruction, the user can input the file to be processed (the file to be written) that needs to be transmitted or stored. Then, when the data write operation enters the kernel layer through a system call, the file to be processed can be transmitted to the kernel layer. The kernel layer can first temporarily store the file to be processed in the memory of the operating system and then divide the file to be processed into blocks; when the first operation instruction is a data read operation instruction, after the user triggers the first operation instruction, the file to be processed can be divided into blocks at the location where the file to be processed (the file to be read) is stored.

[0027] In some embodiments, when dividing the file to be processed into multiple data blocks corresponding to the file to be processed, task blocks corresponding to each data block can also be generated. Each task block can include information related to the data block corresponding to the task block, that is, when generating the data block, information related to each data block can be determined, such as the size of the data block, the location of the data block, and so on. Optionally, the specific method for dividing the file to be processed is not limited in this disclosure. For example, the file to be processed can be divided according to the data volume, and so on.

[0028] Heterogeneous Computing: Heterogeneous computing refers to the use of computing resources with multiple different architectures, processors, or accelerators, combined together to achieve higher performance and higher efficiency in computing. This approach can improve the overall system performance by allocating tasks to the devices most suitable for executing them. Heterogeneous computing typically involves different types of processors such as CPUs, Graphics Processing Units (GPUs), and Field-Programmable Gate Arrays (FPGAs) to accelerate and optimize various applications.

[0029] As a programmable hardware platform, FPGA has high flexibility and parallel processing capabilities, making it very suitable for implementing complex heterogeneous accelerated computing data processing tasks. However, although there are currently some solutions that implement the functions of FPGA encryption / decryption and compression / decompression, such solutions generally only focus on the hardware implementation process, that is, they cannot achieve transparent compression and encryption of files at the software layer, nor can they achieve transparent decryption and decompression of files. Currently, the existing solutions also lack efficient data transmission characteristics at the software layer. Encryption / Decryption: Encryption / decryption refers to the process of protecting or hiding data by using cryptographic techniques. Encryption is the conversion of original data into data processed by a specific algorithm, making it difficult to understand or interpret without authorized access. Decryption is the restoration of encrypted data to its original form. Encryption / decryption is usually used to ensure the confidentiality, integrity, and authenticity of data to prevent unauthorized access and modification. This technology is widely used in the field of information security, such as data transmission, storage, and communication.

[0030] Compression: Without losing or minimizing the loss of data information, data is processed through a specific algorithm to reduce the storage space occupied by the data.

[0031] Decompression: It is the reverse process of compression, that is, restoring the compressed data to the original data. The decompression algorithm needs to decode and restore the compressed data according to the rules and coding methods of the compression algorithm.

[0032] Step 102, perform a first processing operation on the first data block among multiple data blocks, and at the same time perform a second processing operation on the second data block among multiple data blocks.

[0033] In some embodiments, as Figure 9 shown, the user layer and the kernel layer are two different levels in the host-side operating system. After the kernel layer divides the file to be processed into blocks, it can enter the pipeline software stack to perform processing operations on the data blocks. Optionally, the kernel layer can schedule the execution units of the device layer to perform processing operations on the data blocks through a scheduling function.

[0034] In the solution of the present disclosure, the device layer may include a second storage device for storing data and an execution unit for performing processing on data blocks. The second storage device may be, for example, a disk, such as an nvme-ssd disk, a mechanical hard disk, a USB flash drive, etc., and the present disclosure is not limited thereto; the execution unit may be, for example, an accelerator board device, such as an FPGA accelerator board. FPGA is a programmable hardware platform with high flexibility and parallel processing capabilities, and is very suitable for implementing complex data processing tasks.

[0035] Optionally, the execution unit may include a data encryption Intellectual Property Core (IP), a data compression IP, a DEV_DMA-IP, and device memory. At this time, the execution unit can perform read processing, encryption processing, compression processing, and write processing on data blocks; or, the execution unit may include a data decryption IP, a data decompression IP, a DEV_DMA-IP, and device memory. At this time, the execution unit can perform read processing, decryption processing, decompression processing, and write processing on data blocks, where the intellectual property core is an integrated circuit design module with specific functions.

[0036] Optionally, the device layer may include one or more execution units. For example, as Figure 4 shown, it may include a first execution unit (FPGA-1) and a second execution unit (FPGA-2). The first execution unit (compression and encryption accelerator board / FPGA-1 device) may include a data encryption IP, a data compression IP, a DEV_DMA-IP, and device memory. The second execution unit may include a data decryption IP, a data decompression IP, a DEV_DMA-IP, and device memory; or the device layer may include the first execution unit, or include the second execution unit, where the device memory can be used to cache data blocks that need to be processed.

[0037] Optionally, the data encryption IP, the data compression IP, the data decryption IP, and the data decompression IP are logic circuits for data acceleration calculation (the function is to offload the encryption algorithm and the data compression algorithm to hardware), that is, the encryption algorithm and the data compression algorithm are implemented in hardware. Optionally, encryption and decryption refer to the process of protecting or hiding data by using cryptographic techniques. Encryption is to convert the original data into data processed by a specific algorithm, making it difficult to understand or interpret when accessed without authorization. Decryption is to restore the encrypted data to its original form. Encryption and decryption are usually used to ensure the confidentiality, integrity, and authenticity of data to prevent unauthorized access and modification.

[0038] In some embodiments, to address the data transmission latency issue in heterogeneous computing systems and achieve efficient data transmission, the present application adopts the point-to-point communication method in the Peripheral Component Interconnect Express (PCIe) protocol. Specifically, devices are all mounted on the system PCIe bus. In the solution of the present disclosure, direct data transmission is allowed between different devices connected to the PCIe bus without passing through the host CPU and host memory, improving the data transmission efficiency, reducing the CPU load, and significantly enhancing the system performance in compute-intensive tasks. Specifically, in the solution of the present disclosure, point-to-point data transmission can be performed between the second storage device and the host through the PCIe bus (i.e., the DMA transmission process in the point-to-point communication method can be carried out), and point-to-point data transmission can be performed between the second storage device and the execution unit through the PCIe bus (i.e., the DMA transmission process in the point-to-point communication method can be carried out). For example, as Figure 6 shown, point-to-point data transmission can be performed between FPGA-1 and the disk through the bus (Peer-to-Peer is the point-to-point communication method in the PCIe protocol, P2P), and point-to-point data transmission can be performed between FPGA-2 and the disk through the bus. As Figure 7 shown, compared with the non-point-to-point transmission solution, the point-to-point transmission method of the present disclosure does not need to transfer data from the disk to the host-side memory and then transfer the data from the host-side memory to the accelerator board memory. Instead, the point-to-point data transmission method can directly transfer data from the disk to the accelerator board memory, thereby saving a large amount of HOST host-side system memory space. This mechanism also greatly improves the data transmission efficiency and significantly enhances the system performance in compute-intensive tasks (such as data compression / decompression, encryption / decryption).

[0039] In some embodiments, in the solution of the present disclosure, a transparent file system is added on the basis of the original virtual file system and the underlying file system at the kernel layer, and a pipeline software stack is designed in the transparent file system, as Figure 9As shown, the pipelined software stack can process multiple data blocks in a pipelined manner. Specifically, the first data block is read before the second data block, and the first processing operation is an operation that is executed before the second processing operation in the execution order of multiple processing operations. In other words, the first data block can be read first. After reading the first data block, the second data block can be read, realizing pipelined reading of data blocks. During the process of reading the second data block, the first first processing can be performed on the first data block. When the first first processing is performed on the second data block after the second data block is read, the second first processing is simultaneously performed on the first data block, that is, the first processing can be performed on both the first data block and the second data block at the same time, realizing pipelined parallel data processing and improving processing efficiency.

[0040] Optionally, the transparent file system of this solution can include two pipelined software stacks, which are used as read-write interfaces respectively, and then call various hardware computing IPs of two FPGA devices to implement functions of transparent compression and encryption of data and transparent decryption and decompression of data. The meaning of transparent operation is that the user is unaware. For example, when the user initiates a read or write operation, after adding the transparent file system at the kernel layer, the transparent file system will intercept the system call function of the read or write operation, and then trigger the block mechanism and the corresponding pipelined scheduling, and then continue to trigger the FPGA to perform corresponding functions such as data compression and encryption, data decryption and decompression. Therefore, the user only needs to perform regular read and write operations, and can implement the functions of data compression and decryption and decompression without other additional operations.

[0041] Through this application, in response to a first operation instruction, the file to be processed is divided to obtain multiple data blocks corresponding to the file to be processed; a first processing operation is performed on the first data block among the multiple data blocks, and a second processing operation is simultaneously performed on the second data block among the multiple data blocks, where the first data block is read before the second data block, and the first processing operation is an operation that is executed before the second processing operation in the execution order of multiple processing operations. The multiple processing operations include at least two of reading processing, compression processing, encryption processing, and writing processing, or the multiple processing operations include at least two of reading processing, decryption processing, decompression processing, and writing processing. It can realize software-hardware collaborative parallel data processing and can simultaneously meet the requirements of efficient data transmission, efficient transparent compression and encryption, and efficient transparent decryption and decompression.

[0042] Figure 2 Further, a flowchart of another data processing method proposed by the present disclosure is shown. Based on Figure 1 the embodiments shown, the data processing method of the present disclosure may further include the following steps.

[0043] Step 201: Determine multiple task blocks corresponding to multiple data blocks, where each task block includes information related to the data block corresponding to the task block.

[0044] In some embodiments, multiple task blocks corresponding to multiple data blocks can be determined. Each task block includes information related to the data block corresponding to the task block. For example, the task block information describes the size of the data block to be processed, file information where the data file is located (file size, file location, file offset), data block identifier (id), task block identifier (id), and other information.

[0045] Step 202: Determine the identifiers of multiple task blocks.

[0046] In some embodiments, the identifiers of multiple task blocks can be determined. Optionally, the identifier of a task block can be determined based on the identifier of the data block. Multiple data blocks can be sorted according to their positions in the file, and based on the sorting result, the identifier of the task block can be determined. Or the identifier of the task block can be determined based on the data volume contained in the data block, or based on the priority of the data contained in the data block, and so on.

[0047] Step 203: Determine the reading order of multiple data blocks according to the identifiers of multiple task blocks.

[0048] In some embodiments, the reading order of multiple blocks can be determined according to the identifiers of multiple task blocks. For example, multiple data blocks can be read in ascending order of the value of the id, or multiple data blocks can be read in descending order of the value of the id, and so on.

[0049] Optionally, the reading order of multiple data blocks is the order of reading and processing multiple data blocks. The reading order of multiple data blocks can be the order in which multiple data blocks enter the pipeline software stack, or the order in which processing of multiple data blocks starts.

[0050] In summary, in the above embodiments of the present application, multiple task blocks corresponding to multiple data blocks can be determined, and the processing order of multiple data blocks can be determined according to the task block identifier. It is possible to implement pipeline processing of data blocks as needed, which can improve the flexibility and efficiency of data processing.

[0051] Figure 3 Further, a flowchart of another data processing method proposed by the present disclosure is shown. Based on Figure 1 and Figure 2 the embodiments shown, step 102 is further explained and can include the following steps.

[0052] Step 301: Input the first task block into the pipeline software stack in the transparent file system.

[0053] In some embodiments, the first task block may be input into the pipeline software stack in the transparent file system first, that is, the first processing may be performed on the first data block corresponding to the first task block first.

[0054] Step 302, the first-level pipeline in the pipeline software stack responds to the first task block and performs a read process on the first data block.

[0055] In some embodiments, the first-level pipeline in the pipeline software stack responding to the first task block and performing a read process on the first data block includes: determining that the first operation instruction is a data write operation instruction, and the execution order of multiple processing operations is read process, compression process, encryption process, and write process; the first-level pipeline uses the first scheduling function to schedule the execution unit, and obtains the first data block stored in the first storage device through the write operation of the execution unit; writes the first data block into the first device memory of the execution unit.

[0056] In other words, when the first operation instruction is a data write operation instruction, a read process, a compression process, an encryption process, and a write process may be performed on multiple data blocks. At this time, the first scheduling function may instruct the execution unit to obtain the first data block from the first storage device, and the first scheduling function may be used to schedule the execution unit to initiate a write operation to the host side. For example, the FPGA may initiate a DEV_DMA write operation to the host side to obtain the first data block stored in the first storage device, where DEV_DMA represents the DMA of the FPGA, not the DEV_DMA of the disk. The first storage device may be the storage space on the host side. After the execution unit obtains the first data block, the first data block may be temporarily stored in the device memory of the execution unit to facilitate subsequent compression process, encryption process, and write process. Optionally, the first scheduling function may include the first task block, and the execution unit may obtain the first data block from the first storage device according to the first task block.

[0057] In some embodiments, the first-level pipeline in the pipeline software stack responding to the first task block and performing a read process on the first data block includes: determining that the first operation instruction is a data read operation instruction, and the execution order of multiple processing operations is read process, decryption process, decompression process, and write process; the first-level pipeline uses the fifth scheduling function to schedule the execution unit to obtain the first data block stored in the second storage device through the bus; writes the first data block into the first device memory of the execution unit.

[0058] In other words, when the first operation instruction is a data read operation instruction, read processing, decryption processing, decompression processing, and writing processing can be performed on multiple data blocks. At this time, the first scheduling function can instruct the execution unit to obtain the first data block from the second storage device. The first scheduling function can be used to schedule the execution unit to obtain the first data block stored in the second storage device point-to-point through the PCIe bus. After obtaining the first data block, the execution unit can temporarily store the first data block in the device memory of the execution unit to facilitate subsequent decryption processing, decompression processing, and writing processing. Optionally, the first scheduling function may include a first task block, and the execution unit can obtain the first data block from the second storage device according to the first task block.

[0059] Step 303, when the read processing of the first data block is completed in the first-level pipeline, input the second task block into the pipeline software stack.

[0060] In some embodiments, the first-level pipeline can perform a first process on the first data block. When the read processing of the first data block is completed in the first-level pipeline, the second task block can be input into the pipeline software stack, that is, the first process on the second data block can be started, and multiple data blocks can be processed in a pipeline manner. That is, after the first scheduling function completes the read processing of the first data block, the first scheduling function can be used to continue the read processing of the second data block, and so on. Then, the read processing of the third data block can be performed until the read processing of all multiple data blocks is completed.

[0061] Step 304, in response to the second task block, the second-level pipeline in the pipeline software stack performs a second processing operation on the second data block, and at the same time, the first-level pipeline performs a first processing operation on the first data block.

[0062] In some embodiments, in response to the second task block, the second-level pipeline in the pipeline software stack performs a second processing operation on the second data block, that is, the second-level pipeline can perform a first process on the second data block. Optionally, the first scheduling function can be used to perform the read processing on the second data block, that is, the method of performing the read processing on multiple data blocks is the same. The solution of the present disclosure can use the scheduling function to schedule the hardware in the device layer to perform the first process in the kernel layer, which can realize the collaborative processing between the software layer and the hardware layer, reduce the consumption of software resources, and achieve efficient data transmission and transparent operation.

[0063] In some embodiments, the second pipeline stage in the pipeline software stack executes a second processing operation on a second data block in response to a second task block, while the first pipeline stage executes a first processing operation on a first data block, including: at a first time, the first pipeline stage uses a second scheduling function to schedule an execution unit to perform a compression process on the first data block, while the second pipeline stage uses a first scheduling function to schedule the execution unit to obtain the second data block stored in the first storage device through a write operation of the execution unit and store the second data block in the first device memory; at a second time, the first pipeline stage uses a third scheduling function to schedule the execution unit to perform an encryption process on the first data block, while the second pipeline stage uses the second scheduling function to schedule the execution unit to perform a compression process on the second data block; at a third time, the first pipeline stage uses a fourth scheduling function to schedule the execution unit to perform a write process on the first data block, while the second pipeline stage uses the third scheduling function to schedule the execution unit to perform an encryption process on the second data block. In some embodiments, the first pipeline stage using the fourth scheduling function to schedule the execution unit to perform a write process on the first data block includes: the first pipeline stage uses the fourth scheduling function to schedule the execution unit to send the first data block after compression and encryption processes to a second storage device through a bus.

[0064] In other words, when the first operation instruction is a data write operation, after the reading of the first data block is completed and the second task block is input into the pipeline software stack, at the first time, the first scheduling function and the second scheduling function can be used simultaneously to process the second data block and the first data block synchronously. At this time, the second scheduling function schedules the compression IP of the FPGA to perform compression on the first data block, while the first scheduling function schedules the FPGA to obtain the second data block from the first storage device using the DEV-DMA write operation of the FPGA; thereafter, at the second time, the second scheduling function and the third scheduling function can be used simultaneously to process the second data block and the first data block synchronously. At this time, the second scheduling function schedules the compression IP of the FPGA to perform compression on the second data block, while the third scheduling function schedules the encryption IP of the FPGA to perform an encryption process on the first data block; thereafter, at the third time, the third scheduling function and the fourth scheduling function can be used simultaneously to process the second data block and the first data block synchronously. At this time, the third scheduling function schedules the encryption IP of the FPGA to perform encryption on the second data block, while the fourth scheduling function schedules the FPGA to send the first data block after compression and encryption processes to the second storage device through a bus. The second storage device can be a disk for storing user-written data; thereafter, at the fourth time, the processing of the first data block has been completed, and at the fourth time, the fourth scheduling function can be used to schedule the FPGA to send the second data block after compression and encryption processes to the second storage device through a bus.

[0065] When the number of data blocks to be processed is multiple, the pipelining processing method of the above first data block and second data block can be referred to. For example, as Figure 9 shown, first, within the first time, task block 0 is input into the pipeline software stack. Within the first time, the DEV-DMA-IP of FPGA-1 can be scheduled to obtain data block 0 corresponding to task block 0 from the first storage device and store data block 0 in the device memory; afterwards, within the second time, task block 1 can be input into the pipeline software stack. Within the second time, the data compression IP of FPGA-1 can be scheduled to retrieve data block 0 from the device memory and compress data block 0. At the same time, the DEV-DMA-IP obtains data block 1 corresponding to task block 1 and stores data block 1 in the device memory. After data block 0 is compressed, the compressed data block can be stored in the device memory; within the third time, task block 2 can be input into the pipeline software stack. At the same time, the DEV-DMA-IP of FPGA-1 can be scheduled to obtain data block 2 corresponding to task block 2 and store data block 2 in the device memory. At the same time, the data encryption IP of FPGA-1 retrieves the compressed data block 0 from the device memory and encrypts data block 0. At the same time, the data compression IP of FPGA-1 retrieves data block 1 from the device memory and compresses data block 1. Similarly, the encrypted data block 0 is saved in the device memory, and the compressed data block 1 is saved in the device memory; within the fourth time, task block 3 can be input into the pipeline software stack. At the same time, the DEV-DMA-IP of FPGA-1 can be scheduled to obtain data block 3 corresponding to task block 3 and store data block 3 in the device memory. At the same time, FPGA-1 sends the compressed and encrypted data block 0 in the device memory to the disk for storage through the PCIe bus. At the same time, the data encryption IP of FPGA-1 retrieves the compressed data block 1 from the device memory and encrypts data block 1. At the same time, the data compression IP of FPGA-1 retrieves data block 2 from the device memory and compresses data block 2. Similarly, the encrypted data block 1 is saved in the device memory, and the compressed data block 2 is saved in the device memory. Thus, data block 0 is processed. The remaining data blocks repeat the above operations until all multiple data blocks are processed.

[0066] In some embodiments, the second pipeline stage in the pipeline software stack, in response to a second task block, performs a second processing operation on a second data block, while the first pipeline stage performs a first processing operation on a first data block, including: at a first time, the first pipeline stage uses a sixth scheduling function to schedule an execution unit to perform a decryption process on the first data block, while the second pipeline stage uses a fifth scheduling function to schedule the execution unit to obtain the second data block stored in a second storage device through a bus and store the second data block in a first device memory; at a second time, the first pipeline stage uses a seventh scheduling function to schedule the execution unit to perform a decompression process on the first data block, while the second pipeline stage uses a sixth scheduling function to schedule the execution unit to perform a decryption process on the second data block; at a third time, the first pipeline stage uses an eighth scheduling function to schedule the execution unit to perform a write process on the first data block, while the second pipeline stage uses a seventh scheduling function to schedule the execution unit to perform a decompression process on the second data block. In some embodiments, the first pipeline stage using an eighth scheduling function to schedule an execution unit to perform a write process on the first data block includes: the first pipeline stage uses an eighth scheduling function to schedule the execution unit to send the first data block that has undergone decryption and decompression processes to a first cache space for caching through a read operation of the execution unit.

[0067] In other words, when the first operation instruction is a data read operation, after the reading of the first data block is completed and the second task block is input into the pipeline software stack, at the first time, the fifth scheduling function and the sixth scheduling function can be used simultaneously to process the second data block and the first data block synchronously. At this time, the sixth scheduling function schedules the data decryption IP of the FPGA to perform decryption on the first data block, while the fifth scheduling function schedules the FPGA to obtain the second data block from the first storage device through the bus; thereafter, at the second time, the sixth scheduling function and the seventh scheduling function can be used simultaneously to process the second data block and the first data block synchronously. At this time, the sixth scheduling function schedules the data decryption IP of the FPGA to perform decryption on the second data block, while the seventh scheduling function schedules the data decompression IP of the FPGA to perform a decompression process on the first data block; thereafter, at the third time, the seventh scheduling function and the eighth scheduling function can be used simultaneously to process the second data block and the first data block synchronously. At this time, the seventh scheduling function schedules the data decompression IP of the FPGA to perform decompression on the second data block, while the eighth scheduling function schedules the FPGA to cache the first data block that has undergone decryption and decompression processes into the first cache space through the DEV-DMA read operation of the FPGA. The first cache space can be a kernel-layer user cache for the decrypted and compressed data block; thereafter, at the fourth time, the processing of the first data block has been completed, and at the fourth time, the eighth scheduling function can be used to schedule the FPGA to cache the second data block that has undergone decryption and decompression processes into the first cache space through the DEV-DMA read operation of the FPGA.

[0068] When the number of data blocks to be processed is multiple, the pipeline processing method of the first data block and the second data block can be referred to above, for example, Figure 11 As shown, first, at the first time, task block 0 is input into the pipeline software stack, and at the first time, FPGA-2 can be scheduled to obtain data block 0 corresponding to task block 0 from the second storage device through the bus, and data block 0 is stored in the device memory; then, at the second time, task block 1 can be input into the pipeline software stack, and at the second time, the data decryption IP of FPGA-2 can be scheduled to take out data block 0 from the device memory and decrypt data block 0, and at the same time, FPGA-2 obtains data block 1 corresponding to task block 1 from the second storage device through the bus, and stores data block 1 in the device memory, wherein after decrypting data block 0, the decrypted data block can be stored in the device memory; at the third time, task block 2 can be input into the pipeline software stack, and FPGA-2 can be scheduled to obtain data block 2 through the bus, and data block 2 is stored in the device memory, and at the same time, the data decompression IP of FPGA-2 takes out the decrypted data block 0 from the device memory, encrypts data block 0, and at the same time, FPGA -2's data decryption IP takes out data block 1 from the device memory and decrypts data block 1. Similarly, the decompressed data block 0 is saved in the device memory, and the decrypted data block 1 is saved in the device memory; at the fourth time, task block 3 can be input into the pipeline software stack, and FPGA-2 can be scheduled to obtain data block 3 through the bus, and store data block 3 in the device memory. At the same time, DEV-DMA-IP of FPGA-2 sends the decrypted and decompressed data block 0 in the device memory to the first cache space for caching. At the same time, the data decompression IP of FPGA-2 takes out the decrypted data block 1 from the device memory and decompresses data block 1. At the same time, the data decryption IP of FPGA-2 takes out data block 2 from the device memory and decrypts data block 2. Similarly, the decompressed data block 1 is saved in the device memory, and the decrypted data block 2 is saved in the device memory. At this point, data block 0 is processed, and the above operation is repeated for the remaining data blocks until multiple data blocks are all processed.

[0069] In some embodiments, the method further includes: acquiring the first data block that has been decrypted and decompressed from the first cache space; and displaying the first data block that has been decrypted and decompressed.

[0070] Optionally, the data block after decryption and decompression can be sent to the first cache space for caching. When the same data block needs to be read repeatedly later, it can be directly obtained from the cache space without decrypting and decompressing again. That is, the purpose of caching is to reduce the repeated decryption and decompression operations on the same data file. After performing a decryption and decompression operation process on a data block, the decompressed data result is stored in the cache. When the same action is performed next time, the data can be directly obtained from the cache, and the decryption and decompression process of the pipeline will not be executed. The user layer of the host can obtain the cached data block from the first cache space, which can improve the acquisition efficiency and complete the user's data read operation.

[0071] In some embodiments, at least one of the first scheduling function, the second scheduling function, the third scheduling function, the fourth scheduling function, the fifth scheduling function, the sixth scheduling function, the seventh scheduling function, and the eighth scheduling function includes the first task block or the second task block. In other words, when using the scheduling function to schedule the FPGA to process the data block, the scheduling function can include the information of the data block processed by the execution unit, so that the execution unit can determine the data block to be processed according to the information related to the data block.

[0072] In summary, in this application, a transparent file system is embedded in the kernel layer, and a task block mechanism and two four-level pipeline software stacks for reading and writing are designed inside it. The point-to-point data transfer, DEV_DMA data read and write transfer, data transparent compression and encryption, and data transparent decompression are called successively. The efficient cooperative parallel call of software and hardware is realized, which greatly improves the computing throughput rate of the heterogeneous acceleration system. And by using the memory address mapping technology, the device memory is inserted, managed, and allocated in the kernel layer of the Linux operating system. In this way, the files on the disk can be transferred point-to-point to the accelerator device through the kernel layer interface function. This data transfer method significantly reduces the data transfer latency and improves the overall performance of the system.

[0073] Based on Figure 1 、 Figure 2 and Figure 3 The embodiments shown, the above method will be further explained by a specific example below. This example provides a method for implementing an efficient data transfer, data storage, data encryption and decryption heterogeneous acceleration computing software scheduling system. It is introduced in the following aspects: 1. The principle of PCIe point-to-point transmission. 2. The schematic diagram of the data processing process of this design scheme. 3. The design scheme of the four-level pipeline software stack for point-to-point data file transmission, data compression and encryption, and data decryption and decompression. 4. The block diagram of the actual scheme composition.

[0074] 1. The principle of point-to-point data transmission In a traditional heterogeneous computing system, the disk storage device needs to perform multiple DMA transfers to transfer data to the accelerator board device. First, the disk data is transferred to the host memory through the disk's DMA, and then the data is directly transferred from the host side to the accelerator board through the device's DMA. This method has a relatively large transmission delay.

[0075] To solve the data transmission delay problem in heterogeneous computing systems and achieve efficient data transmission, the present invention adopts a peer-to-peer data transmission method. Peer-to-Peer is a peer-to-peer communication method in the PCIe protocol.

[0076] The PCIe protocol supports multiple data transmission modes, including peer-to-peer and direct memory access (DMA), etc. When a device has the ability to act as a bus master, it can initiate data transmission independently and send data directly to another PCIe device without relying heavily on the central processing unit (CPU) and other system resources. This direct data transmission method between devices is called a peer-to-peer transaction. Since peer-to-peer transactions only occur between devices within the system, it can significantly reduce the data transmission delay and improve the overall system performance.

[0077] Figure 4 This is a schematic diagram of the topology structure of a PCIe bus protocol provided herein. The peer-to-peer - DMA (abbreviated as peer-to-peer) technology allows direct data transmission between different devices connected to the PCIe bus without transferring data from the disk to the host-side memory and then from the host-side memory to the accelerator board memory. By using the peer-to-peer data transmission method, data can be directly transferred from the disk to the accelerator board memory, thereby saving a large amount of HOST host-side system memory space. This mechanism also greatly improves the data transmission efficiency and significantly enhances the system performance in computationally intensive tasks (such as data compression / decompression, encryption / decryption).

[0078] As Figure 5 shown, it is a schematic diagram of a traditional DMA data transmission scheme. As Figure 6 shown, it is a schematic diagram of a peer-to-peer DMA data transmission scheme provided by this solution. In heterogeneous computing scenarios, such as data compression / decompression, encryption / decryption, etc., the heterogeneous computing of data itself is particularly time-consuming. For small amounts of data, the impact is relatively small, but for larger data, if the data transmission path is particularly long and multiple copies are made, the data delay will be extremely large, which will ultimately affect the performance of the entire system. Figure 6 This is a schematic diagram when the file performs peer-to-peer data transmission in a data heterogeneous computing system. Figure 7 This is a schematic diagram of realizing peer-to-peer data stream transmission between the disk and the FPGA.

[0079] 2. Schematic Diagram of the Data Processing Process of This Design Scheme Figure 8 It is a schematic diagram of the single-step execution steps when the data encryption and decryption algorithm and the query algorithm are offloaded to the FPGA hardware. It is explained in two parts, left and right, and each part has four steps to complete. Figure 8 On the left side, the device FPGA-1 executes the data compression and encryption tasks. The device FPGA-1 on the left side will transfer the compressed and encrypted data results through point-to-point data transmission to the disk (disk DMA write operation); on the right side, the device FPGA-2 executes the data decryption and decompression tasks. First, the data is transferred from the disk to the FPGA-2 side through point-to-point data transmission (disk DMA read operation), and then the data decryption and decompression steps are carried out.

[0080] (1)Data Compression and Encryption (Steps ①②③④ Executed by the FPGA-1 on the Left Side Device) First, a write operation is executed at the user layer. FPGA-1 initiates a DEV_DMA write operation to the host side (DEV_DMA represents the DMA of the FPGA. Note: The DEV_DMA here is different from the DMA of the disk), writes the data into the memory of the FPGA_1 device, and then sequentially executes the data compression and encryption operations on FPGA-1. Finally, FPGA-1 will transfer the compressed and encrypted data results through point-to-point data transmission to the disk (disk DMA write operation). Therefore, the user's write operation will trigger the data compression and encryption operation.

[0081] (2)Data Decryption and Decompression (Steps ①②③④ Executed by the FPGA-2 on the Right Side Device) First, a read operation is executed at the user layer, and then the data is transferred from the disk to the memory on the FPGA-2 side through point-to-point data transmission (disk DMA read operation). Then, the data decryption and decompression operations are sequentially executed on FPGA-2. Finally, FPGA-2 initiates a DEV_DMA read operation to the host side (DEV_DMA represents the DMA of the FPGA. Note: The DEV_DMA here is different from the DMA of the disk), and finally returns the result to the user layer, and the read operation ends. Therefore, the user's read operation will trigger the data decryption and decompression operation.

[0082] 3. Design Scheme of the Four-Level Pipeline Software Stack for Data Compression and Encryption and Data Decryption and Decompression (1)Data Compression and Encryption Pipeline Software Stack As Figure 9 shown, it is the design scheme of the compression and encryption pipeline software stack provided by this scheme. Figure 9It is described in three levels: the user level, the kernel level, and the device level. The user level and the kernel level are two different levels in the host-side operating system. The device level includes disks and FPGA accelerator boards, and these devices are all mounted on the PCIe bus of the entire system. The compression and encryption accelerator board in the device level is the FPGA-1 device, and inside the FPGA-1 device, there are data encryption IP, data compression IP, device memory, etc. Among them, the data encryption IP and the data compression IP are logic circuits for data acceleration calculation (the function is to offload the encryption algorithm and the data compression algorithm to hardware), that is, the encryption algorithm and the data compression algorithm are implemented in hardware.

[0083] The pipelined software stack is to improve the parallel scheduling ability of the system software to adapt to the data processing throughput rate of the FPGA hardware parallel computing. If a serial system call method (such as Figure 10 ) is adopted, it will consume more time, resulting in a large delay and reducing the overall system performance. For example, in this solution, a four-level pipeline is adopted, which respectively performs DEV_DMA write data transfer, data compression, data encryption, and point-to-point data transfer. Each level of the pipeline has to execute the scheduling of these four steps, and each step represents a relevant scheduling function or algorithm. For example, in this solution, corresponding functions are designed respectively to schedule DEV_DMA write data transfer, data compression, data encryption, and point-to-point data transfer.

[0084] If a serial scheduling method is adopted, it will take 16 time periods to complete 4 tasks, while only 7 time periods are consumed if a four-level pipeline is adopted. The delay is greatly reduced.

[0085] After the user initiates a write operation, the write operation enters the kernel layer through a system call. After passing through the write operating system call in the kernel layer, the data file to be written is task-blocked. Subsequently, the data file to be written is evenly divided into several task blocks (task block 0, task block 1, … task block n-1). The task block information describes the size of the data block to be processed, the file information of the data file (file size, file location, file offset), the data block id, the task block id, etc. Subsequently, when the first stage of the four-stage pipeline receives task block 0, it starts to execute DEV_DMA write data transfer, data compression, data encryption, and point-to-point data transfer in sequence. According to the characteristics of the pipeline, the DEV_DMA write data transfer in the second stage of the pipeline and the data compression in the first stage of the pipeline occur at the same time. The second stage of the pipeline starts the DEV_DMA write data transfer... and so on. The task block ids corresponding to each stage of the four-stage pipeline are 4i, 4i+1, 4i+2, 4i+3 respectively. i (starting from 0 and incrementing by 1 successively) represents the number of cycles of task distribution. After each distribution of tasks to the fourth stage of the pipeline, the value of i will increase by 1. When each stage of the pipeline is completed, the result will be transferred from the device memory on the FPGA-1 side to the disk.

[0086] (2) Data Decryption and Decompression Pipeline Software Stack As Figure 11 shown, it is the design of the four-stage pipeline software stack for data decryption and decompression, which is also described in three layers: the user layer, the kernel layer, and the device layer. Among them, the user layer and the kernel layer are two different layers in the host-side operating system. The device layer includes a disk and an FPGA accelerator board, and these devices are all mounted on the PCIe bus of the entire system. The decryption and decompression accelerator board in the device layer is the FPGA-2 device, and inside the FPGA-2 device, there are a data decryption IP, a data decompression IP, and device memory. Among them, the data decryption IP and the data decompression IP are logic circuits for performing data acceleration calculations (the function is to offload the decryption algorithm and the data decompression algorithm to hardware), that is, the decryption algorithm and the data decompression algorithm are implemented in hardware.

[0087] When the user initiates a read operation, the read operation enters the kernel layer through a system call. After the kernel layer reads the operating system call, it divides the data file to be read into task blocks, and then divides the data file to be read into several task blocks (task block 0, task block 1, ... task block n-1). The task block information describes the size of the data block to be processed, the file information of the data file (file size, file location, file offset), data block id, task block id and other information. Then, when the first-level pipeline of the four-level pipeline receives task block 0, it starts to execute point-to-point data transmission, data decryption, data decompression, and DEV_DMA write data transmission in sequence. According to the characteristics of the pipeline, the pipeline reading process is similar to the writing performed by FPGA-1 above, and decryption and decompression are the inverse operations of compression and encryption. The task block ids corresponding to each level of the four-level pipeline are 4i, 4i+1, 4i+2, and 4i+3, respectively. i (i starts from 0 and increases by 1) represents the number of cycles of task distribution. Each time the fourth-level pipeline distributes tasks, the value of i increases by 1. When each level of the pipeline is completed, the result will be moved from the device memory on the FPGA-2 side to the host side cache (the purpose of the cache is to reduce the repeated decryption and decompression operations on the same data file. After performing a decryption and decompression operation on a data block, the decompressed data result is stored in the cache. The next time the same action is performed, the data can be directly obtained from the cache without executing the decryption and decompression process of the pipeline). Finally, the host will read the final result from the cache, and the read operation ends.

[0088] 4. Structural diagram of the actual solution like Figure 12 As shown in the figure, the structure of the compression encryption, decryption and decompression heterogeneous computing data transmission solution is shown. In this solution, the disk and FPGA accelerator board are connected to the server host through the PCIe bus. The FPGA accelerator board contains multiple key modules, such as DEV_DMA-IP, encryption IP, decryption IP, compression IP, decompression IP, memory controller and device memory. Among them: DEV_DMA-IP: responsible for DMA data transfer on the PCIe bus.

[0089] Encryption IP: responsible for data encryption operations.

[0090] Decryption IP: responsible for data decryption operations.

[0091] Compression IP: responsible for implementing the compression operation of data files.

[0092] Decompression IP: responsible for decompressing data files The server host runs the Linux operating system, which is divided into a kernel layer and an application layer. In the application layer, users can initiate read and write operations on data files, including data compression and encryption, and data decryption and decompression. In the kernel layer, the memory management mechanism is used to manage and allocate memory resources, and then operations such as point-to-point data transmission, data encryption and decryption, and data compression and decompression are issued, and the software stack workflow is called. The FPGA driver in the kernel layer is responsible for managing various devices.

[0093] In this solution, the focus of the kernel layer is to implement the data chunking mechanism, the four-stage pipeline software stack, and the FPGA driver. These components work together to ensure the efficient transfer of data between the disk and the FPGA, and to complete operations such as compression and encryption, decryption and decompression required by users.

[0094] It should be particularly emphasized that in this solution, a transparent file system is added on the basis of the original virtual file system and the underlying file system in the kernel, and two four-stage pipelines are designed in the transparent file system, which are used as read and write interfaces respectively, and then various hardware computing IPs of two FPGA devices are called to implement the functions of transparent compression and encryption of data and transparent decryption and decompression of data. The meaning of transparent operation is that it is imperceptible to users. For example, when a user initiates a read or write operation, after adding the transparent file system in the kernel layer, the transparent file system will intercept the system call function of the read or write operation, and then trigger the chunking mechanism and the corresponding pipeline scheduling, and then continue to trigger the FPGA to perform corresponding data compression and encryption, data decryption and decompression and other functions. Therefore, users only need to perform regular read and write operations, and can realize the functions of data compression and decryption and decompression without other additional operations. This is the transparent operation.

[0095] As Figure 13 shown, it is the specific implementation process of heterogeneous computing - transparent compression and encryption of data, including: 1) The user initiates a data write operation.

[0096] 2) The operating system kernel receives the data write system call.

[0097] 3) The chunking mechanism in the kernel layer divides the data file to be written into n equal task chunks, and each task chunk contains information such as the size of the data block to be processed, the file information where the data is located (file size, file location, file offset), data block id, task block id, etc.

[0098] 4) The for loop mechanism is used to allocate task chunk information to the four-stage pipeline software stack.

[0099] 5) The four-stage pipeline software stack starts to execute, and the pipeline software stack can improve the software and hardware collaborative data processing ability.

[0100] 5-1) Start the DEV_DMA write operation to transfer data into FPGA-1.

[0101] 5-2) FPGA-1 performs data compression operation; 5-3) FPGA-1 performs data encryption operation; 5-4) Start the disk DMA write operation, perform point-to-point data transfer, and transfer the data in the FPGA-1 memory to the disk.

[0102] 6) The transparent compression and encryption operation of the data is completed, and the entire write operation process ends here.

[0103] As Figure 14 shown, it is the specific implementation process of heterogeneous computing - transparent decryption and decompression of data, including: 1) The user initiates a data read operation.

[0104] 2) The operating system kernel receives the data read system call.

[0105] 3) The kernel layer chunking mechanism divides the data file to be read into n equal task chunks. Each task chunk contains information such as the size of the data chunk to be processed, file information of the data (file size, file location, file offset), data chunk id, task chunk id, etc.

[0106] 4) Use the for loop mechanism to allocate task chunk information to the four-level pipeline software stack.

[0107] 5) The four-level pipeline software stack starts to execute. The pipeline software stack can improve the software and hardware collaborative data processing ability.

[0108] 5-1) Start the disk DMA read operation, perform point-to-point data transfer, and transfer the data in the disk to the FPGA-2 memory; 5-2) FPGA-2 performs data compression operation; 5-3) FPGA-2 performs data encryption operation; 5-4) Start the DEV_DMA read operation to transfer data into the host cache.

[0109] 6) The user layer reads the decompressed result from the cache. The transparent decryption and decompression operation of the data is completed, and the entire read operation process ends here.

[0110] In summary, this example proposes a solution that simultaneously achieves efficient data transmission, transparent data compression and encryption, and transparent data decryption and decompression in a heterogeneous computing environment. This solution uses memory address mapping technology to insert, manage, and allocate device memory at the kernel layer of the Linux operating system. In this way, files on the disk can be transferred point-to-point to the accelerator device through the kernel layer interface function. This data transmission method significantly reduces data transmission latency and improves the overall performance of the system. By embedding a transparent file system at the kernel layer and designing a task chunking mechanism and two four-stage pipelined software stacks for reading and writing inside it, point-to-point data transmission, DEV_DMA data read and write transmission, data transparent compression and encryption, and data transparent decompression are successively called. It realizes efficient cooperative parallel calls of software and hardware, greatly improving the computing throughput of the heterogeneous acceleration system.

[0111] The solution of this application: 1. Through memory address mapping technology, map device memory to the host memory address space, enabling point-to-point efficient data transmission, reducing the number of data copies, and reducing data transmission latency.

[0112] 2. Designed DEV_DMA write data transmission, data encryption, data compression, and point-to-point data transmission, implemented a four-stage pipelined software stack for write operations, and thus achieved efficient software and hardware system scheduling.

[0113] 3. Implement the task chunking mechanism in the heterogeneous acceleration computing solution.

[0114] 4. Designed point-to-point data transmission, data decryption, data decompression, and DEV_DMA read data transmission, implemented a four-stage pipelined software stack for read operations, and thus achieved efficient software and hardware system scheduling.

[0115] 5. The data compression and encryption operations and data decryption and decompression operations in this solution are transparent operations. Users only need to perform data writing operations, and the system will automatically implement data compression and encryption operations and decryption and decompression operations without users performing additional operations.

[0116] 6. The API interfaces of point-to-point data transmission, data compression and encryption operations, and data decryption and decompression operations functions are all implemented in the kernel state, reducing system calls and process switches between the kernel state and the user state.

[0117] 7. Achieve compatibility with storage disk types. The disks used in this application do not distinguish disk types, and the disks can be nvme-ssd disks, mechanical hard disks, USB flash drives, etc.

[0118] In summary, the implementation method of an efficient heterogeneous acceleration computing software scheduling system for data transmission, data storage, data encryption and decryption designed in this solution not only overcomes the deficiencies of the prior art, but also significantly improves the performance and security of data processing, effectively saves storage space, greatly reduces the storage resource occupancy rate of the data center, and has broad application prospects.

[0119] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0120] An embodiment of the present application also provides a data processing device 1500, Figure 15 which is a schematic structural diagram of a data processing device provided by an embodiment of the present disclosure. As Figure 15 shown, it includes: A first processing unit 1510, configured to divide a file to be processed in response to a first operation instruction, to obtain a plurality of data blocks corresponding to the file to be processed; A second processing unit 1520, configured to perform a first processing operation on a first data block among the plurality of data blocks, and at the same time perform a second processing operation on a second data block among the plurality of data blocks, where the first data block is read before the second data block is read, and the first processing operation is an operation that is executed before the second processing operation according to the execution order of a plurality of processing operations, and the plurality of processing operations include at least two of a read processing, a compression processing, an encryption processing, and a write processing, or the plurality of processing operations include at least two of a read processing, a decryption processing, a decompression processing, and a write processing.

[0121] Further, in a possible implementation manner of an embodiment of the present disclosure, the data processing device further includes a third processing unit, configured to: determine a plurality of task blocks corresponding to the plurality of data blocks, where each task block includes information related to the data block corresponding to the task block; determine identifiers of the plurality of task blocks; and determine a read order of the plurality of data blocks according to the identifiers of the plurality of task blocks.

[0122] Further, in a possible implementation manner of an embodiment of the present disclosure, the third processing unit is configured to: determine that the first operation instruction is a data write operation instruction when the plurality of processing operations include at least two of a read processing, a compression processing, an encryption processing, and a write processing; and determine that the first operation instruction is a data read operation instruction when the plurality of processing operations include at least two of a read processing, a decryption processing, a decompression processing, and a write processing.

[0123] Further, in a possible implementation manner of the embodiment of the present disclosure, the second processing unit 1520 is configured to: input the first task block into the pipeline software stack in the transparent file system; the first - stage pipeline in the pipeline software stack responds to the first task block and performs a read process on the first data block; when the first - stage pipeline finishes performing the read process on the first data block, input the second task block into the pipeline software stack; the second - stage pipeline in the pipeline software stack responds to the second task block and performs a second processing operation on the second data block, and at the same time, the first - stage pipeline performs a first processing operation on the first data block.

[0124] Further, in a possible implementation manner of the embodiment of the present disclosure, the second processing unit 1520 is configured to: determine that the first operation instruction is a data write operation instruction, and the execution order of multiple processing operations is read process, compression process, encryption process, and write process; the first - stage pipeline uses the first scheduling function to schedule the execution unit to obtain the first data block stored in the first storage device through the write operation of the execution unit; write the first data block into the first device memory of the execution unit.

[0125] Further, in a possible implementation manner of the embodiment of the present disclosure, the second processing unit 1520 is configured to: at the first time, the first - stage pipeline uses the second scheduling function to schedule the execution unit to perform a compression process on the first data block, and at the same time, the second - stage pipeline uses the first scheduling function to schedule the execution unit to obtain the second data block stored in the first storage device through the write operation of the execution unit and store the second data block in the first device memory; at the second time, the first - stage pipeline uses the third scheduling function to schedule the execution unit to perform an encryption process on the first data block, and at the same time, the second - stage pipeline uses the second scheduling function to schedule the execution unit to perform a compression process on the second data block; at the third time, the first - stage pipeline uses the fourth scheduling function to schedule the execution unit to perform a write process on the first data block, and at the same time, the second - stage pipeline uses the third scheduling function to schedule the execution unit to perform an encryption process on the second data block.

[0126] Further, in a possible implementation manner of the embodiment of the present disclosure, the second processing unit 1520 is configured to: the first - stage pipeline uses the fourth scheduling function to schedule the execution unit to send the first data block after compression processing and encryption processing to the second storage device through the bus.

[0127] Further, in a possible implementation manner of the embodiment of the present disclosure, the second processing unit 1520 is configured to: determine that the first operation instruction is a data read operation instruction, and the execution order of multiple processing operations is read process, decryption process, decompression process, and write process; the first - stage pipeline uses the fifth scheduling function to schedule the execution unit to obtain the first data block stored in the second storage device through the bus; write the first data block into the first device memory of the execution unit.

[0128] Further, in a possible implementation manner of the embodiment of the present disclosure, the second processing unit 1520 is configured to: at a first time, the first-level pipeline uses a sixth scheduling function to schedule an execution unit to perform decryption processing on a first data block, and at the same time, the second-level pipeline uses a fifth scheduling function to schedule the execution unit to obtain a second data block stored in a second storage device through a bus and store the second data block in a first device memory; at a second time, the first-level pipeline uses a seventh scheduling function to schedule the execution unit to perform decompression processing on the first data block, and at the same time, the second-level pipeline uses a sixth scheduling function to schedule the execution unit to perform decryption processing on the second data block; at a third time, the first-level pipeline uses an eighth scheduling function to schedule the execution unit to perform writing processing on the first data block, and at the same time, the second-level pipeline uses a seventh scheduling function to schedule the execution unit to perform decompression processing on the second data block.

[0129] Further, in a possible implementation manner of the embodiment of the present disclosure, the second processing unit 1520 is configured to: the first-level pipeline uses an eighth scheduling function to schedule the execution unit to send the first data block that has undergone decryption processing and decompression processing to a first cache space for caching through a read operation of the execution unit.

[0130] Further, in a possible implementation manner of the embodiment of the present disclosure, the data processing device further includes a fourth processing unit, configured to: obtain the first data block that has undergone decryption processing and decompression processing from the first cache space; display the first data block that has undergone decryption processing and decompression processing.

[0131] Further, in a possible implementation manner of the embodiment of the present disclosure, at least one of the first scheduling function, the second scheduling function, the third scheduling function, the fourth scheduling function, the fifth scheduling function, the sixth scheduling function, the seventh scheduling function, and the eighth scheduling function includes a first task block or a second task block.

[0132] For the description of the features in the corresponding embodiment of the data processing device, reference may be made to the relevant description of the corresponding embodiment of the data processing method, which will not be elaborated here one by one.

[0133] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above embodiments of the data processing method.

[0134] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any one of the above embodiments of the data processing method when running.

[0135] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as USB flash drives, read-only memory (ROM), random access memory (RAM), mobile hard disks, magnetic disks, or optical discs.

[0136] The embodiments of the present application also provide a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the data processing method.

[0137] The embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the data processing method.

[0138] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0139] The above has introduced in detail a data processing method, an electronic device, a storage medium, and a product provided by the present application. Specific examples are used herein to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A data processing method, characterized in that, The method includes: In response to a first operation instruction, dividing a file to be processed to obtain a plurality of data blocks corresponding to the file to be processed; Performing a first processing operation on a first data block among the plurality of data blocks, and simultaneously performing a second processing operation on a second data block among the plurality of data blocks, where the first data block is read before the second data block is read, and the first processing operation is an operation that is executed before the second processing operation according to the execution order of a plurality of processing operations, and the plurality of processing operations include at least two of a read processing, a compression processing, an encryption processing, and a write processing, or the plurality of processing operations include at least two of a read processing, a decryption processing, a decompression processing, and a write processing.

2. The method according to claim 1, wherein The method further includes: Determining a plurality of task blocks corresponding to the plurality of data blocks, where each task block includes information related to the data block corresponding to the task block; Determining identifiers of the plurality of task blocks; Determining a read order of the plurality of data blocks according to the identifiers of the plurality of task blocks.

3. The method according to claim 1, characterized in that, The method further includes: In a case where the plurality of processing operations include at least two of the read processing, the compression processing, the encryption processing, and the write processing, determining that the first operation instruction is a data write operation instruction; In a case where the plurality of processing operations include at least two of the read processing, the decryption processing, the decompression processing, and the write processing, determining that the first operation instruction is a data read operation instruction.

4. The method according to claim 3, characterized in that The performing a first processing operation on a first data block among the plurality of data blocks and simultaneously performing a second processing operation on a second data block among the plurality of data blocks includes: Inputting a first task block corresponding to the first data block into a pipeline software stack in a transparent file system; In response to the first task block, a first-level pipeline in the pipeline software stack performs the read processing on the first data block; When the first-level pipeline finishes performing the read processing on the first data block, inputting a second task block corresponding to the second data block into the pipeline software stack; In response to the second task block, a second-level pipeline in the pipeline software stack performs the second processing operation on the second data block, and simultaneously the first-level pipeline performs the first processing operation on the first data block.

5. The method according to claim 4, wherein The first-level pipeline in the pipeline software stack performing the read processing on the first data block in response to the first task block includes: Determining that the first operation instruction is a data write operation instruction, and the execution order of the plurality of processing operations is the read processing, the compression processing, the encryption processing, the write processing; The first-level pipeline uses a first scheduling function to schedule an execution unit, and obtains the first data block stored in a first storage device through a write operation of the execution unit; Writing the first data block into a first device memory of the execution unit.

6. The method according to claim 5, wherein The second-level pipeline in the pipeline software stack performing the second processing operation on the second data block in response to the second task block and simultaneously the first-level pipeline performing the first processing operation on the first data block includes: At the first time, the first - stage pipeline uses the second scheduling function to schedule the execution unit to perform the compression processing on the first data block. At the same time, the second - stage pipeline uses the first scheduling function to schedule the execution unit to obtain the second data block stored in the first storage device through the write operation of the execution unit and store the second data block in the first device memory; At the second time, the first - stage pipeline uses the third scheduling function to schedule the execution unit to perform the encryption processing on the first data block. At the same time, the second - stage pipeline uses the second scheduling function to schedule the execution unit to perform the compression processing on the second data block; At the third time, the first - stage pipeline uses the fourth scheduling function to schedule the execution unit to perform the writing processing on the first data block. At the same time, the second - stage pipeline uses the third scheduling function to schedule the execution unit to perform the encryption processing on the second data block.

7. The method according to claim 6, characterized in that, The first - stage pipeline using the fourth scheduling function to schedule the execution unit to perform the writing processing on the first data block includes: The first - stage pipeline uses the fourth scheduling function to schedule the execution unit to send the first data block after the compression processing and the encryption processing to the second storage device through the bus.

8. The method according to claim 4, characterized in that The first - stage pipeline in the pipeline software stack, in response to the first task block, performing the reading processing on the first data block includes: Determining that the first operation instruction is a data read operation instruction, and the execution order of the multiple processing operations is the reading processing, the decryption processing, the decompression processing, and the writing processing; The first - stage pipeline uses the fifth scheduling function to schedule the execution unit to obtain the first data block stored in the second storage device through the bus; Write the first data block into the first device memory of the execution unit.

9. The method according to claim 8, characterized in that, The second - stage pipeline in the pipeline software stack, in response to the second task block, performing the second processing operation on the second data block, while the first - stage pipeline performing the first processing operation on the first data block includes: At the first time, the first - stage pipeline uses the sixth scheduling function to schedule the execution unit to perform the decryption processing on the first data block. At the same time, the second - stage pipeline uses the fifth scheduling function to schedule the execution unit to obtain the second data block stored in the second storage device through the bus and store the second data block in the first device memory; At the second time, the first - stage pipeline uses the seventh scheduling function to schedule the execution unit to perform the decompression processing on the first data block. At the same time, the second - stage pipeline uses the sixth scheduling function to schedule the execution unit to perform the decryption processing on the second data block; At the third time, the first - stage pipeline uses the eighth scheduling function to schedule the execution unit to perform the writing processing on the first data block. At the same time, the second - stage pipeline uses the seventh scheduling function to schedule the execution unit to perform the decompression processing on the second data block.

10. The method according to claim 9, wherein The first - stage pipeline uses the eighth scheduling function to schedule the execution unit to perform the writing process on the first data block, including: The first - stage pipeline uses the eighth scheduling function to schedule the execution unit to send the first data block that has undergone the decryption process and the decompression process to the first cache space for caching through the read operation of the execution unit.

11. The method according to claim 10, wherein The method further includes: Obtain the first data block that has undergone the decryption process and the decompression process from the first cache space; Display the first data block that has undergone the decryption process and the decompression process.

12. The method according to any one of claims 5 to 11, characterized in that, At least one of the first scheduling function, the second scheduling function, the third scheduling function, the fourth scheduling function, the fifth scheduling function, the sixth scheduling function, the seventh scheduling function, and the eighth scheduling function includes the first task block or the second task block.

13. An electronic device, characterized in that, Includes: A memory for storing a computer program; A processor for implementing the steps of the data - processing method according to any one of claims 1 to 12 when executing the computer program.

14. A computer-readable storage medium, characterized in that, A computer program is stored in the computer - readable storage medium, wherein the computer program, when executed by a processor, implements the steps of the data - processing method according to any one of claims 1 to 12.

15. A computer program product, comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the steps of the data - processing method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • File writing method and device

    CN113641643A

  • Processing data stream modifications to reduce power effects during parallel processing

    CN115315688A

  • Data query method, system and device, medium and program product

    CN119885247A

  • Concurrent computations operating on same data for CPU cache efficiency

    US20190318016A1