A data processing method and apparatus
Patent Information
- Application Number
- CN202510387922.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2026-09-29
AI Technical Summary
这种并行化机制虽然在压缩阶段展现出卓越的时间效益,但是在实际应用中仍面临诸多挑战,如等距切割文件引起部分数据块压缩比急剧劣化(极端场景可能有5-10%的压缩比损失),导致文件压缩效果不佳
Smart Images

Figure CN122845667A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and more particularly to a data processing method and apparatus. Background Technology
[0002] In today's internet age, with its surging digital wave, information technology has deeply penetrated every corner of society, driving the unprecedented rapid development of numerous fields such as e-commerce, social media, online education, and cloud computing services. As the data foundation of these complex and massive business systems, databases shoulder a crucial mission, and their position in the entire internet ecosystem architecture is becoming increasingly pivotal.
[0003] A deeper exploration of the internal operating mechanisms of database storage engines reveals that their read and write performance is influenced by numerous factors, with compression and decompression performance undoubtedly playing a crucial and decisive role. In large-scale data storage scenarios, compression technology can effectively reduce data storage space and lower storage costs, while also reducing network bandwidth consumption and accelerating transmission efficiency during data transfer. However, when data needs to be read and written, the time consumed by decompression and compression operations becomes a bottleneck restricting the improvement of read and write speeds. Currently, mainstream storage engines widely used in the industry, such as Hadoop and Ceph, provide solid support for the processing of massive amounts of data thanks to their excellent distributed architecture and powerful storage management capabilities, and they all unanimously choose Zstandard (ZSTD) as their core compression component. ZSTD demonstrates certain advantages in many scenarios due to its excellent compression ratio, relatively fast compression speed, and good compatibility. However, it is undeniable that ZSTD's traditional compression and decompression mode is mostly a serial processing method, that is, compressing files one by one, and cannot perform parallel compression of a single file. This makes its processing efficiency inadequate when dealing with massive datasets. The lack of parallel processing capabilities prevents full utilization of the hardware resources of modern multi-core processors, resulting in a significant amount of CPU computation cycles being wasted waiting for individual data block compression and decompression to complete, which severely hinders the improvement of database storage engine throughput.
[0004] To overcome these challenges, a parallel compression mechanism has emerged. This mechanism innovatively introduces the concept of thread parallelism, aiming to fully leverage the advantages of multi-core processors and improve compression and decompression efficiency. Specifically, this mode divides the original single file into multiple data blocks at equal intervals, and then allocates a separate block for each thread to compress. In this way, multiple threads can compress the same file in parallel. Afterward, these scattered file fragments, compressed separately by each thread, are spliced and integrated to finally generate a complete compressed package. Although this parallel mechanism demonstrates excellent time efficiency in the compression stage, it still faces many challenges in practical applications. For example, equal-interval file division can cause a sharp deterioration in the compression ratio of some data blocks (in extreme cases, there may be a 5-10% loss in compression ratio), resulting in poor file compression performance. Summary of the Invention
[0005] This application provides a data processing method and apparatus for improving the compression performance of a parallel compression mechanism.
[0006] Firstly, this application provides a data processing method, which can be executed by a server or by a component of the server (such as a processor). Taking server execution as an example, the method includes: the server acquiring data to be compressed and dividing the data to be compressed into multiple data blocks; after obtaining multiple data blocks, compressing multiple first data blocks among the multiple data blocks in parallel, wherein compressing any first data block includes: compressing the first data block using a dictionary of second data blocks, wherein the second data block is a data block adjacent to the first data block among the multiple data blocks, for example, the second data block is the preceding data block or the following data block of the first data block, the dictionary of second data blocks is used to indicate the correspondence between the content identifier of each data unit in the second data block and the compressed data, the content identifier is determined based on the content of the data unit, and the compressed data is used to compress the data required by the data unit.
[0007] The above design divides the data to be compressed into multiple data blocks for parallel compression, ensuring compression efficiency. Furthermore, when compressing multiple data blocks of the data to be compressed in parallel, the data dependency between adjacent data blocks is cleverly utilized to improve the compression ratio, thereby improving compression performance while ensuring parallel compression efficiency.
[0008] In one possible design of the first aspect, the server compresses the first data block using a dictionary of the second data block, including: the server, based on the dictionary of the second data block, determines data units in the second data block whose content is identical to that of data units in the first data block. For example, if the first data block includes data unit A, the server determines data units in the second data block whose content is identical to that of data unit A, such as data unit B. For example, the content identifier of data unit B is the same as that of data unit A. Then, the server compresses data unit A using the compressed data corresponding to data unit B from the dictionary.
[0009] The above design utilizes the data dependencies between adjacent data blocks for data compression, which improves the compression ratio compared to compressing each data block individually.
[0010] In one possible design of the first aspect, in the dictionary of the second data block, the compressed data of a data unit in the second data block is the location information of that data unit, which can be used to indicate the position of the data unit in the second data block. Based on this, a specific way for the server to compress data unit A using the compressed data of data unit B is as follows: the server uses the location information of data unit B as the compressed data unit A; or, based on the location information of data unit B and the location information of data unit A, a position offset information is determined, which indicates the position offset between data unit B and data unit A. The server uses this position offset information as the compressed data unit A.
[0011] The above design allows for data compression using location information, which is compatible with existing compression algorithms and provides a wide range of application scenarios.
[0012] In one possible design of the first aspect, a specific way for the server to determine that the contents of data unit A and data unit B are the same is as follows: The server determines the content identifier of data unit A in the first data block, and determines whether the dictionary of the second data block contains the content identifier. Assuming that the second data block contains the content identifier, and the content identifier is the content identifier of data unit A in the second data block, then it is determined that the contents of data unit A and data unit B are the same. Alternatively, after determining that the content identifiers of the two data units are the same, the server obtains the content data of data unit B from the second data block based on the position information corresponding to the content identifier in the dictionary of the second data block. Then, the server compares whether the content data of data unit A and data unit B are consistent. If they are consistent, then it is determined that the contents of data unit A and data unit B are the same.
[0013] The above design uses content identifiers with relatively small data volume to determine whether the contents of two data units are the same, or uses content identifiers with relatively small data volume for preliminary matching, in order to reduce the storage overhead of the dictionary and the computational overhead of generating the dictionary and the matching steps. By comparing the contents, the compressed data of the data units with the same content is used for compression, ensuring the accuracy of the compressed data and achieving a high-performance lossless compression.
[0014] In one possible design of the first aspect, the server adds length indication information to the compressed data of each data block in the data to be compressed. For example, the server adds length indication information to the beginning of the compressed data corresponding to each data block. This length indication information is used to indicate the length of the compressed data of that data block. Then, the server combines the compressed data of each data block with the length indication information of each data block to obtain the compressed data of the data to be compressed. For example, the length indication information of each data block is concatenated before the compressed data of that data block to obtain the compressed data of the data to be compressed.
[0015] In one possible design of the first aspect, the server obtains the compressed data of the data to be compressed, for example, the complete compressed data of a file, and then reads multiple length indication information included in the complete compressed data; then, based on the multiple length indication information, the complete compressed data is divided into multiple compressed data blocks; the server performs parallel decompression on the multiple compressed data blocks to obtain the decompressed data.
[0016] With the above design, when the data to be compressed is the data of a file, the length indication information of each data block in the complete compressed data of the file can be used to divide the complete compressed data into blocks, thereby realizing parallel decompression of a file and improving decompression efficiency.
[0017] Secondly, this application also provides a data processing apparatus that has the function of implementing the behaviors in the first aspect and possible implementation methods of the first aspect. The beneficial effects can be found in the description of the first aspect and will not be repeated here. The functions can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions. In one possible design, the apparatus structure includes a segmentation module and a compression module. Optionally, it may also include a data processing module and a decompression module. In one possible design, some or all of the modules may be the same module; for example, the dictionary generation module and the compression module may be the same module. These modules can perform the functions of the behaviors in the method examples of the first aspect, as detailed in the method examples, and will not be repeated here.
[0018] Thirdly, this application also provides a computing device that has the functionality to implement the behavior in the method example of the first aspect described above. The beneficial effects are described in the first aspect and will not be repeated here. The computing device includes a processor and a memory. The processor is configured to support the computing device in executing the first aspect or any possible implementation of the first aspect. The memory is coupled to the processor and stores the necessary program instructions and data of the computing device. The computing device also includes a communication interface for communicating with other devices.
[0019] Fourthly, this application also provides a computing device cluster, which includes at least one computing device. This at least one computing device has the functionality to implement the behavior described in the method example of the first aspect above. The beneficial effects can be found in the description of the first aspect and will not be repeated here. Each computing device includes a processor and a memory. The processor is configured to support the computing device in executing the first aspect or any possible implementation of the first aspect. The memory is coupled to the processor and stores the necessary program instructions and data of the computing device. The computing device also includes a communication interface for communicating with other devices.
[0020] Fifthly, this application also provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the first aspect or any possible implementation of the first aspect.
[0021] Sixthly, this application also provides a computer program product containing instructions that, when run on a computer, causes the computer to perform the first aspect or any possible implementation of the first aspect.
[0022] In a seventh aspect, this application also provides a computer chip connected to a memory, which is used to read and execute a software program stored in the memory to perform the first aspect or any possible implementation method of the first aspect.
[0023] For the beneficial effects of aspects two through seven, please refer to the beneficial effects of aspect one, which will not be repeated here. Attached Figure Description
[0024] Figure 1 This application provides an illustration of an application scenario.
[0025] Figure 2 This application provides another application scenario illustration;
[0026] Figure 3 A schematic diagram of the structure of a computing device provided in this application;
[0027] Figure 4 A schematic diagram of the architecture of a storage system provided in this application;
[0028] Figure 5 A schematic diagram of another storage system architecture provided in this application;
[0029] Figure 6 A flowchart illustrating a data compression method provided in this application;
[0030] Figure 7 A schematic diagram of a multi-block parallel compression method provided in this application;
[0031] Figure 8 A schematic diagram illustrating another multi-block parallel compression method provided in this application;
[0032] Figure 9A A schematic diagram of a multi-threaded parallel compressed data block provided in this application;
[0033] Figure 9B This application provides a schematic diagram illustrating the relationship between data blocks and data units.
[0034] Figure 10 A schematic diagram of the structure of the compressed data provided in this application;
[0035] Figure 11 A flowchart illustrating a data decompression method provided in this application;
[0036] Figure 12A A schematic diagram of another compressed data structure provided in this application;
[0037] Figure 12B A schematic diagram illustrating the parallel decompression method provided in this application;
[0038] Figure 13 A schematic diagram of the structure of a data processing device provided in this application;
[0039] Figure 14 A schematic diagram of the structure of a computing device provided in this application. Detailed Implementation
[0040] The data processing methods in this application include data compression methods and data decompression methods, which are applicable to various application scenarios, such as data storage scenarios and data transmission scenarios.
[0041] Figure 1This diagram illustrates a data storage scenario provided by an embodiment of this application. In this data storage scenario, when a computing device acquires data that needs to be stored, such as data generated by a user using the computing device, the data compression method of this application can be used to compress the data and generate compressed data in order to reduce the storage space occupied by the data. The compressed data is stored in the memory or hard disk of the computing device to replace the original data before compression, thereby saving storage space. The computing device can be a device with computing and storage resources, such as a server, high-performance computer, laptop, desktop computer, mobile phone, etc.
[0042] When the computing device needs to obtain raw data, such as when it receives a request from another device to access raw data or when it receives a decompression request triggered by a user on the computing device, the computing device can obtain the compressed data from the hard disk or memory and decompress the compressed data into raw data using the data decompression method of this application.
[0043] Figure 2 This application provides a data transmission scenario. In this scenario, for a first device and a second device with data interaction needs, the first device can compress the original data using the data compression method of this application to generate compressed data. Then, the first device can send the compressed data to the second device. After receiving the compressed data, the second device can decompress the compressed data using the data decompression method of this application to obtain the original data. Optionally, the first device and the second device can communicate bidirectionally. The second device, as the sending end, compresses the original data using the data compression algorithm of this application and sends it to the first device. The first device, as the receiving end, decompresses the compressed data using the data decompression method of this application, as described above.
[0044] Both the first and second devices can be the aforementioned computing devices, as detailed below in the appendix. Figure 3 The internal components of the computing device are described.
[0045] Figure 3 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application.
[0046] like Figure 3 As shown, in terms of hardware, the computing device 10 includes at least a processor 123 and memory 124. Optionally, the computing device 10 may also include a hard disk 125 and a network card 126.
[0047] Processor 123 can be a central processing unit (CPU) used to process requests generated within the computing device and also to process data access requests from external devices.
[0048] For example, an internal request might involve the CPU 123 receiving a data compression request triggered by the user interface provided by the computing device 10. The CPU 123 then uses the data compression method of this application to compress the data to be compressed in the request, obtaining compressed data, and temporarily storing the compressed data in memory 124. When the total amount of data in memory 124 reaches a certain threshold, the processor 123 migrates the data stored in memory 124 to hard disk 125 for persistent storage. Alternatively, the CPU 123 might receive a data decompression request triggered by the user interface provided by the computing device 10. The CPU 123 retrieves the compressed data involved in the request from an external device, from memory 124, or from hard disk 125, and uses the data decompression method of this application to decompress the compressed data, obtaining the original data.
[0049] For example, when computing device 10 is the first device in a data transmission scenario, CPU 123 receives a data access request from the second device via network interface card 126. This data access request can be a write data request, carrying data to be written. CPU 123 compresses the data to be written using the data compression method of this application, generating compressed data, and temporarily stores the compressed data in memory 124. When the total amount of data in memory 124 reaches a certain threshold, processor 123 migrates the data stored in memory 124 to hard disk 125 for persistent storage. Alternatively, the data access request can be a read data request. CPU 123 retrieves the compressed data involved in the read data request from memory 124 or hard disk 125 and sends it to the second device via network interface card 126. Or, CPU 123 decompresses the compressed data using the data decompression method of this application to obtain the original data. Then, CPU 123 sends the original data to the second device via network interface card 126.
[0050] Figure 3Taking a CPU as an example, in practical applications, processors 123 can also be other general-purpose processors, digital signal processors (DSPs), hardware logic circuits, processing cores, application-specific integrated circuits (ASICs), AI chips, or programmable logic devices (PLDs). The aforementioned PLDs can be complex programmable logical devices (CPLDs), field-programmable gate arrays (FPGAs), generic array logic (GALs), systems-on-chips (SoCs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or any combination thereof. Furthermore, Figure 3 Only one CPU123 is shown in this embodiment. In practical applications, there are often multiple CPU123s, and each CPU123 has one or more processor cores. This embodiment does not limit the number of processors or the number of processor cores.
[0051] Memory 124 refers to the internal memory that directly exchanges data with processor 123. It can read and write data at any time at high speed, serving as temporary data storage for the operating system or other running programs. Memory 124 includes, but is not limited to: random access memory (RAM), read-only memory (ROM), static random access memory (SRAM), dynamic random access memory (DRAM), double data rate memory (DDR), and storage class memory (SCM). Read-only memory can be programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), etc. Memory 124 can be used as main memory. In practical applications, controller 0 can be configured with multiple memory modules 124, and different types of memory modules 124. This embodiment does not limit the number or type of memory module 124.
[0052] Hard disk 125 is used for persistent data storage. When data is not found in memory 124, data can be read from hard disk 125 and loaded into memory 124. Hard disk 125 can be non-volatile memory, such as read-only memory (ROM), hard disk drive (HDD), or solid state disk (SSD).
[0053] The network interface card 126 is used for the computing device 10 to communicate with other devices. For example, when the computing device 10 is the first device, it can communicate with the second device through the network interface card 126.
[0054] These components can be connected to each other via computer bus 127. Computer bus 127 includes, but is not limited to: peripheral component interconnect express (PCIe) bus, double data rate (DDR) bus, multi-protocol interconnect bus, serial advanced technology attachment (SATA) bus, serial attached SCSI (SAS) bus, controller area network (CAN) bus, computer express link (CXL) standard bus, etc., without specific limitations.
[0055] In addition to computing devices, embodiments of this application can also be applied to storage systems. Figure 4 This is a schematic diagram of a storage system architecture provided in an embodiment of this application.
[0056] exist Figure 4 In this system, users access data through applications. The computing devices running these applications are called "application servers." Application server 100 can be a physical machine or a virtual machine. Physical application servers include, but are not limited to, desktop computers, servers, laptops, and mobile devices. The application server accesses storage system 120 to access data via fiber optic switch 150. However, switch 150 is only an optional device; application server 100 can also communicate directly with storage system 120 via the network.
[0057] Figure 4The storage system 120 shown is a centralized storage system. A key feature of a centralized storage system is a unified entry point through which all data from external devices passes. This entry point is the engine 121 of the centralized storage system. The engine 121 is the most crucial component of the centralized storage system, and many advanced functions of the storage system are implemented within it.
[0058] like Figure 4 As shown, engine 221 contains one or more controllers. Figure 4 Taking engine 221 with two controllers as an example, controller 0 and controller 1 have a mirror channel. When controller 0 writes data to its memory 224, it can send a copy of that data to controller 1 via the mirror channel. Controller 1 then stores the copy in its local memory 224. Thus, controller 0 and controller 1 act as backups for each other. When controller 0 fails, controller 1 can take over its operations, and vice versa, preventing hardware failures from rendering the entire storage system 220 unavailable. When engine 221 has four controllers, any two controllers have a mirror channel, thus any two controllers act as backups for each other.
[0059] Engine 221 also includes a front-end interface 225 and a back-end interface 226. The front-end interface 225 is used to communicate with the application server 100 to provide storage services to the application server 100. The back-end interface 226 is used to communicate with the hard disk 105 to expand the capacity of the storage system. Through the back-end interface 226, engine 221 can connect to more hard disks 105, thus forming a very large storage resource pool.
[0060] Controller 0 can execute the data processing method provided in the embodiments of this application. In terms of hardware, such as... Figure 4 As shown, controller 0 includes at least processor 223 and memory 224. Controller 1 (and others) Figure 4 The hardware components and software structure of the controller (not shown) are similar to those of controller 0. Here, we will take controller 0 as an example for introduction.
[0061] Processor 223 is a central processing unit (CPU) used to process data access requests from outside the storage system (application server 100 or other storage systems), and also to process requests generated within the storage system. For example, when processor 223 receives a write data request from application server 100 through front-end port 225, it uses the data compression method provided in this embodiment to compress the data to be written, generating compressed data, which is then temporarily stored in memory 224. When the total amount of data in memory 224 reaches a certain threshold, processor 223 sends the data stored in memory 224 to hard disk 105 for persistent storage through back-end port. When processor 223 receives a read data request from application server 100 through front-end port 225, it uses the data decompression method provided in this embodiment to retrieve the compressed data from memory 224, decompress the compressed data into original data, and send the original data to application server 100.
[0062] The components and functions of processor 223, memory 224, and hard disk 105 can be found in the foregoing introduction to processor 123, memory 124, and hard disk 125, and will not be repeated here.
[0063] Figure 4 The diagram illustrates a centralized storage system with integrated disk and controller. In this system, engine 221 may not have hard drive bays; hard drive 105 needs to be placed in hard drive enclosure 130, and back-end interface 226 communicates with hard drive enclosure 130. The back-end interface 226 exists in the form of an adapter card within engine 221, and two or more back-end interfaces 226 can be used simultaneously on one engine 221 to connect multiple hard drive enclosures. Alternatively, the adapter card can be integrated onto the motherboard, in which case it can communicate with processor 223 via a high-speed Serial Computer Interconnect Express (PCI-E) bus. Optionally, in another centralized storage system architecture, storage system 220 can also be a storage system with integrated disk and controller. In this case, engine 221 has hard drive bays, and hard drive 105 can be directly deployed within engine 221, meaning hard drive 105 and engine 221 are deployed on the same device.
[0064] The above describes a centralized storage system. In addition to being applicable to centralized storage systems, the embodiments of this application are also applicable to distributed storage systems.
[0065] Figure 5 This is a schematic diagram of a distributed storage system architecture provided in an embodiment of this application. The distributed storage system includes a server cluster. Figure 5 As shown, the server cluster includes one or more servers 140 ( Figure 5 The illustration shows three servers 140, but this application is not limited to three servers 140, and each server 140 can communicate with each other.
[0066] In terms of software, each server 140 has an operating system. Optionally, virtual machines 107 can be created on server 140. The computing resources required by virtual machine 107 come from the local processor 112 and memory 113 of server 140, while the storage resources required by virtual machine 107 can come from the local hard disk 105 of server 140 or from the hard disk 105 of other servers 140. In addition, various applications can run in virtual machine 107, and users can trigger read / write data requests through the applications in virtual machine 107.
[0067] In terms of hardware, such as Figure 5 As shown, server 140 includes at least processor 112, memory 113, network interface card 114, and hard disk 105. Processor 112, memory 113, network interface card 114, and hard disk 105 are connected via a computer bus. Processor 112 can execute the data processing method provided in this application. For details on the functions of components such as processor 112, memory 113, network interface card 114, and hard disk 105, please refer to... Figure 4 or Figure 5 The relevant information will not be repeated here.
[0068] It should be noted that, Figure 4 , Figure 5 These are possible examples of centralized storage systems and distributed storage systems, respectively. The embodiments in this application are not limited to these. Figure 4 , Figure 5 The system architecture shown is applicable to any system architecture with data storage and processing capabilities as described in this application. Furthermore, Figures 3-5 The system architecture shown is for illustrative purposes only; real-world systems may include more complex architectures. Figures 3-5 More or fewer devices, computing devices or storage nodes can have relative Figures 3-5 The embodiments of this application do not limit the number of components, whether more or fewer.
[0069] Both the data compression and decompression methods in this application employ parallelization. In the data compression method, the data to be compressed is divided into multiple data blocks. Each data block is independently compressed by an execution unit (such as a thread or process) to obtain compressed data for that data block. Multiple execution units then concatenate their respective compressed data to form complete compressed data. Unlike existing parallel compression mechanisms, in this application, each execution unit compresses its own data block based on the data dependencies between multiple data blocks, increasing data stickiness during the compression process and thus improving compression performance. Based on the data decompression method provided in this application, compressed data blocks within the compressed data can be identified during decompression, enabling parallel decompression of multiple compressed data blocks and improving decompression efficiency.
[0070] The following describes the data processing method provided in this application embodiment, using the above-mentioned system as an example. This method can be divided into two stages: the first stage is a data compression process (see...). Figure 6 The method flow), on the other hand, is the data decompression process (see the method flow). Figure 11 (Methods and procedures).
[0071] Figure 6 This is a flowchart illustrating the data compression method provided in an embodiment of this application. The method can be... Figures 1-3 The method is executed by the computing device in the system, or by a component of the computing device (such as CPU123), or the method can be executed by... Figure 4 or Figure 5 The data compression method provided in this application is executed by a device in the storage system (such as controller 0 or server 140), or by a component of these devices (such as CPU 223 or CPU 112). The following description uses a server as the execution subject to illustrate the data compression method provided in this application. It should be understood that when other devices, such as the aforementioned computing devices, execute the data compression method, their execution process is similar to that of the server executing the data compression method; the main difference lies in the execution subject. This method may include some or all of the following steps:
[0072] Step 601: The server obtains the data to be compressed.
[0073] There are several ways for a server to obtain the data to be compressed; it can obtain it from internal devices or external devices. For example, in Figure 1 In this context, users use computing devices to generate data to be compressed, such as Word or Excel files created by the user on the computing device, or files imported by the user. The computing device receives the data to be compressed, either input or selected by the user. For example, in... Figure 2 In this process, the first device receives the data to be compressed from the second device. For example, in... Figure 4The storage system receives data to be compressed from the application server, or... Figure 5 In this context, the server receives data to be compressed from external devices or other servers.
[0074] When the data to be compressed is obtained by the server from an external device, the acquisition method can be either passive reception by the server or active acquisition by the server, for example... Figure 2 When a second device needs to use the storage resources of a first device, the second device sends the data to be written to the first device. The first device passively receives the data to be written from the second device and uses it as data to be compressed. Alternatively, it could be a backup scenario where the first device periodically requests backup data from the second device according to a predetermined plan and uses the backup data obtained from the first device as data to be compressed.
[0075] The data to be compressed can be determined by the server or specified by the user. For example, after receiving data to be written from another device, the server may proactively initiate a compression process before writing the data to save storage overhead, treating the data to be written as the data to be compressed. Alternatively, the server may receive a data compression request triggered by the user, for example... Figure 1 When a user selects one or more files in a pop-up window, the user selects one or more files to compress. The computing device receives the compression request triggered by the user and selects one or more files as files to be compressed.
[0076] In other words, the data to be compressed includes data from one or more data objects. These data objects can be of various types, such as files (e.g., text, images, audio, video, executable files, or program files), data sets (e.g., database records, tables, training sets, or validation sets), log file sets, data streams (e.g., network data streams, various data structures stored in memory, cached data), or other data types. There are no specific limitations on these types.
[0077] Step 602: The server divides the data to be compressed into multiple data blocks, with each data block assigned to an execution unit.
[0078] For ease of explanation, we will use a file as an example here. In step 601, the data to be compressed obtained by the server may include one or more files.
[0079] When the data to be compressed includes multiple files, a coarse-grained approach allows the server to compress these files in parallel, improving overall compression efficiency. A fine-grained approach allows the server to divide each file into multiple data blocks and use the data compression algorithm of this application to compress each file's multiple data blocks in parallel, further improving the compression efficiency of each file. Similarly, when the data to be compressed includes only one file, the data compression method of this application is used to process that single file.
[0080] The following section uses a single file as an example to introduce this data compression method.
[0081] The server divides a file containing data to be compressed into multiple data blocks, which are then assigned to multiple execution units (such as processes or threads). Each execution unit is responsible for compressing one of the data blocks. For example, the server selects multiple processes and assigns the data blocks to them, or it selects multiple threads and assigns the data blocks to them. These processes / threads can be selected from idle processes / threads or created specifically for this data compression task. For clarity, the following explanation uses threads as an example.
[0082] There are several ways for a server to divide a file into multiple data blocks. In one method, the server divides the file's content data into multiple blocks at equal intervals, meaning the file's content data is evenly divided into multiple blocks, each of equal length. This allows for a more balanced load distribution among threads during parallel compression, preventing less heavily loaded threads from waiting for excessively loaded execution, thereby improving the utilization of computing resources. Alternatively, non-equal interval partitioning can be used, where the file data is divided into multiple data blocks of varying or completely different lengths.
[0083] Step 603: The server uses adjacent data blocks to perform parallel compression on the multiple data blocks.
[0084] After obtaining multiple data blocks, the server compresses multiple first data blocks in parallel. Compression of any first data block includes compressing it using its adjacent data blocks. Adjacent data blocks refer to data blocks that are adjacent to the first data block among the multiple data blocks; these can be either the preceding or following data block. These multiple first data blocks can be all or a subset of the data blocks in the multiple data blocks.
[0085] For example, see Figure 7As shown, the content data of a certain file is divided into block1, block2, block3, ..., block n. Taking block2 as an example, the adjacent data blocks of block2 are block1 and block3, where block1 is the data block before block2 and block3 is the data block after block2.
[0086] In one example, see [link to example]. Figure 7 As shown, the server compresses the first data block using the preceding data block. For example, it uses block1 to compress block2, block2 to compress block3, block3 to compress block4, and so on. Optionally, block1 can be compressed independently since it has no preceding data block. Specifically, in this example, the server can assign block1 to thread 1, block2 and block3 to thread 2, block3 and block4 to thread 4, and so on. Thread 1 compresses block1, thread 2 uses block1 to compress block2, thread 3 uses block2 to compress block3, and so on.
[0087] In another example, see Figure 8 As shown, the server compresses the first data block using the subsequent data block. For example, block1 is compressed using block2, block2 using block3, block3 using block4, and so on. Optionally, blockn can be compressed independently since it has no subsequent data block. Specifically, in this example, the server can assign block1 and block2 to thread 1, block2 and block3 to thread 2, block3 and block4 to thread 3, and so on. Thread 1 compresses block1 using block2, thread 2 compresses block2 using block3, thread 3 compresses block3 using block4, and so on.
[0088] The specific method for compressing the first data block based on its adjacent data blocks may include: the server generating a dictionary of the adjacent data blocks, and using the dictionary of the adjacent data blocks to compress the first data block.
[0089] by Figure 7 For example, in Figure 7There are various ways to perform parallel compression of multiple data blocks using multiple threads. In some embodiments, the server includes a main execution unit that can send two adjacent data blocks together to their respective threads, and each thread can generate its own dictionary of adjacent data blocks. For example, in... Figure 7 In the main execution unit, block1 and block2 are sent to thread 2, block2 and block3 to thread 3, block3 and block4 to thread 4, and so on. Then, thread 2 generates a dictionary of block1, thread 3 generates a dictionary of block2, thread 4 generates a dictionary of block3, and so on. Afterwards, participants... Figure 9A As shown, thread 1 directly compresses block 1, thread 2 uses the dictionary of block 1 to compress block 2, thread 3 uses the dictionary of block 2 to compress block 3, thread 4 uses the dictionary of block 3 to compress block 4, and so on, with thread n using block(n-1) to compress block n. This parallel approach to dictionary generation effectively improves dictionary generation efficiency, thereby shortening compression time and improving compression performance.
[0090] In other embodiments, the main execution unit is responsible for generating dictionaries for multiple data blocks and distributing the dictionary for data block i and data block i+1 to the corresponding threads. For example, in Figure 7 In the main execution unit, the dictionary of block1 and block2 are sent to thread 2, the dictionary of block2 and block3 are sent to thread 3, and so on. Optionally, the main execution unit can also send the dictionary of data block i, data block i, and data block i+1 together to the corresponding thread, for example, in... Figure 7 In the example, the main thread sends the dictionary of block1, block1, and block2 to thread2, the dictionary of block2, block2, and block3 to thread3, and so on.
[0091] The main execution unit can be one of the multiple threads responsible for compression, or it can be a thread or a process outside of those multiple threads; there is no specific limitation.
[0092] Next, we will introduce the dictionary.
[0093] Taking a dictionary of a data block as an example, a dictionary is typically a data structure used to store and manipulate key-value pairs. In this embodiment, the dictionary of a data block may include the correspondence between the content identifier (which can be used as a key) and the compressed data (as a value) of each data unit in the data block.
[0094] Taking a data block as an example, a data block consists of multiple data units, each of equal length. For instance, each data unit is 4 bytes long, meaning each data block is divided into multiple data units of 4 bytes each. Figure 9B As shown, assuming block1 includes ABDEFHIGHKLO..., where each letter represents one byte, if divided into 4-byte units, the data units in block1 are: ABDE, FHIG, HKLO, ..., XXXX. The dictionary of block1 includes: the correspondence between the content identifier of ABDE and the compressed data, the correspondence between the content identifier of FHIG and the compressed data, the correspondence between the content identifier of HKLO and the compressed data, and so on.
[0095] The content identifier is determined based on the content of a data unit and is used to identify that content. For example, the content identifier is a data fingerprint. A data fingerprint (usually based on a hash algorithm, such as SHA-256) is a fixed-length value that uniquely identifies the data content. In other words, the content identifier can be the hash value of the content data of the data unit, or a data identifier generated based on other algorithms and content data; there is no specific limitation.
[0096] Compressed data refers to the data used to compress this data unit. This compressed data can be the compressed version of the data unit, or it can be the data used to generate the compressed data for this data unit. For example, it could be the location information of the data unit, indicating its position within the data block. This location information could be the starting address of the data unit, such as the offset of the first byte of the data unit within the block. For instance, in block1 of the example above, the location information of data unit ABDE could be the offset of A within block1, the location information of data unit FHIG could be the offset of F within block1, and the location information of data unit HKLO could be the offset of H within block1.
[0097] Optionally, in this example, the dictionary of block1 may include: the correspondence between the hash values of ABDE and the offsets of ABDE, the correspondence between the hash values of FHIG and the offsets of FHIG, and the correspondence between the hash values of HKLO and the offsets of HKLO.
[0098] It should be noted that the offset of a data unit within a data block can be determined by its position within the entire file. For example, if the offset of the first data unit in data block 1 is 0, and the length of data block 1 is 4n bytes, then the offset of the first data unit in data block 2 is 4n, the offset of the next data unit is 4(n+1), and so on. Alternatively, it can be a relative offset of the data unit within the data block. For example, if the offset of the first data unit in data block 1 is 0, the offset of the second data unit is 4 bytes, and so on, the offset of the first data unit in data block 2 is also 0, the offset of the second data unit is 4 bytes, the offset of the third data unit is 8 bytes, and so on.
[0099] The above description is based on a data unit length of 4 bytes. This application also supports other lengths, such as each data unit having a length of 2 bytes, 8 bytes, or other values. There is no specific limitation, which can be determined by the computer system architecture. For example, in a 32-bit computer system, a word is usually 32 bits (equivalent to 4 bytes), and in a 64-bit computer system, a word is usually 64 bits (equivalent to 8 bytes).
[0100] The following section will explain how to use a dictionary to compress data blocks.
[0101] For ease of explanation, the following example illustrates the compression of the second data block using the dictionary of the first data block. The process may include:
[0102] Based on the dictionary of the first data block, identify the data units in the second data block that have the same content as the first data block. Use the position information of the data unit in the dictionary of the first data block to compress the data units in the second data block that have the same content. For ease of explanation, the data units with the same content in the first data block are referred to as the first data unit, and the data units with the same content in the second data block are referred to as the second data unit.
[0103] In one example, if two data units have the same content identifier, the two data units are considered to have the same content. This determination process may include: generating a content identifier for each data unit in the second data block; for each data unit, taking the first data unit as an example, determining whether the dictionary of the first data block includes the content identifier of the first data unit; if it does, it indicates that there exists a data unit (i.e., the second data unit) in the first data block with the same content as the first data unit. Otherwise, it indicates that there is no data unit in the first data block with the same content as the first data unit.
[0104] In another example, after determining that the content identifiers of two data units are the same, the content data of the two data units are compared. If they match, the content of the two data units is considered the same; otherwise, the content of the two data units is considered different. This process may include: generating a content identifier for each data unit in the second data block; for each data unit, taking the first data unit as an example, determining whether the dictionary of the first data block includes the content identifier of the first data unit; if it does, then reading the data unit (i.e., the second data unit) at the corresponding position in the first data block according to the position information corresponding to the content identifier in the dictionary; comparing the first data unit and the second data unit; if their content data is completely identical, it means that the content of the first data unit and the second data unit is the same; if they are inconsistent, it means that their content data is completely different or not completely identical.
[0105] If the contents of the first data unit and the second data unit are the same, then when compressing the second data unit, the compressed data of the first data unit in the dictionary of the first data block is used to compress the second data unit. It should be noted that there may be multiple second data units in the second data block, and the compression method for each second data unit is the same; in other words, the compressed data of each second data unit is the same.
[0106] For example, continue with Figure 7 Taking thread 2 as an example, thread 2 uses the dictionary of block1 to compress block2. At this time, block1 is used as the first data block mentioned above, and block2 is used as the second data block.
[0107] Assume that the dictionary of block1 may include: the correspondence between the hash values of ABDE (let's say h1) and their offsets (let's say f1), the correspondence between the hash values of FHIG (let's say h2) and their offsets (let's say f2), and the correspondence between the hash values of HKLO (let's say h3) and their offsets (let's say f3). In short, the dictionary of block1 includes (h1, f1), (h2, f2), and (h3, f3).
[0108] Thread 2 generates the content identifier for each data unit in block 2. Assuming block 2 includes NKFCABDEHKLOABDE..., thread 2 will obtain the content identifiers for NKFC (h4), ABDE (h1), HKLO (h3), and so on. These content identifiers may be duplicated. Taking one content identifier as an example, thread 2 matches the content identifier (h1) of ABDE (as the first data unit) with the dictionary of block 1. If the dictionary of block 1 contains h1, it searches for the data unit at position f1 in block 1 based on the offset f1 corresponding to h1 in the dictionary, such as obtaining ABDE (as the second data unit). After this, thread 2 compares the ABDE in block 2 with the ABDE in block 1. Since their content data is completely identical, the two data units are confirmed to be the same. When compressing the ABDE in block 2, thread 2 uses the offset (e.g., f1) corresponding to h1 in the dictionary of block 1 to compress the ABDE. For example, the compressed data is f1. Alternatively... In an alternative approach, the compressed data can also be f1', where f1' indicates the offset between the first and second data units. For example, assuming the offset of ABDE in block2 is f4, and the offset of ABDE in block1 is known to be f1, then the compressed data of ABDE can be f1-f4 (or f4-f1). This indicates that the data units at positions f1-f4 in the file have the same content as the data unit at position f1. The data units at positions f1-f4 can be used to recover the data units at positions f1-f4, allowing for faster recovery of the uncompressed data.
[0109] Thread 2 compresses each data unit in block 2 using the method described above, obtaining the compressed data for each data block in block 2. The remaining threads perform parallel compression on their respective data blocks based on the method used by thread 2, and each thread obtains the compressed data for its assigned data block; details will not be elaborated here.
[0110] The compression method described above can achieve lossless compression. Lossless compression means that any set of original data can be processed into compressed data with a smaller memory footprint using this algorithm, and the content after data restoration is completely consistent with the original data without any information loss, thus ensuring the integrity of the original data. The embodiments of this application are compatible with existing compression algorithms. For example, the ZSTD algorithm is also a lossless compression algorithm, and the data compression method of this application can be applied to the ZSTD algorithm.
[0111] Step 604: The server concatenates the compressed data from multiple data blocks into complete compressed data.
[0112] The server sequentially concatenates the compressed data blocks of a file to obtain the complete compressed data of the file. Continuing with this... Figure 7 The process involves each thread writing the compressed data of its assigned data block into a designated buffer, which can be done serially or in parallel. The compressed data from multiple data blocks is then concatenated sequentially to form the complete compressed data for the file.
[0113] It should be noted that, Figure 7 The compressed data shown is for illustrative purposes only. The complete compressed data of a file may contain more information, depending on the compression algorithm or protocol used.
[0114] For example, the RFC 8478 protocol requires that files compressed using the Zstandard (ZSTD) algorithm meet the protocol's format requirements. For instance, the complete compressed data of a file consists of multiple compressed frames, each of which may contain one or more compressed blocks. Here, compressed frames and compressed blocks are concepts in the ZSTD algorithm. A compressed frame is a complete logical unit of data compressed using the ZSTD algorithm, and may include one or more compressed blocks, as well as frame headers and trailers. Multiple compressed frames can be sequentially concatenated to form a continuous compressed data stream. Each compressed frame is independent and has clearly defined start and end markers. A compressed block is the smallest unit of compression and a component of a frame. For example, in one example, the compressed data from each of the aforementioned data blocks can be considered a compressed block. When a compressed frame includes multiple compressed blocks, the complete compressed data of a file may include, for example, […]. Figure 10 The information shown in (a).
[0115] In another example, a compressed frame may contain only one compressed block. In this case, the compressed data of each data block constitutes a compressed frame, and the complete compressed data of a file may include, for example, ... Figure 10 The information shown in (b).
[0116] It should be noted that, Figure 10 Examples (a) and (b) are merely illustrative; in practice, frame structures may have more or less information, for example, they may lack a frame trailer. Optionally, embodiments of this application may also utilize existing compression algorithms (such as...) Figure 10 (a) or Figure 10 Based on the compressed data structure shown in (b)), modifications are made to obtain the complete compressed data of a file. The following will combine... Figure 12A This will be introduced in detail here, but will not be repeated.
[0117] When the data to be compressed includes multiple files, each file is compressed separately as described above to obtain the complete compressed data for each file. Optionally, the server can package the compressed data of multiple files into a single compressed archive.
[0118] The above section used files as an example to introduce parallel compression methods for files. If the data to be compressed is of other types, such as tables or datasets, taking tables as an example, the server can divide a table into multiple data blocks and perform parallel compression on each table to obtain compressed data for each table. See the description above regarding files as the block object; it will not be repeated here.
[0119] With the above design, after multiple threads divide a complete file into multiple data blocks, they can use adjacent data blocks for compression. By leveraging the data dependencies between adjacent data blocks, the compression ratio can be improved, thereby improving compression performance while ensuring parallel compression efficiency.
[0120] Figure 11 This is a flowchart illustrating the data decompression method provided in this application embodiment. This method can be performed by... Figures 1-3 The method is executed by the computing device in the system, or by a component of the computing device (such as CPU123), or the method can be executed by... Figure 4 or Figure 5 The data decompression method provided in this application embodiment is executed by a device in the storage system (such as controller 0 or server 140), or by a component of these devices (such as CPU 223 or CPU 112). The following description uses a server as the execution subject. It should be understood that when other devices, such as the aforementioned computing devices, execute the data decompression method, their execution process is similar to that of the server executing the data decompression method; the main difference lies in the execution subject. This method may include some or all of the following steps:
[0121] Step 1101: The server obtains the compressed data (i.e., the data to be decompressed).
[0122] There are several ways for a server to obtain compressed data; it can obtain it from the device's internal storage or from external devices. For example, for Figure 3 The computing device 10 shown can retrieve compressed data from memory 124 or hard disk 125. For example, in... Figure 2 In this process, the first device obtains compressed data from the second device, etc.
[0123] Continuing with the file example, compressed data refers to data compressed using a compression algorithm. In the example above, compressed data could be the compressed data of a single file, or it could be a compressed archive containing the compressed data of one or more files. Compressed data can be in a specific format, such as zip or rar file formats, but there are no specific limitations.
[0124] Step 1102: The server divides the compressed data into multiple compressed data blocks.
[0125] Taking the complete compressed data of a file as an example, the complete compressed data includes multiple length indication information. Based on these multiple length indication information, the server divides the complete compressed data of a file into multiple data segments, and each data segment is recorded as a compressed data block.
[0126] For example, in step 604, when the server concatenates the compressed data from multiple data blocks into complete compressed data, it can... Figure 10 Based on the structure shown in (b), further information can be added, such as adding specific boundary magic words and custom data before each compressed frame (as allowed by the RFC 8478 protocol), resulting in... Figure 12A The file shown contains complete compressed data. Boundary magic words can be used to indicate the interval between compressed frames. Custom data can be the length of the compressed frames (denoted as length indication information).
[0127] The server reads the length indication information after each compressed frame's boundary magic word. This length indication information can be used to indicate the length of time until the next data frame (e.g., ...). Figure 12A In the L), the server reads the next compressed frame based on the length indication information, and then reads the subsequent compressed frames based on the length indication information of the next compressed frame, thereby dividing the complete compressed data of a file into multiple compressed frames. When a compressed frame includes a compressed block, and each compressed block is a block of compressed data, the compressed data of multiple blocks of a file can be obtained in the above way.
[0128] For example, in Figure 10 Based on (b), a magic word 1 and custom data 1 are inserted at the beginning of compressed frame 1 (containing the compressed data of block 1), and custom data 1 is the length of compressed frame 1; a magic word 2 and custom data 2 are inserted at the beginning of compressed frame 2 (containing the compressed data of block 2), and custom data 2 is the length of compressed frame 2; a magic word 3 and custom data 3 are inserted at the beginning of compressed frame 3 (containing the compressed data of block 2), and custom data 3 is the length of compressed frame 3, and so on.
[0129] Based on this data structure, after recognizing magic word 1, the server reads custom data 1 and determines the length of compressed frame 1, thereby extracting the compressed data of block 1 from compressed frame 1. Then, based on the length of compressed frame 1, it determines the starting address of magic word 2, and sequentially reads custom data 2 according to this address, determining the length of compressed frame 2 and extracting the compressed data of block 2 from compressed frame 2; then, based on the length of compressed frame 2, it determines the starting address of magic word 3, and sequentially reads custom data 3 according to this address, determining the length of compressed frame 3 and extracting the compressed data of block 3 from compressed frame 3, and so on, until compressed frame n is obtained. In this way, the server can divide the complete compressed data of a file into multiple data segments (i.e., compressed data blocks) to be decompressed. For example, in... Figure 12A In the diagram, compressed data block 1 includes the compressed data of block 1, compressed data block 2 includes the compressed data of block 2, compressed data block 3 includes the compressed data of block 3, and so on.
[0130] It should be noted that, Figure 12A This is merely an example; in practical applications, magic text and custom data can also be inserted after each compressed frame, and this embodiment does not limit this. Alternatively, magic text can be added only before the first compressed frame. There are no specific limitations.
[0131] For another example, the server (or the party compressing the data) records the offset and length of each compressed frame within the compressed data. When decompressing the compressed data, this recorded information is used to divide the compressed data into blocks. Alternatively, the server and the party compressing the data may pre-agree on the length of the compressed data blocks, or the compressed data blocks may have a uniform length. The server then divides the compressed data according to this pre-defined length.
[0132] Step 1103: The server performs parallel decompression on multiple compressed data blocks.
[0133] join Figure 12B As shown, similar to parallel compression mechanisms, the server uses multiple execution units (such as threads or processes) to decompress multiple compressed data blocks in parallel. Each execution unit corresponds to one compressed data block, and each execution unit performs parallel decompression on its assigned compressed data block. For example, in Figure 12BIn the example, thread 1 decompresses compressed data block 1 to obtain the original content data of block 1; thread 2 decompresses compressed data block 2 to obtain the original content data of block 2; thread 3 decompresses compressed data block 3 to obtain the original content data of block 3, and so on. Existing decompression algorithms can be used during decompression. For example, the data content of the corresponding uncompressed data unit can be found based on the compressed data (e.g., offset) of each data unit, and the compressed data can be replaced to restore the original data of each data block.
[0134] Step 1104: The server concatenates the decompressed data from multiple compressed data blocks into complete original data.
[0135] Each thread writes the decompressed data to a specified buffer, and the original data of the file is obtained by sequentially concatenating multiple decompressed data.
[0136] If the data to be decompressed includes compressed data from multiple files, the above method can be used to decompress each file in parallel to obtain the original data of each file.
[0137] Through the above design, custom data (such as length indication information) is identified by boundary magic words in the compressed data, and the length indication information is read to determine the length of the current compressed frame. Based on this length, the length indication information of the next compressed frame can be read directly without skipping the current compressed frame. Multiple compressed data blocks can be extracted from the compressed data without reading all the compressed data. Compared with the existing technology that can only decompress multiple files in parallel, the embodiments of this application can realize parallel decompression at the file granularity. That is, the compressed data of a file can be divided into multiple data blocks for parallel decompression, thereby improving the decompression efficiency of a file.
[0138] It should be noted that, Figure 6 and Figure 11 The method embodiments can be executed on the same device or in the same embodiment, or, Figure 6 and Figure 11 It can be executed on two separate devices or in two separate embodiments, and this application does not limit this to any particular method.
[0139] Based on the same inventive concept as the method embodiments, this application also provides a data processing apparatus for performing the above-described... Figure 6 or Figure 11 The server executes data processing methods. For example... Figure 13As shown, in one example, the data processing device 1300 includes a segmentation module 1301 and a compression module 1302. Optionally, it also includes a data processing module 1303, an acquisition module 1304, and a decompression module 1305. Specifically, in the data processing device 1300, the modules are connected to each other through a communication path.
[0140] The segmentation module 1301 is used to divide the data to be compressed into multiple data blocks; see details below. Figure 6 The description of step 602 in the method embodiment will not be repeated here.
[0141] Compression module 1302 is used to compress multiple first data blocks among the plurality of data blocks in parallel. Compression of any one of the first data blocks includes: compressing the first data block using a dictionary of second data blocks, where the second data block is a data block adjacent to the first data block among the plurality of data blocks. The dictionary of the second data block is used to indicate the correspondence between the content identifier and the compressed data for each data unit in the second data block. The content identifier is determined based on the content of the data unit, and the compressed data is used to compress the data required by the data unit. See details... Figure 6 The description of step 603 in the method embodiment will not be repeated here.
[0142] In one possible design, the acquisition module 1304 is used to acquire the data to be compressed; see details below. Figure 6 The description of step 601 in the method embodiment will not be repeated here.
[0143] In one possible design, when the compression module 1302 compresses the first data block using the dictionary of the second data block, it is specifically used to: determine multiple data units with the same content in the second data block and the first data block according to the dictionary of the second data block; the multiple data units include the first data unit in the first data block and the second data unit in the second data block; and compress the first data unit using the compressed data of the second data unit.
[0144] In one possible design, the compressed data of the data unit in the second data block is the location information of the data unit, which is used to indicate the position of the data unit in the second data block;
[0145] When the compression module compresses the first data unit using the compressed data of the second data unit, it is specifically used to: use the position information of the second data unit as the compressed first data unit; or use the position offset information as the compressed first data unit, wherein the position offset information indicates the position offset between the first data unit and the second data unit.
[0146] In one possible design, when the compression module 1302 determines multiple data units with the same content in the second data block and the first data block based on the dictionary of the second data block, it is specifically used to: determine the first data unit based on the content identifier of each data unit in the first data block and the dictionary of the second data block, wherein the content identifier of the first data unit is included in the dictionary of the second data block; determine the position information in the dictionary corresponding to the content identifier of the first data unit, and obtain the second data unit at the position indicated by the position information in the second data block, wherein the second data unit has the same content as the first data unit.
[0147] In one possible design, the second data block is the preceding data unit adjacent to the first data block, or the second data block is the following data unit adjacent to the first data block.
[0148] In one possible design, the data processing module 1303 is used to add length indication information to the compressed data of each data block in the data to be compressed. The length indication information is used to indicate the length of the compressed data of each data block. Based on the compressed data of each data block and the length indication information of each compressed data block, the compressed data is obtained. See details... Figure 6 The description of step 604 in the method embodiment will not be repeated here.
[0149] In one possible design, the acquisition module 1304 is used to acquire the compressed data;
[0150] The decompression module 1305 is used to read multiple length indication information included in the compressed data; divide the compressed data into multiple data blocks based on the multiple length indication information; and perform parallel decompression on the multiple data blocks to obtain the decompressed data. See details below. Figure 11 The description of the method implementation examples will not be repeated here.
[0151] In another example, embodiments of this application also provide another data processing apparatus, which may include only an acquisition module 1304 and a decompression module 1305, these two modules being used to execute independently. Figure 11 The process of the method implementation will not be described again here.
[0152] It should be noted that the module division in this embodiment is illustrative and represents only one logical functional division. In actual implementation, other division methods may be used. The functional modules in this embodiment can be integrated into one module, or each module can exist physically separately, or two or more modules can be integrated into one module. For example, the compression module 1302 and the data processing module 1303 can be integrated into one module. The integrated units described above can be implemented in hardware or as software functional units.
[0153] This application also provides a computing device 1400. For example... Figure 14 As shown, the computing device 1400 includes a bus 1402, a processor 1404, a memory 1406, and a communication interface 1408. The processor 1404, the memory 1406, and the communication interface 1408 communicate with each other via the bus 1402. The computing device 1400 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 1400.
[0154] Bus 1402 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 14 The bus 1402 may be represented by a single line, but this does not mean that there is only one bus or one type of bus. The bus 1402 may include a path for transmitting information between various components of the computing device 1400 (e.g., memory 1406, processor 1404, communication interface 1408).
[0155] The processor 1404 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0156] The memory 1406 may include volatile memory, such as random access memory (RAM). The processor 1404 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0157] In one possible design, the memory 1406 stores executable program code, which the processor 1404 executes to implement the functions of the aforementioned segmentation module 1301, compression module 1302, data processing module 1303, acquisition module 1304, and decompression module 1305, thereby realizing the data processing method. That is, the memory 1406 stores instructions for the data processing device 1300 to execute the data processing method provided in this application.
[0158] The communication interface 1408 uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between the computing device 1400 and other devices or communication networks.
[0159] This application also provides a computing device cluster. The computing device cluster includes at least one computing device 1400. The computing device 1400 may be a server. In some embodiments, the computing device 1400 may also be a desktop computer, a laptop computer, or a smartphone, or other terminal device.
[0160] Based on the above embodiments, this application also provides a computer program that, when run on a computer, causes the computer to perform... Figure 6 or Figure 11 The data processing method provided in the illustrated embodiment.
[0161] Based on the above embodiments, this application also provides a computer-readable storage medium storing a computer program, which, when executed by a computer, causes the computer to perform... Figure 6 or Figure 11 The illustrated embodiment provides a data processing method. The storage medium can be any available medium accessible to a computer. For example, but not limited to, a computer-readable medium can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code having an instruction or data structure form and accessible to a computer.
[0162] Based on the above embodiments, this application also provides a chip, which is used to read a computer program stored in a memory and implement... Figure 6 or Figure 11 The data processing method provided in the illustrated embodiment.
[0163] Based on the above embodiments, this application provides a chip system including a processor for supporting computer devices to implement... Figure 6 or Figure 11 The illustrated embodiment provides a data processing method. In one possible design, the chip system further includes a memory for storing programs and data necessary for the computer device. The chip system may consist of chips or may include chips and other discrete components.
[0164] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0165] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0166] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0167] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0168] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the scope of protection of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A data processing method, characterized in that, The method includes: Divide the data to be compressed into multiple data blocks; Multiple first data blocks among the plurality of data blocks are compressed in parallel, wherein compressing any one of the first data blocks includes: compressing the first data block using a dictionary of a second data block, wherein the second data block is a data block adjacent to the first data block among the plurality of data blocks, and the dictionary of the second data block is used to indicate the correspondence between the content identifier and the compressed data of each data unit in the second data block, wherein the content identifier is determined based on the content of the data unit, and the compressed data is the data required to compress the data unit.
2. The method as described in claim 1, characterized in that, The compression of the first data block using the dictionary of the second data block includes: Based on the dictionary of the second data block, a plurality of data units with the same content in the second data block and the first data block are identified; the plurality of data units include the first data unit in the first data block and the second data unit in the second data block; The first data unit is compressed using the compressed data from the second data unit.
3. The method as described in claim 2, characterized in that, The compressed data of the data unit in the second data block is the location information of the data unit, which is used to indicate the position of the data unit in the second data block; Compressing the first data unit using the compressed data from the second data unit includes: The location information of the second data unit is used as the compressed first data unit; or The position offset information is used as the compressed first data unit, and the position offset information indicates the position offset between the first data unit and the second data unit.
4. The method as described in claim 3, characterized in that, The step of determining multiple data units with identical content in the second data block and the first data block based on the dictionary of the second data block includes: The first data unit is determined based on the content identifier of each data unit in the first data block and the dictionary of the second data block, wherein the content identifier of the first data unit is contained in the dictionary of the second data block; Determine the location information in the dictionary corresponding to the content identifier of the first data unit, and obtain the second data unit at the location indicated by the location information in the second data block, wherein the content of the second data unit is the same as that of the first data unit.
5. The method according to any one of claims 1-4, characterized in that, The second data block is the preceding data unit adjacent to the first data block, or the second data block is the following data unit adjacent to the first data block.
6. The method according to any one of claims 1-5, characterized in that, The method further includes: Add length indication information to the compressed data of each data block in the data to be compressed, the length indication information being used to indicate the length of the compressed data of the data block; Based on the compressed data of each data block and the length indication information of the compressed data of each data block, the compressed data of the data to be compressed is obtained.
7. The method as described in claim 6, characterized in that, The method further includes: Obtain the compressed data of the data to be compressed; Read the multiple length indication information included in the compressed data; The compressed data is divided into multiple compressed data blocks based on the multiple length indication information; The multiple compressed data blocks are decompressed in parallel to obtain the decompressed data.
8. A data processing apparatus, characterized in that, The device includes: The segmentation module is used to divide the data to be compressed into multiple data blocks; A compression module is used to compress multiple first data blocks among the plurality of data blocks in parallel. Compression of any one of the first data blocks includes: compressing the first data block using a dictionary of a second data block, where the second data block is a data block adjacent to the first data block among the plurality of data blocks. The dictionary of the second data block is used to indicate the correspondence between the content identifier and the compressed data for each data unit in the second data block. The content identifier is determined based on the content of the data unit, and the compressed data is used to compress the data required by the data unit.
9. The apparatus as claimed in claim 8, characterized in that, When the compression module compresses the first data block using the dictionary of the second data block, it is specifically used to: determine multiple data units with the same content in the second data block and the first data block according to the dictionary of the second data block; the multiple data units include the first data unit in the first data block and the second data unit in the second data block; The first data unit is compressed using the compressed data from the second data unit.
10. The apparatus as claimed in claim 9, characterized in that, The compressed data of the data unit in the second data block is the location information of the data unit, which is used to indicate the position of the data unit in the second data block; When the compression module compresses the first data unit using the compressed data of the second data unit, it is specifically used to: use the position information of the second data unit as the compressed first data unit; or use the position offset information as the compressed first data unit, wherein the position offset information indicates the position offset between the first data unit and the second data unit.
11. The apparatus as claimed in claim 10, characterized in that, When the compression module determines multiple data units with the same content in the second data block and the first data block based on the dictionary of the second data block, it is specifically used to: determine the first data unit based on the content identifier of each data unit in the first data block and the dictionary of the second data block, wherein the content identifier of the first data unit is included in the dictionary of the second data block; determine the location information in the dictionary corresponding to the content identifier of the first data unit, and obtain the second data unit at the location indicated by the location information in the second data block, wherein the second data unit has the same content as the first data unit.
12. The apparatus according to any one of claims 8-11, characterized in that, The second data block is the preceding data unit adjacent to the first data block, or the second data block is the following data unit adjacent to the first data block.
13. The apparatus according to any one of claims 8-12, characterized in that, The device also includes a data processing module; The data processing module is used to add length indication information to the compressed data of each data block in the data to be compressed, and the length indication information is used to indicate the length of the compressed data of the data block. Based on the compressed data of each data block and the length indication information of the compressed data of each data block, the compressed data of the data to be compressed is obtained.
14. The apparatus as claimed in claim 13, characterized in that, The device also includes a decompression module; The acquisition module is used to acquire the compressed data of the data to be compressed; The decompression module is used to read multiple length indication information included in the compressed data; and to divide the compressed data into multiple compressed data blocks based on the multiple length indication information. The multiple compressed data blocks are decompressed in parallel to obtain the decompressed data.
15. A computing device, characterized in that, It includes at least one processor and at least one memory; wherein the one or more memories store one or more computer programs, the one or more computer programs including instructions that, when executed by the one or more processors, cause the computing device to perform the method as described in any one of claims 1-7.
16. A cluster of devices, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of devices to perform the method as described in any one of claims 1-7.
17. A computer-readable storage medium, characterized in that, The device stores computer-executable instructions for causing a computer to perform the method as described in any one of claims 1-7.
18. A computer program product, characterized in that, It includes computer-executable instructions for causing a computer to perform the method as described in any one of claims 1-7.