Data writing method and system, storage medium, electronic equipment and computer program product

By encoding and storing data streams in the storage client, the memory resource preemption problem caused by processing other services when the CPU processes encoding operations, and improves system performance.

CN120010792AActive Publication Date: 2025-05-16JINAN INSPUR DATA TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510491043.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-05-16
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

In the prior art, the CPU runs software to perform encoding operations tasks in data stream processing, resulting in the CPU need to handle other services, resulting in the preemption of memory resources and affecting the overall performance.

Method used

The first processing unit in the storage client encodes and calculates the data stream corresponding to the data write request, and stores the encoded data stream in the address space of the cache unit, thereby constructing a write control message and sending it to the storage node for writing.

Benefits of technology

It avoids memory resource preemption caused by the CPU when processing other services when processing encoding operations, and improves the overall performance of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010792A_ABST
    Figure CN120010792A_ABST
Patent Text Reader

Abstract

The invention discloses a data writing method and system, a storage medium, electronic equipment and a computer program product, and relates to the field of storage systems.The data writing method includes the steps that coding calculation is conducted on a first data stream through a first processing unit in a storage client side, and the first data stream obtained after coding calculation is stored in a cache unit in the first processing unit; according to the first address of the first solid state disk to which the first data stream is to be written and the first identifier of the second address of the cache unit where the first data stream is located, the first storage node containing the first solid state disk writes the first data stream after coding calculation into the first solid state disk. That is to say, according to the embodiment of the invention, the coding calculation task in the data stream processing is carried out through the storage client instead of the CPU, so that the problem of preemption of memory resources caused by the fact that the coding calculation task in the data stream processing is executed through CPU running software and other services need to be processed when the CPU processes the coding calculation in related technologies is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of storage systems, and in particular to a data writing method and system, a storage medium, an electronic device, and a computer program product. Background Art

[0002] Distributed storage often uses a multi-copy backup mechanism and an erasure code mechanism to ensure the reliability of stored data. Because erasure code technology is accompanied by a large number of matrix operations and cross-node data transmission, it will cause a lot of computing and transmission overhead to the system. However, the current mainstream usage is to use the CPU to run software to perform erasure encoding and decoding calculations, relying on the computing power of the central processing unit (CPU) rather than a dedicated erasure hardware accelerator. When the input / output (I / O) bandwidth of the system is large, erasure calculations will occupy more CPU resources. At the same time, the CPU also has to process the business of other software modules. The CPU and memory resource preemption caused by erasure calculations will inevitably affect the overall performance.

[0003] Therefore, in the related art, the CPU runs software to execute encoding operation tasks in data stream processing. When processing encoding operations, the CPU also needs to process other services, resulting in the preemption of memory resources, which has not been effectively solved. Summary of the invention

[0004] The present application provides a data writing method and system, a storage medium, an electronic device and a computer program product, so as to at least solve the problem of memory resource preemption caused by executing encoding operation tasks in data stream processing through CPU running software in the related technology, and the CPU also needs to process other services when processing encoding operations.

[0005] The present application provides a data writing method, which is applied to a storage client, comprising: upon receiving a data write request, performing a first calculation type of encoding calculation on a first data stream corresponding to the data write request by a first processing unit in the storage client, and storing the first data stream after the encoding calculation in a second address space of a cache unit in the first processing unit, wherein the first calculation type includes at least one of the following: data erasure, data encryption, and the first calculation type is the calculation type of the first data stream; constructing a write control message according to a first address of a first solid-state hard disk corresponding to the first data stream and a first identifier of a second address corresponding to the second address space, wherein the first address is an address corresponding to a first address space in the first solid-state hard disk for writing the first data stream after the encoding calculation; sending the write control message to a first storage node including the first solid-state hard disk, so that the first storage node obtains the first data stream after the encoding calculation according to the write control message, and writes the first data stream after the encoding calculation to the first solid-state hard disk.

[0006] The present application also provides a data writing system, comprising: a storage client and a target storage node, wherein: the storage client is used to, upon receiving a data write request, perform a decoding calculation of a first calculation type on a first data stream corresponding to the data write request through a first processing unit in the storage client, and store the decoded and calculated first data stream in a second address space of a cache unit in the first processing unit; construct a write control message according to a first address of a first solid-state hard disk corresponding to the first data stream and a first identifier of a second address corresponding to the second address space, and send the write control message to a first storage node including the first solid-state hard disk, wherein the first calculation type includes at least one of the following: data erasure, data encryption, the first calculation type is the calculation type of the first data stream, and the first address is an address corresponding to a first address space in the first solid-state hard disk for writing the decoded and calculated first data stream; the target storage node comprises: the first storage node is used to obtain the decoded and calculated first data stream according to the write control message, and write the decoded and calculated first data stream into the first solid-state hard disk.

[0007] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned data writing methods when executing the computer program.

[0008] The present application also provides a computer-readable storage medium, in which a computer program is stored, wherein when the computer program is executed by a processor, the steps of any of the above-mentioned data writing methods are implemented.

[0009] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned data writing methods when executed by a processor.

[0010] Through this application, since the first data stream is encoded and calculated by the first processing unit in the storage client, and the first data stream after encoding and calculation is stored in the cache unit in the first processing unit; then according to the first address of the first solid-state hard disk to which the first data stream is to be written and the first identifier of the second address of the cache unit where the first data stream is located, the first storage node including the first solid-state hard disk can write the first data stream after encoding and calculation into the first solid-state hard disk. In other words, the embodiment of the present application performs encoding and calculation tasks in data stream processing by the storage client instead of the CPU, thereby solving the problem of memory resource preemption caused by the CPU running software to execute encoding and calculation tasks in data stream processing in the related art. When the CPU processes encoding and calculation, it also needs to process other services, thereby avoiding the problem of memory resource preemption. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0012] Figure 1 It is a hardware structure block diagram of a computer terminal of a data writing method according to an embodiment of the present application;

[0013] Figure 2 is a flow chart of a data writing method according to an embodiment of the present application;

[0014] Figure 3 It is a flow chart of a method for performing data stream processing by a PCIE expansion card in the related art;

[0015] Figure 4 is an architecture diagram corresponding to a write processing flow according to an optional embodiment of the present application;

[0016] Figure 5 is a flowchart of write data processing according to an optional embodiment of the present application;

[0017] Figure 6 is an architecture diagram corresponding to a read processing flow according to an optional embodiment of the present application;

[0018] Figure 7 is a flowchart of a read data process according to an optional embodiment of the present application;

[0019] Figure 8 It is a framework diagram of a data writing system according to an embodiment of the present application. DETAILED DESCRIPTION

[0020] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0021] It should be noted that, in the description of this application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects, and are not used to describe a specific order or sequence.

[0022] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0023] In conjunction with the specific application environment architecture or the specific hardware architecture on which the execution of the data writing method depends, the specific application environment architecture or the specific hardware architecture is described herein.

[0024] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Taking running on a computer terminal as an example, Figure 1 1 is a hardware structure block diagram of a computer terminal of a data writing method according to an embodiment of the present application. Figure 1 As shown, the computer terminal may include one or more ( Figure 1 Only one is shown in the figure) a processor 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data, wherein the above-mentioned computer terminal may also include a transmission device 106 and an input and output device 108 for communication functions. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above-mentioned computer terminal. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown.

[0025] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the method for determining the interactive state in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, to implement the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories can be connected to the computer terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0026] The transmission device 106 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of a computer terminal. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0027] The embodiment of the present application provides a data writing method. The following is an explanation of the technical terms involved in the embodiment of the present application:

[0028] Data erasure correction (or erasure code correction), where erasure coding (EC) can split data into fragments, expand and encode redundant data blocks, and store them in different locations, such as disks, storage nodes, or other geographical locations.

[0029] Figure 2 is a flow chart of a data writing method according to an embodiment of the present application, which can be applied to Figure 1 In a computer terminal, such as Figure 2 As shown, the process includes the following steps:

[0030] Step S202: upon receiving a data write request, performing encoding calculation of a first calculation type on a first data stream corresponding to the data write request by a first processing unit in the storage client, and storing the first data stream after the encoding calculation in a second address space of a cache unit in the first processing unit, wherein the first calculation type includes at least one of the following: data erasure and data encryption, and the first calculation type is a calculation type of the first data stream;

[0031] Among them, the above-mentioned first processing unit can be a data processing unit (DPU) in the storage client, or a smart network card in the storage client, or a combination of a DPU and a smart network card in the storage system, expressed as DPU / smart network card.

[0032] Step S204, constructing a write control message according to a first address of the first solid-state hard disk corresponding to the first data stream and a first identifier of a second address corresponding to the second address space, wherein the first address is an address corresponding to a first address space in the first solid-state hard disk for writing the first data stream after the encoding calculation;

[0033] Step S206, sending the write control message to the first storage node including the first solid state drive, so that the first storage node obtains the first data stream after the encoding calculation according to the write control message, and writes the first data stream after the encoding calculation to the first solid state drive.

[0034] Through the data writing method of the present application, the first data stream is encoded and calculated by the first processing unit in the storage client, and the first data stream after encoding and calculation is stored in the cache unit in the first processing unit; then, according to the first address of the first solid-state hard disk to be written by the first data stream and the first identifier of the second address of the cache unit where the first data stream is located, the first storage node including the first solid-state hard disk can write the first data stream after encoding and calculation into the first solid-state hard disk. In other words, the embodiment of the present application performs the encoding and calculation tasks in the data stream processing by the storage client instead of the CPU, thereby solving the problem of preemption of memory resources caused by the CPU running software to execute the encoding and calculation tasks in the data stream processing in the related technology. When the CPU processes the encoding and calculation, it also needs to process other services, thereby avoiding the problem of preemption of memory resources.

[0035] Optionally, before constructing a write control message according to the first address of the first solid-state hard disk corresponding to the first data stream and the first identifier of the second address corresponding to the second address space in the above step S204, the method further includes: obtaining the second data stream corresponding to the data write request, and performing data preprocessing operations on the second data stream to obtain the first data stream; determining whether there is a reserved address space corresponding to the first data stream in multiple solid-state hard disks according to a distributed data layout algorithm, wherein each storage node includes a solid-state hard disk, and the multiple storage nodes include: the first storage node, the reserved address space is an address space reserved for writing the first data stream after the encoding calculation; in determining whether there is a reserved address space corresponding to the first data stream in the multiple solid-state hard disks In a case where the reserved address space exists in one or more second solid-state hard disks, the addresses corresponding to the one or more second solid-state hard disks are determined as the first address, wherein the first solid-state hard disk includes: the one or more second solid-state hard disks; in a case where it is determined that the reserved address space does not exist in the multiple solid-state hard disks, consistency negotiation is performed with the multiple storage nodes so that the multiple storage nodes select one or more third solid-state hard disks from the multiple solid-state hard disks and allocate address space for the first data stream in the one or more third solid-state hard disks; the addresses corresponding to the one or more third solid-state hard disks are determined as the first address, wherein the first solid-state hard disk includes: the one or more third solid-state hard disks.

[0036] It is understandable that before constructing the write control message, it is necessary to obtain the first data stream and determine the first address of the address space where the first data stream is to be stored in the first solid state drive. Specifically:

[0037] When the storage client receives a data write request, it first processes the data write request to form a second data stream. The second data stream is the original, unprocessed data stream. Then, the storage client preprocesses the second data stream, such as data striping, metadata processing, etc., to obtain the first data stream. The first data stream is an optimized data stream that has been preliminarily processed and is ready for erasure coding (i.e., coding calculation of erasure calculation) or encryption (i.e., coding calculation of encryption calculation).

[0038] According to the distributed data layout algorithm, check whether there is address space reserved for the first data stream in multiple solid-state drives (SSDs). Each storage node contains an SSD, and there may be multiple storage nodes. The reserved address space refers to the space prepared in advance for storing the encoded data to reduce the delay when writing data.

[0039] If the reserved address space is not found in the SSD, the storage system client will negotiate consistency with multiple storage nodes. The purpose of consistency negotiation is to determine which SSD (third SSD) in the storage node can allocate address space for the first data stream. Consistency negotiation ensures that all nodes participating in storage can reach a consensus on the storage location of the first data stream, avoiding conflicts and inconsistencies when writing data.

[0040] After consistency negotiation, the address corresponding to the determined first solid state hard disk is determined as the first address.

[0041] Optionally, the first calculation type in the encoding calculation of the first calculation type performed by the first processing unit in the storage client on the first data stream corresponding to the data write request in step S202 may include: data erasure correction, data encryption, data deduplication and other situations. The following is a process of performing encoding calculation when it is determined that the first calculation type is data erasure correction and data encryption:

[0042] (1) When it is determined that the first calculation type is data erasure correction, calling a data erasure correction calculation unit in the first processing unit through the first processing unit, and encoding the first data stream according to the data erasure correction calculation unit to generate first erasure correction data and first verification data corresponding to the first data stream; and generating the first data stream after the encoding calculation according to the first erasure correction data and the first verification data.

[0043] It is understandable that the coding calculation of data erasure correction can be implemented by the first processing unit, and the specific steps are:

[0044] Determine whether the calculation type is data erasure: Determine whether the current data processing request involves data erasure operation. Data erasure is a data protection technology that divides the original data into multiple data blocks and calculates and generates additional check data blocks. Even if some data blocks are lost or damaged, the original data can be restored through the remaining intact data blocks and check data blocks.

[0045] Calling the data erasure correction calculation unit: Once the calculation type is confirmed to be data erasure correction, the storage client calls the preset data erasure correction calculation unit through the first processing unit.

[0046] Perform coding calculation: After receiving the first data stream from the storage client, the data erasure calculation unit divides the first data stream and performs coding calculation on the divided data blocks according to the erasure coding algorithm. This process generates first erasure data and first verification data, where the first erasure data contains the divided data blocks, and the first verification data is an additional verification block generated for data recovery.

[0047] Generate a data stream after coding calculation: After completing the coding calculation, the data erasure correction calculation unit combines the first erasure correction data and the first verification data to generate a first data stream after coding calculation.

[0048] (2) When it is determined that the first calculation type is data encryption, the data encryption calculation unit in the first processing unit is called by the first processing unit, and an encryption algorithm and a first key corresponding to the first data stream are determined according to the data encryption calculation unit; the encryption calculation unit encrypts the first data stream according to the encryption algorithm and the first key to generate the first data stream after the encoding calculation.

[0049] It is understandable that, when the calculation type is determined to be data encryption, data encryption processing can be implemented by the first processing unit. The specific steps are:

[0050] Determine the calculation type as data encryption: Determine whether the calculation type of the current data operation is data encryption. Data encryption is an important means of protecting data security. It converts raw data into ciphertext to prevent unauthorized access to data during transmission or storage.

[0051] Calling the data encryption calculation unit: If it is determined that the current operation requires data encryption, the storage client will call the built-in data encryption calculation unit through the first processing unit. The data encryption calculation unit is hardware-level and is specially designed to execute data encryption algorithms. It can complete complex data encryption tasks in a very short time, improving data security and processing efficiency.

[0052] Determine the encryption algorithm and key: After receiving the first data stream, the data encryption calculation unit determines the applicable encryption algorithm and first key according to preset or client-specified conditions.

[0053] Perform encryption processing: The data encryption calculation unit uses a determined encryption algorithm and key to perform encryption processing on the first data stream. The encryption process converts the original data into ciphertext data that cannot be directly recognized, and generates the first data stream after encoding calculation.

[0054] Optionally, after storing the first data stream after encoding calculation in the second address space of the cache unit in the first processing unit in the above step S202, the method further includes: determining the first identifier corresponding to the second address space when the first processing unit determines that the first data stream after encoding calculation has been stored in the second address space; sending the completion status and the first identifier corresponding to the first data stream after encoding calculation to the storage client through the write preparation interface corresponding to the first processing unit, wherein the completion status is used to indicate that the first data stream after encoding calculation has been stored in the second address space.

[0055] It is understandable that after the data processing is completed, the storage status and key identification information can be fed back to the storage client through the first processing unit. Specifically:

[0056] Confirming data storage: After the first processing unit completes the encoding calculation (such as erasure coding or encryption operation) of the first data stream, and the first data stream after the encoding calculation has been successfully stored in the second address space in the cache unit, the first processing unit will confirm that the storage status of the first data stream after the encoding calculation is successful storage.

[0057] Determine an identifier: After confirming that the data storage is successful, the first processing unit determines a first identifier associated with the storage location (second address). The first identifier is key information for uniquely identifying the storage location of the first data stream after the encoding calculation in the cache unit in the first processing unit.

[0058] Sending completion status and first identifier: The first processing unit sends the completion status of data processing and the first identifier to the storage client through a predefined write preparation interface. The completion status is usually a Boolean value or a status code that clearly indicates whether the first data stream after the encoding calculation has been successfully stored in the second address space.

[0059] The client receives feedback: After receiving the completion status and the first identifier sent by the first processing unit, the storage system client determines whether the data storage is successful according to the completion status. If the storage is successful, the storage client can use the first identifier to perform subsequent data operations, such as writing data, reading data, retrieving data, etc.; if it fails, the storage client needs to retry or handle the exception according to the error information.

[0060] Optionally, the above-mentioned step S204 constructs a write control message according to the first address of the first solid-state hard disk corresponding to the first data stream and the first identifier of the second address corresponding to the second address space, including: obtaining the first data information of the first data stream after the encoding calculation, and determining the message format of the write control message to be constructed, wherein the message format includes: header information and one or more data segments of the write control message, the first data information includes at least one of the following: the data stream size of the first data stream after the encoding calculation, the creation time of the first data stream after the encoding calculation, the data version corresponding to the first data stream after the encoding calculation, and the data write instruction corresponding to the first data stream after the encoding calculation; filling the header information according to the first address and the first identifier, and filling the one or more data segments according to the first data information, so as to construct the write control message according to the filled header information and the filled one or more data segments.

[0061] It is understandable that the process of constructing a write control message may include:

[0062] Obtaining first data information: After the first processing unit completes the encoding calculation (such as erasure correction and encryption) of the first data stream, the first data information of the first data stream after the encoding calculation is obtained. The first data information may include the size of the data stream, the creation time, the data version, or the data writing instruction.

[0063] Determine the message format: Determine the message format of the write control message to be constructed. The write control message usually consists of header information and one or more data segments. The header information contains metadata required for control and location of data, while the data segment contains the actual data information or instructions.

[0064] Filling header information: When constructing a message, the header information will be filled according to the first address (i.e., the target address of the storage device) and the first identifier (such as the Key value, which is used to uniquely identify the storage location).

[0065] Filling the data segment: The data segment is filled according to the first data information. The information in the data segment is used to describe the characteristics and storage requirements of the data stream, such as the data size to help determine the storage space, the creation time or data version for data version control and consistency management, and the data write instruction to directly guide the specific execution of the storage operation.

[0066] Constructing a write control message: Constructing a complete write control message based on the filled header information and data segment. Once the write control message is constructed, it will be sent to the first storage node to notify the first storage node to store data.

[0067] Optionally, the above-mentioned data writing method may include not only a data writing process, but also a data reading process, specifically: when a data reading request is received, second data information of the third data stream corresponding to the data reading request is obtained through the first processing unit, wherein the second data information includes at least one of the following: a third address of the second solid-state hard disk storing the third data stream, and calculation parameters corresponding to the third data stream, and the calculation parameters include at least one of the following: data correction and erasure parameters, data encryption parameters, and data deduplication parameters; a third address space is allocated to the third data stream in the cache unit through the first processing unit, and a second identifier of the fourth address corresponding to the third address space is sent to the storage client; and the data reading operation corresponding to the data reading request is performed according to the third address and the second identifier.

[0068] Wherein, the data read operation corresponding to the data read request is performed according to the third address and the second identifier, including: constructing a read control message according to the third address and the second identifier, and sending the read control message to the second storage node corresponding to the second solid-state hard disk, so that the second storage node obtains the third data stream in the second solid-state hard disk according to the read control message, and sends the third data stream to the third address space; controlling the first processing unit to perform a second calculation type of decoding calculation on the third data stream stored in the third address space according to the calculation parameters, and updating the data in the third address space to the third data stream after decoding calculation, so that the target object sending the data read request reads the third data stream after decoding calculation, wherein the second calculation type includes at least one of the following: data erasure, data encryption, and the second calculation type is the calculation type of the third data stream.

[0069] It is understandable that the storage client can also process data read requests, specifically:

[0070] Receiving a data read request: When the storage client receives a data read request sent by the target object, the first processing unit processes the data read request. The data read request usually includes the third data stream, the read position (offset), and the length information.

[0071] Obtaining read data information: The first processing unit analyzes the data read request and extracts the second data information therefrom, including the third address of the second solid-state hard disk storing the third data stream (i.e., the storage address of the data in the SSD) and the calculation parameters corresponding to the third data stream. The calculation parameters may be data erasure parameters, data encryption parameters, or data deduplication parameters, depending on the processing type of the encoding calculation performed when the data is written.

[0072] Allocating cache address space: The first processing unit allocates a third address space for the third data stream in the cache unit. The third address space is used to temporarily store data read from the second solid-state hard disk and subsequent decoding calculation results. At the same time, the second identifier of the fourth address corresponding to the allocated third address space is sent to the storage client.

[0073] Construct and send a read control message: The storage client constructs a read control message according to the received second identifier and third address, and sends the read control message to the second storage node where the second solid-state hard disk storing the third data stream is located. The read control message contains information such as a read instruction, a key value, and a target device address (i.e., a third address), and is used to instruct the second storage node to read data from the second solid-state hard disk and transmit it to the third address space.

[0074] Execute data reading: After the second storage node receives the read control message, it reads the third data stream from the third address of the second solid-state hard disk according to the instructions and information contained in the read control message, and directly writes the data to the third address space through the second processing unit in the second storage node, avoiding additional data movement of the CPU and memory.

[0075] Decoding calculation: After the third data stream is read into the third address space, the first processing unit will perform decoding calculation of the second calculation type on the third data stream according to the calculation parameters corresponding to the data read request. If the data is erasure coded when written, then erasure decoding calculation will be performed at this time; if the data is encrypted when written, then data decryption calculation will be performed at this time. The third data stream after decoding calculation will be updated to the third address space.

[0076] Completing the read operation: After the decoding calculation is completed, the decoded data stored in the third address space can be directly read by the target object.

[0077] Optionally, before obtaining the second data information of the third data stream corresponding to the data read request through the first processing unit, the method also includes: obtaining a data stream identifier corresponding to the third data stream according to the data read request, and determining distribution information corresponding to the third data stream according to the data stream identifier, wherein the distribution information is used to indicate the distribution of the third data stream in multiple storage nodes; parsing the distribution information, and determining the second storage node storing the third data stream among the multiple storage nodes according to the parsed distribution information; and determining the address corresponding to the second solid-state hard disk in the second storage node as the third address.

[0078] It is understandable that this description reveals that in the technical solution of the present invention, the storage client needs to determine the specific location (ie, the third location) of the third data stream to be read in the solid state drive. Specifically:

[0079] Obtaining a data stream identifier: When the storage client receives a data read request, it extracts a data stream identifier corresponding to the third data stream from the data read request. The data stream identifier is usually a unique identifier of the data stream in the storage system, such as a hash value of a file or an ID of an object, which is used to accurately identify and locate data.

[0080] Determine data distribution information: Using the data stream identifier, the storage client queries the metadata of the storage system to determine the distribution information of the third data stream. The distribution information describes the storage layout of the third data stream in multiple storage nodes, including which storage node each data block after data striping is stored in, and the specific storage location in the storage node (such as the address in the SSD).

[0081] Parsing distribution information: The distribution information obtained by parsing the storage client is used to identify each second storage node storing the third data stream and the storage location inside the second storage node. The parsing process may involve converting the distribution information into a format that is easy to process, such as mapping the node ID and address in the distribution information to the actual network address and storage device address.

[0082] Determine the third address: From the parsed distribution information, the storage client can determine the second storage node storing the third data stream. The second storage node is the node that the data read request needs to directly access. It contains the third address required to read the data stream, that is, the specific location of the third data stream in the second solid-state drive. This specific location is the third address.

[0083] Optionally, obtaining second data information of the third data stream corresponding to the data read request through the first processing unit includes: calling a read preparation interface corresponding to the first processing unit, and sending the second data information to the first processing unit according to the read preparation interface.

[0084] It is understandable that the data interaction process between the storage client and the first processing unit can be implemented through a read preparation interface.

[0085] Optionally, after updating the data in the third address space to the third data stream after decoding and calculation, the method also includes: determining an application cache address corresponding to the data read request, wherein the application cache address is an address corresponding to the application cache space in the target object for reading the third data stream after decoding and calculation; calling a read interface corresponding to the first processing unit, and sending the second identifier and the application cache address to the first processing unit according to the read interface; upon receiving the second identifier and the application cache address, the first processing unit sends the third data stream after decoding and calculation to the application cache space corresponding to the application cache address, so that the target object reads the third data stream after decoding and calculation.

[0086] It is understandable that the application cache address can be determined, and the first processing unit can be used to implement the process of quickly transmitting the decoded and calculated third data stream to the address space corresponding to the application cache address after data decoding. The specific steps include:

[0087] Determine the application cache address: When the storage client receives a data read request, it determines that the target object that sends the data read request needs to read the data to its internal application cache address. The application cache address refers to the specific address of the application cache space in the target object used to read and process the third data stream after decoding and calculation.

[0088] Calling the read interface: The storage client can call the read interface corresponding to the first processing unit. The read interface is a hardware interface for controlling data reading and transmission. The storage client sends the second identifier (i.e., the Key value in the third address space) and the application cache address to the first processing unit through the read interface.

[0089] Sending the second identifier and the application cache address: The storage client sends the second identifier and the application cache address to the first processing unit through the read interface. The second identifier is used to locate the decoded and calculated third data stream in the cache unit of the first processing unit, and the application cache address is used to indicate the application cache space inside the target object.

[0090] Data transfer to application cache: After the first processing unit receives the second identifier and the application cache address, it immediately starts the data transfer process. The decoded and calculated third data stream is directly written from the address space corresponding to the third address to the cache space corresponding to the application cache address of the target object through DMA (Direct Memory Access). DMA technology allows data to be transferred directly between hardware without CPU intervention, which can significantly increase the data transfer speed and reduce the use of CPU and memory resources.

[0091] The target object reads data: Once the data transmission is completed, the target object can directly read the decoded and calculated third data stream from the application cache space.

[0092] Optionally, the first processing unit performs a decoding calculation of a second calculation type on the third data stream stored in the third address space according to the calculation parameters, including: when the second calculation type is data erasure correction, calling a data erasure correction calculation unit in the first processing unit through the first processing unit, so that the data erasure correction calculation unit determines the second erasure correction data and second verification data corresponding to the third data stream according to the calculation parameters; and performing a decoding calculation on the third data stream according to the second erasure correction data and the second verification data.

[0093] It can be understood that the first processing unit can perform data erasure decoding calculation on the third data stream, specifically:

[0094] Calling the data erasure correction calculation unit: When the first processing unit determines that the second calculation type corresponding to the data read request is data erasure correction decoding, calling its own built-in data erasure correction calculation unit. The data erasure correction calculation unit is part of the hardware acceleration module and is specifically used to perform decoding calculations of erasure codes to restore original data.

[0095] Determine erasure data and check data: After receiving the call, the data erasure calculation unit determines the second erasure data and the second check data in the third data stream according to the pre-provided calculation parameters (including but not limited to the erasure algorithm, data block size, redundancy ratio, etc.). Erasure data refers to data blocks that are dispersedly stored after erasure coding, while check data refers to additional data blocks generated for data redundancy and recovery.

[0096] Perform decoding calculation: After determining the second erasure correction data and the second verification data, the data erasure correction calculation unit performs decoding calculation on the third data stream to restore the original third data stream and obtain the third data stream after decoding calculation.

[0097] Update decoded data: After the decoding calculation is completed, the third data stream stored in the cache unit in the first processing unit will be updated to the decoded third data stream.

[0098] Optionally, the first processing unit performs a decoding calculation of a second calculation type on the third data stream stored in the third address space according to the calculation parameters, including: when the second calculation type is data encryption, calling the data encryption calculation unit in the first processing unit through the first processing unit, and obtaining a second key and a second encryption algorithm for encrypting the third data stream according to the data encryption calculation unit; determining a decryption algorithm corresponding to the second encryption algorithm, and decrypting the third data stream according to the decryption algorithm and the second key to generate the third data stream after the decoding calculation.

[0099] It can be understood that the first processing unit can also perform a decoding operation corresponding to the data encryption, that is, decrypt the third data stream. Specifically:

[0100] Calling the data encryption calculation unit: After determining that the second calculation type of the data read request is data decryption, the first processing unit calls its built-in data encryption calculation unit.

[0101] Obtaining encryption information: Through the data encryption calculation unit, the first processing unit can obtain the second key and the second encryption algorithm used when encrypting the third data stream.

[0102] Determine the decryption algorithm: The first processing unit determines the corresponding decryption algorithm according to the acquired second encryption algorithm. The encryption algorithm and the decryption algorithm are one-to-one.

[0103] Decryption processing: The first processing unit uses the determined decryption algorithm and the second key to decrypt the encrypted third data stream pulled from the storage node, and restores the data to a plaintext state. The decryption process may include key expansion, data block decoding, and algorithm-specific reverse operations.

[0104] Generate the decoded third data stream: After decryption, the originally encrypted third data stream will be restored to an unencrypted data stream, that is, generate the decoded third data stream. The decoded third data stream will be stored in the cache unit of the first processing unit.

[0105] In order to better understand the process of the above-mentioned data writing method, the implementation process of the above-mentioned data writing method is described below in combination with an optional embodiment, but it is not used to limit the technical solution of the embodiment of the present application.

[0106] The following is an explanation of the technical terms that may appear in the optional embodiments of this application:

[0107] 1) Single-stream bandwidth: the read and write bandwidth of a single thread of an application;

[0108] 2) High performance computing (HPC): refers to the computing mode that uses supercomputers and parallel processing technology to solve large-scale complex computing problems in science, engineering, business and other fields;

[0109] 3) Data striping: Split a file into many small data blocks, i.e. data striping, and then distribute the data stripes to each node of distributed storage;

[0110] 4) Kernel state: In the kernel running state of the computing processor, modern operating systems such as Linux and Windows have kernel state, which is used to run the operating system management process, resource scheduling, memory management and other processes;

[0111] 5) User state: the user running state of the computing processor. Modern operating systems such as Linux and Windows all have user state, which is used to run user processes.

[0112] 6) Data copy: In modern computer processors, memory addresses are divided into user state and kernel state. User state is the address space of user processes, and each user process is independently allocated; the address space of kernel state belongs to the kernel process of the operating system and cannot be used by user state.

[0113] 7) Computer Operating System (OS): It is a computer program that manages computer hardware and software resources, controls the operation of the computer, manages all system resources, and provides users with a friendly operating interface to perform various tasks.

[0114] 8) Kernel client refers to the client of distributed storage, which is deployed in the user state of the client host OS to realize the interconnection access between the client host and the distributed storage system;

[0115] 9) Virtual file system interface (VFS for short). VFS implements the file interface of the operating system, and various file systems implement statistical interface docking;

[0116] 10) Memory Management Unit (MMU): Responsible for the virtual address space management of each process in the operating system, and responsible for the mapping and management of the address space and physical memory of the user-mode process.

[0117] The optional embodiments of the present application relate to the field of storage systems, specifically subdivided into the field of independent client performance optimization. A data stream acceleration method for a storage system is proposed. It is proposed to use computational erasure correction technology in the process of network data stream processing, merge data stream processing and hardware acceleration, and improve the processing performance of a single stream.

[0118] Single-stream is a relatively demanding application scenario in the storage field, and is mainly suitable for high-performance computing applications, especially data writing applications such as satellites, astronomical telescopes, and cryo-electron microscopes. With the increasing amount of data from satellites and astronomical telescopes, distributed storage clients are required to provide higher reception performance, that is, single-stream performance.

[0119] Generally speaking, distributed storage is a storage system composed of multiple storage nodes connected through a network to achieve a unified namespace, which can be accessed in parallel by clients. Distributed storage single-threaded performance is mainly achieved by striping files and sending them to multiple nodes in parallel for parallel processing.

[0120] In order to solve the problem of memory resource preemption caused by the CPU running software to execute encoding calculation tasks in data stream processing, the CPU needs to process other services while processing encoding calculations. A method for realizing data stream processing through a peripheral component interconnection extension (Peripheral Component Interconnect Express, referred to as PCIE) expansion card is also proposed in the related technology. Specifically: the PCIE expansion card is directly connected to the computing node, and the data is processed by transmitting the data to the acceleration card for processing, and then returned to the original processing flow. The advantage of this processing is that the acceleration card and the original processing flow are decoupled and plug-and-play. The problem is that the interactive resources are consumed a lot, which limits the processing performance of the acceleration card.

[0121] Figure 3 It is a flow chart of a method for processing data streams by a peripheral component interconnect express (PCIE) expansion card in the related art, such as Figure 3 As shown:

[0122] Step S301, the application process in user mode sends a request to read or write data;

[0123] Step S302, after receiving the processed read / write data request sent by the VFS interface in the kernel state, performing other processing operations on the file corresponding to the processed read / write request, for example: preprocessing operation;

[0124] Step S303, copy and migrate the data to migrate the data to the PCIE expansion card;

[0125] Step S304, performing redundancy correction and erasure calculation on the data through the PCIE expansion card, and sending the data obtained after iosa to the acceleration card;

[0126] Step S305: Send the data processed by the acceleration card to the backend storage system, that is, any one or more storage nodes among the storage nodes 31-36.

[0127] The optional embodiment of the present application proposes a method for accelerating the data flow of a storage system, which realizes the fusion of the main data flow processing and the erasure calculation sub-processing through the coordinated processing of the data flow processing and the erasure calculation sub-processing of the distributed storage system client, avoids the data migration and interaction from the main data to the erasure sub-processing flow, and realizes hardware acceleration within the data flow. Specifically:

[0128] The method for accelerating the data flow of the storage system involves three parts: a storage system client (i.e., storage client), a data processing unit (DPU) / intelligent network card (i.e., a first processing unit and a second processing unit), and a storage system node (i.e., a target storage node). These three parts can execute the data writing process and the data reading process. Specifically:

[0129] Figure 4 This is an architecture diagram corresponding to the write processing flow according to an optional embodiment of the present application. Specifically, the write processing flow may include the following process steps:

[0130] (1) The storage system client completes the metadata processing, data striping and addressing operations of the file system client, negotiates the consistency of the write address space, interacts with the DPU / Smart NIC for data stream processing, and transmits the write data stream to the DPU / Smart NIC for calculation through the hardware acceleration interface;

[0131] (2) The client's DPU / Smart NIC contains an FPGA that can perform accelerated calculations on the write data stream for erasure correction / encryption (erasure correction calculations can generate erasure codes). After the calculation is completed, it only returns the address of the cache unit in the client's DPU / Smart NIC where the data calculated by the client is stored (i.e. Figure 4The DPU internal address in ), does not return the data stream (where, Figure 4 Key: addr is the address identifier corresponding to the address in DPU);

[0132] An optional embodiment of the present application can implement data communication between the DPU and storage nodes, other DPUs or network devices through a communication module.

[0133] (3) The storage client and the storage node complete the coordination of the write control message of the data flow, that is, the storage client sends a write control message to the storage node, so that the storage node pulls the data in the cache unit in the DPU smart network card through the hardware acceleration interface based on the write control message.

[0134] (4) Through the DPU smart network card and storage medium device (SSD) in the storage node, the direct memory access (DMA) writing of data is completed.

[0135] Among them, the hardware acceleration interfaces involved in the data writing process may include: preprocessing interface: transmitting data, erasure calculation parameters, returning the DPU key and cache address; pulling interface: storage node DPU / intelligent pulling data, etc.

[0136] Among them, NIC is used as a communication bridge between the storage node and DPU / Smart NIC when the DPU and Smart NIC are used in combination to receive and send data streams.

[0137] In an alternative embodiment, Figure 5 is a flowchart of write data processing according to an optional embodiment of the present application, such as Figure 5 As shown:

[0138] Step S501, the user-mode application performs a write operation on the specified file, offset and length through system calls and VFS, and transfers the data to the storage system client cache;

[0139] Step S502: the storage system client (ie, storage client) completes the processing of metadata and data, and finds the target node and target device according to the distributed data layout algorithm.

[0140] Step S503: before sending the write request, the storage system client determines whether there is a reserved destination address space, and if so, uses it directly; if there is no reserved address space, it negotiates consistency with the target node (i.e., the first storage node) and the target device, and allocates the destination device address space (i.e., the first address space);

[0141] Step S504: The storage system client calls the hardware acceleration interface - write preparation interface of the DPU / intelligent network card (i.e., the first processing unit) to pass the cached data (i.e., the first data stream), the address length of the cached data, the calculation type of the cached data (i.e., the first calculation type), etc. to the DPU / intelligent network card. The calculation type may include erasure correction, encryption, deduplication, etc.

[0142] Step S505: The DPU / intelligent network card performs erasure / encryption coding calculation on the data cache that has been input according to the calculation type (including erasure, encryption, etc.) and data information (if it is erasure calculation, the erasure processing interface can be called to send the data related information to the erasure calculation module in the DPU / intelligent network card; if it is encryption calculation, the data related information can be sent to the encryption calculation module), and writes the processed data into the data cache (i.e., cache unit) in the DPU / intelligent network card;

[0143] The erasure calculation module is a data erasure calculation unit, and the encryption calculation module is a data encryption calculation unit.

[0144] Furthermore, the DPU / intelligent network card returns the result to the storage system client through the hardware acceleration interface-write preparation interface, and the interface returned by the interface includes: the completion status of the interface, and the key value (i.e., the first identifier) ​​of the address of the processed data in the data cache, and the key value is the key representing the address of the processed data cached by the cache unit in the DPU / intelligent network card.

[0145] The network module is responsible for sending and receiving data packets, for example, receiving cache data, the address length of cache data, and the calculation type of cache data sent by the storage system client.

[0146] The storage system client can also pre-apply for write space in the DPU's data cache to store processed data.

[0147] Step S506: The storage client sends the calculated data cache key (i.e., the first identifier) ​​to the storage node in the form of a write control message according to the address of the target node and the target device (i.e., the first address). Figure 5 One or more of the storage nodes 51-53 in the storage node may be the first storage node);

[0148] In step S507, the storage node receives the write control message sent by the storage client, calls the DPU / smart NIC hardware acceleration interface - write pull interface, and transmits the Key to the DPU / smart NIC of the storage node; then the DPU / smart NIC of the storage node pulls data according to the Key and the second address; the DPU / smart NIC of the storage client receives the data pull request, and transmits the data in the data cache (i.e., the first data stream after encoding and calculation) to the DPU / smart NIC of the storage node; the DPU / smart NIC of the storage node writes the first data stream after encoding and calculation into the corresponding address space in the SSD medium.

[0149] The storage node responds to the write control message successfully; after receiving all responses, the storage system client completes the write request and returns the application's system call.

[0150] Figure 6 This is an architecture diagram corresponding to the read processing flow according to an optional embodiment of the present application. Specifically, the read processing flow may include the following process steps:

[0151] (1) The storage system client completes the metadata processing, data striping and addressing operations of the file system client. The data stream processing pre-applies for the cache address and key (second identifier) ​​of the DPU / Smart NIC through the hardware acceleration interface.

[0152] (2) The storage system client sends the pre-applied read cache address and key, the target device and the address of the target node to read the data to the storage node through the read control message of the data stream;

[0153] (3) The storage node transmits the target address of the data, the address and key of the read cache to the DPU / smart network card of the storage node;

[0154] (4) The DPU / Smart NIC triggers a read operation on the storage medium / device, and reads the data from the storage medium device DMA to the cache space of the DPU / Smart NIC;

[0155] (5) The DPU / Smart NIC of the storage node pushes the data to the DPU / Smart NIC cache of the client, completing the reading process of the data stream.

[0156] In an alternative embodiment, Figure 7 is a flowchart of a read data process according to an optional embodiment of the present application, such as Figure 7 As shown:

[0157] Step S701, the application performs a read operation on the specified file, offset and length through a system call and VFS, and passes in a read cache address;

[0158] Step S702, the storage client processes a read request (ie, a data read request);

[0159] Step S703, finding the address of the target node and target device of the data (i.e., the third address) according to the file and other information;

[0160] Step S704, calling the hardware acceleration interface - read preparation interface, and sending the data length, erasure correction, encryption and other parameters (i.e., calculation parameters) of the data to be read to the DPU / intelligent network card (i.e., the first processing unit);

[0161] The DPU / intelligent network card calls the corresponding erasure processing interface or encryption processing interface according to the incoming data length, erasure correction, encryption and other parameters, and allocates internal cache and computing units (i.e. Figure 7 The erasure calculation module and encryption calculation module in the data) perform erasure / encryption decoding operations on the data;

[0162] The DPU / intelligent network card returns the key value (i.e., the second identifier) ​​corresponding to the data cache in the DPU / intelligent network card through the hardware acceleration interface-read preparation interface;

[0163] Step S705: After receiving the return value of the Key value, the storage client sends the Key to the target node and target device (i.e. Figure 7 Storage nodes 71-storage nodes 73 in the storage node, that is, the second storage node).

[0164] Step S706: The storage system node receives the read control message, calls the DPU / intelligent network card hardware acceleration interface-read interface, and transmits the Key, client address, and target storage device address to the DPU / intelligent network card (i.e., the second processing unit) in the storage node; the DPU / intelligent network card in the storage node pushes the read data to the DPU / intelligent network card of the client;

[0165] Step S707: The client's DPU / intelligent network card receives the pushed data and puts it into the cache of the specified Key. After all the data is received, it performs erasure correction / encryption decoding calculation and stores the decoded data in the data cache.

[0166] After the DPU / intelligent network card of the storage system node completes pushing the data, it returns a successful read control message to the storage client; the client receives the read control message successfully, calls the hardware acceleration interface - read interface, and sends the Key and application cache address to the client's DPU / intelligent network card. The client's DPU / intelligent network card copies the calculated cache data (that is, the third data stream after decoding) to the application cache space based on the Key and application cache address; the client returns the system call of the application to complete the data reading operation.

[0167] In summary, the optional embodiment of the present application proposes a hardware erasure acceleration technology based on data stream processing, which realizes the fusion of main data stream processing and erasure / encryption calculation sub-processing through the coordinated processing of data stream processing and erasure calculation on the client of the distributed storage system, avoids data migration and interaction from the main data to the erasure sub-processing stream, and realizes hardware acceleration within the data stream; on high-performance data streams, traditional hardware acceleration cards have the problem that the consumption of moving PICe data back and forth between the CPU and the acceleration card is far greater than the improvement of hardware acceleration, and the data interaction bottleneck problem causes the hardware acceleration to fail to achieve the expected performance improvement. The optional embodiment of the present application does not change the direction of data flow, accelerates the data stream, and improves the performance of data processing.

[0168] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0169] In this embodiment, a data writing system is also provided, which is used to implement the above embodiments and preferred implementation modes, and the descriptions that have been made are not repeated. Although the systems described in the following embodiments are preferably implemented in software, the implementation in hardware, or a combination of software and hardware, is also possible and conceivable.

[0170] Figure 8 is a framework diagram of a data writing system according to an embodiment of the present application, such as Figure 8 As shown, the system includes: a storage client 82 and a target storage node 84, wherein:

[0171] The storage client 82 is used for, when receiving a data write request, performing a decoding calculation of a first calculation type on a first data stream corresponding to the data write request through a first processing unit in the storage client, and storing the first data stream after the decoding calculation in a second address space of a cache unit in the first processing unit; constructing a write control message according to a first address of a first solid-state hard disk corresponding to the first data stream and a first identifier of a second address corresponding to the second address space, and sending the write control message to a first storage node including the first solid-state hard disk, wherein the first calculation type includes at least one of the following: data erasure, data encryption, the first calculation type is a calculation type of the first data stream, and the first address is an address corresponding to a first address space in the first solid-state hard disk for writing the first data stream after the decoding calculation;

[0172] The target storage node 84 includes: the first storage node is used to obtain the first data stream after decoding and calculation according to the write control message, and write the first data stream after decoding and calculation into the first solid state hard disk.

[0173] Through the data writing system of the present application, the first data stream is encoded and calculated by the first processing unit in the storage client, and the first data stream after encoding and calculation is stored in the cache unit in the first processing unit; then, according to the first address of the first solid-state hard disk to be written by the first data stream and the first identifier of the second address of the cache unit where the first data stream is located, the first storage node including the first solid-state hard disk can write the first data stream after encoding and calculation into the first solid-state hard disk. In other words, the embodiment of the present application performs the encoding and calculation tasks in the data stream processing by the storage client instead of the CPU, thereby solving the problem of preemption of memory resources caused by the CPU running software to execute the encoding and calculation tasks in the data stream processing in the related technology. When the CPU processes the encoding and calculation, it also needs to process other services, thereby avoiding the problem of preemption of memory resources.

[0174] In an exemplary embodiment, the first storage node is also used to parse the write control message to determine the first address and the first identifier when receiving the write control message; call the write pull interface corresponding to the second processing unit in the first storage node, and send the first address and the first identifier to the second processing unit according to the write pull interface; the second processing unit is also used to determine the second address corresponding to the first data stream after decoding and calculation according to the first identifier, and send a data pull request to the first processing unit according to the second address, so that the first processing unit sends the first data stream after decoding and calculation to the second processing unit according to the data pull request; write the first data stream after decoding and calculation into the first address space corresponding to the first address.

[0175] In an exemplary embodiment, the storage client 82 is further used to obtain a second data stream corresponding to the data write request, and perform data preprocessing operations on the second data stream to obtain the first data stream; determine whether there is a reserved address space corresponding to the first data stream in multiple solid-state hard disks according to a distributed data layout algorithm, wherein each storage node includes a solid-state hard disk, and the multiple storage nodes include: the first storage node, and the reserved address space is an address space reserved for writing the first data stream after the encoding calculation; when it is determined that the reserved address space exists in one or more second solid-state hard disks among the multiple solid-state hard disks, determine the address corresponding to the one or more second solid-state hard disks as the first address, wherein the first solid-state hard disk includes: the one or more second solid-state hard disks; when it is determined that the reserved address space does not exist in the multiple solid-state hard disks, perform consistency negotiation with the multiple storage nodes so that the multiple storage nodes select one or more third solid-state hard disks from the multiple solid-state hard disks, and allocate address space for the first data stream in the one or more third solid-state hard disks; determine the address corresponding to the one or more third solid-state hard disks as the first address, wherein the first solid-state hard disk includes: the one or more third solid-state hard disks.

[0176] In an exemplary embodiment, the storage client 82 is also used to, when determining that the first calculation type is data erasure, call the data erasure calculation unit in the first processing unit through the first processing unit, and encode the first data stream according to the data erasure calculation unit to generate first erasure data and first verification data corresponding to the first data stream; and generate the first data stream after the encoding calculation according to the first erasure data and the first verification data.

[0177] In an exemplary embodiment, the storage client 82 is also used to call the data encryption calculation unit in the first processing unit through the first processing unit when it is determined that the first calculation type is data encryption, and determine the encryption algorithm and the first key corresponding to the first data stream according to the data encryption calculation unit; the encryption calculation unit encrypts the first data stream according to the encryption algorithm and the first key to generate the first data stream after the encoded calculation.

[0178] In an exemplary embodiment, when the first processing unit determines that the first data stream after the encoding calculation has been stored in the second address space, the first identifier corresponding to the second address space is determined; the completion status corresponding to the first data stream after the encoding calculation and the first identifier are sent to the storage client 82 through the write preparation interface corresponding to the first processing unit, wherein the completion status is used to indicate that the first data stream after the encoding calculation has been stored in the second address space.

[0179] In an exemplary embodiment, the storage client 82 is also used to obtain the first data information of the first data stream after the encoding calculation, and determine the message format of the write control message to be constructed, wherein the message format includes: header information and one or more data segments of the write control message, the first data information includes at least one of the following: the data stream size of the first data stream after the encoding calculation, the creation time of the first data stream after the encoding calculation, the data version corresponding to the first data stream after the encoding calculation, and the data write instruction corresponding to the first data stream after the encoding calculation; fill in the header information according to the first address and the first identifier, and fill in the one or more data segments according to the first data information, so as to construct the write control message according to the filled header information and the filled one or more data segments.

[0180] In an exemplary embodiment, the storage client 82 is also used to determine the second data information of the third data stream corresponding to the data read request when receiving the data read request, wherein the second data information includes at least one of the following: a third address of the second solid-state hard disk storing the third data stream, and a calculation parameter corresponding to the third data stream, and the calculation parameter includes at least one of the following: a second data correction and erasure parameter, a second data encryption parameter, and a second data deduplication parameter; the first processing unit in the storage client 82 is also used to allocate a third address space for the third data stream in the cache unit in the first processing unit, and send a second identifier of a fourth address corresponding to the third address space to the storage client; the storage client 82 is also used to construct a read control message based on the third address and the second identifier, and send the read control message to the second storage node corresponding to the second solid-state hard disk, wherein the target storage node includes: the second storage node.

[0181] In an exemplary embodiment, the second storage node is used to obtain the third data stream in the second solid-state hard disk according to the read control message, and send the third data stream to the third address space; the first processing unit is also used to perform a second calculation type of decoding calculation on the third data stream stored in the third address space according to the calculation parameters, and update the data in the third address space to the decoded third data stream, so that the storage client can read the decoded third data stream according to the second identifier, wherein the second calculation type includes at least one of the following: data correction and erasure, data encryption, and data deduplication, and the second calculation type is the calculation type of the third data stream.

[0182] In an exemplary embodiment, the storage client 82 is further configured to call a read preparation interface corresponding to the first processing unit, and send the second data information to the first processing unit according to the read preparation interface.

[0183] In an exemplary embodiment, the storage client 82 is also used to determine the application cache address corresponding to the data read request, wherein the application cache address is the address corresponding to the application cache space in the target object for reading the third data stream after decoding and calculation; call the read interface corresponding to the first processing unit, and send the second identifier and the application cache address to the first processing unit according to the read interface; upon receiving the second identifier and the application cache address, the first processing unit sends the third data stream after decoding and calculation to the application cache space corresponding to the application cache address, so that the target object reads the third data stream after decoding and calculation.

[0184] In an exemplary embodiment, the storage client 82 is also used to call the data erasure correction calculation unit in the first processing unit through the first processing unit when the second calculation type is data erasure correction, so that the data erasure correction calculation unit determines the second erasure correction data and the second verification data corresponding to the third data stream according to the calculation parameters; and decodes and calculates the third data stream according to the second erasure correction data and the second verification data.

[0185] In an exemplary embodiment, the storage client 82 is also used to call the data encryption calculation unit in the first processing unit through the first processing unit when the second calculation type is data encryption, and obtain the second key and the second encryption algorithm for encrypting the third data stream according to the data encryption calculation unit; determine the decryption algorithm corresponding to the second encryption algorithm, and decrypt the third data stream according to the decryption algorithm and the second key to generate the third data stream after the decoding calculation.

[0186] The description of the features in the embodiment corresponding to the data writing system can refer to the relevant description of the embodiment corresponding to the data writing method, which will not be repeated here.

[0187] An embodiment of the present application further provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above data writing method embodiments.

[0188] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned data writing method embodiments when running.

[0189] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0190] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above data writing method embodiments are implemented.

[0191] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned data writing method embodiments are implemented.

[0192] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0193] The above is a detailed introduction to a data writing method provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core idea of ​​the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A data writing method, characterized in that: Applicable to storage clients, including: In the case of receiving a data write request, performing a first calculation type of encoding calculation on a first data stream corresponding to the data write request by a first processing unit in the storage client, and storing the first data stream after the encoding calculation in a second address space of a cache unit in the first processing unit, wherein the first calculation type includes at least one of the following: data erasure, data encryption, and the first calculation type is a calculation type of the first data stream; Constructing a write control message according to a first address of the first solid-state hard disk corresponding to the first data stream and a first identifier of a second address corresponding to the second address space, wherein the first address is an address corresponding to a first address space in the first solid-state hard disk for writing the first data stream after the encoding calculation; The write control message is sent to the first storage node including the first solid state drive, so that the first storage node obtains the first data stream after the encoding calculation according to the write control message, and writes the first data stream after the encoding calculation into the first solid state drive.

2. The data writing method according to claim 1, characterized in that: Before constructing a write control message according to the first address of the first solid-state hard disk corresponding to the first data stream and the first identifier of the second address corresponding to the second address space, the method further includes: Acquire a second data stream corresponding to the data write request, and perform a data preprocessing operation on the second data stream to obtain the first data stream; Determine whether there is a reserved address space corresponding to the first data stream in a plurality of solid-state hard disks according to a distributed data layout algorithm, wherein each storage node includes a solid-state hard disk, and the plurality of storage nodes include: the first storage node, the reserved address space being an address space reserved for writing the first data stream after the encoding calculation; In the case where it is determined that the reserved address space exists in one or more second solid-state hard disks among the plurality of solid-state hard disks, determining addresses corresponding to the one or more second solid-state hard disks as the first address, wherein the first solid-state hard disk includes: the one or more second solid-state hard disks; When it is determined that the reserved address space does not exist in the plurality of solid-state hard disks, performing consistency negotiation with the plurality of storage nodes so that the plurality of storage nodes select one or more third solid-state hard disks from the plurality of solid-state hard disks, and allocate address space for the first data stream in the one or more third solid-state hard disks; The addresses corresponding to the one or more third solid-state hard disks are determined as the first address, wherein the first solid-state hard disk includes: the one or more third solid-state hard disks.

3. The data writing method according to claim 1, characterized in that: Performing a first computing type of encoding calculation on a first data stream corresponding to the data write request by a first processing unit in the storage client includes: In a case where it is determined that the first calculation type is data erasure correction, calling a data erasure correction calculation unit in the first processing unit through the first processing unit, and encoding the first data stream according to the data erasure correction calculation unit to generate first erasure correction data and first verification data corresponding to the first data stream; The first data stream after the encoding calculation is generated according to the first erasure correction data and the first verification data.

4. The data writing method according to claim 1, characterized in that: Performing a first computing type of encoding calculation on a first data stream corresponding to the data write request by a first processing unit in the storage client includes: In the case where it is determined that the first calculation type is data encryption, calling a data encryption calculation unit in the first processing unit through the first processing unit, and determining an encryption algorithm and a first key corresponding to the first data stream according to the data encryption calculation unit; The encryption calculation unit performs encryption processing on the first data stream according to the encryption algorithm and the first key to generate the first data stream after the encoding calculation.

5. The data writing method according to claim 1, characterized in that: After storing the first data stream after the encoding calculation in the second address space of the cache unit in the first processing unit, the method further includes: When the first processing unit determines that the first data stream after the encoding calculation has been stored in the second address space, determining the first identifier corresponding to the second address space; The completion status corresponding to the first data stream after the encoding calculation and the first identifier are sent to the storage client through the write preparation interface corresponding to the first processing unit, wherein the completion status is used to indicate that the first data stream after the encoding calculation has been stored in the second address space.

6. The data writing method according to claim 1, characterized in that: Constructing a write control message according to a first address of a first solid-state hard disk corresponding to the first data stream and a first identifier of a second address corresponding to the second address space, including: Acquire first data information of the first data stream after the encoding calculation, and determine the message format of the write control message to be constructed, wherein the message format includes: header information and one or more data segments of the write control message, and the first data information includes at least one of the following: data stream size of the first data stream after the encoding calculation, creation time of the first data stream after the encoding calculation, data version corresponding to the first data stream after the encoding calculation, and data write instruction corresponding to the first data stream after the encoding calculation; The header information is filled according to the first address and the first identifier, and the one or more data segments are filled according to the first data information, so as to construct the write control message according to the filled header information and the filled one or more data segments.

7. The data writing method according to claim 1, characterized in that: The method further comprises: In the case of receiving a data read request, obtaining, through the first processing unit, second data information of a third data stream corresponding to the data read request, wherein the second data information includes at least one of the following: a third address of a second solid-state hard disk storing the third data stream, and a calculation parameter corresponding to the third data stream, wherein the calculation parameter includes at least one of the following: a data erasure parameter, a data encryption parameter, and a data deduplication parameter; Allocate a third address space for the third data stream in the cache unit by the first processing unit, and send a second identifier of a fourth address corresponding to the third address space to the storage client; A data read operation corresponding to the data read request is performed according to the third address and the second identifier.

8. The data writing method according to claim 7, characterized in that: Executing a data read operation corresponding to the data read request according to the third address and the second identifier includes: Constructing a read control message according to the third address and the second identifier, and sending the read control message to the second storage node corresponding to the second solid-state hard disk, so that the second storage node obtains the third data stream in the second solid-state hard disk according to the read control message, and sends the third data stream to the third address space; Control the first processing unit to perform a decoding calculation of a second calculation type on the third data stream stored in the third address space according to the calculation parameters, and update the data in the third address space to the third data stream after the decoding calculation, so that the target object sending the data read request can read the third data stream after the decoding calculation, wherein the second calculation type includes at least one of the following: data erasure, data encryption, and the second calculation type is the calculation type of the third data stream.

9. The data writing method according to claim 8, characterized in that: Before acquiring, by the first processing unit, second data information of the third data stream corresponding to the data read request, the method further includes: Acquire a data stream identifier corresponding to the third data stream according to the data read request, and determine distribution information corresponding to the third data stream according to the data stream identifier, wherein the distribution information is used to indicate distribution of the third data stream in multiple storage nodes; parsing the distribution information, and determining the second storage node storing the third data stream among the plurality of storage nodes according to the parsed distribution information; An address corresponding to the second solid state hard disk in the second storage node is determined as the third address.

10. The data writing method according to claim 8, characterized in that: Acquiring, by the first processing unit, second data information of the third data stream corresponding to the data read request, includes: A read preparation interface corresponding to the first processing unit is called, and the second data information is sent to the first processing unit according to the read preparation interface.

11. The data writing method according to claim 10, characterized in that: After updating the data in the third address space to the third data stream after decoding and calculation, the method further includes: Determine an application cache address corresponding to the data read request, wherein the application cache address is an address corresponding to an application cache space in the target object for reading the third data stream after decoding and calculation; calling a read interface corresponding to the first processing unit, and sending the second identifier and the application cache address to the first processing unit according to the read interface; When receiving the second identifier and the application cache address, the first processing unit sends the decoded and calculated third data stream to the application cache space corresponding to the application cache address, so that the target object reads the decoded and calculated third data stream.

12. The data writing method according to claim 10, characterized in that: The first processing unit performs a decoding calculation of a second calculation type on the third data stream stored in the third address space according to the calculation parameter, including: In a case where the second calculation type is data erasure correction, calling a data erasure correction calculation unit in the first processing unit through the first processing unit, so that the data erasure correction calculation unit determines second erasure correction data and second verification data corresponding to the third data stream according to the calculation parameters; The third data stream is decoded and calculated according to the second erasure correction data and the second check data.

13. The data writing method according to claim 10, characterized in that: The first processing unit performs a decoding calculation of a second calculation type on the third data stream stored in the third address space according to the calculation parameter, including: In the case where the second calculation type is data encryption, calling a data encryption calculation unit in the first processing unit through the first processing unit, and obtaining a second key and a second encryption algorithm for encrypting the third data stream according to the data encryption calculation unit; A decryption algorithm corresponding to the second encryption algorithm is determined, and the third data stream is decrypted according to the decryption algorithm and the second key to generate the third data stream after the decoding calculation.

14. A data writing system, characterized in that: It includes storage client and target storage node, where: The storage client is used for, upon receiving a data write request, performing a decoding calculation of a first calculation type on a first data stream corresponding to the data write request through a first processing unit in the storage client, and storing the first data stream after the decoding calculation in a second address space of a cache unit in the first processing unit; constructing a write control message according to a first address of a first solid-state hard disk corresponding to the first data stream and a first identifier of a second address corresponding to the second address space, and sending the write control message to a first storage node including the first solid-state hard disk, wherein the first calculation type includes at least one of the following: data erasure, data encryption, the first calculation type is a calculation type of the first data stream, and the first address is an address corresponding to a first address space in the first solid-state hard disk for writing the first data stream after the decoding calculation; The target storage node includes: the first storage node is used to obtain the first data stream after decoding and calculation according to the write control message, and write the first data stream after decoding and calculation into the first solid state hard disk.

15. The data writing system according to claim 14, characterized in that: include: The first storage node is further configured to parse the write control message upon receiving the write control message to determine the first address and the first identifier; calling a write-pull interface corresponding to a second processing unit in the first storage node, and sending the first address and the first identifier to the second processing unit according to the write-pull interface; The second processing unit is further configured to determine the second address corresponding to the first data stream after decoding and calculation according to the first identifier, and send a data pull request to the first processing unit according to the second address, so that the first processing unit sends the first data stream after decoding and calculation to the second processing unit according to the data pull request; The first data stream after decoding and calculation is written into a first address space corresponding to the first address.

16. The data writing system according to claim 14, characterized in that: include: The storage client is further used to determine, when receiving a data read request, second data information of a third data stream corresponding to the data read request, wherein the second data information includes at least one of the following: a third address of a second solid-state hard disk storing the third data stream, and a calculation parameter corresponding to the third data stream, wherein the calculation parameter includes at least one of the following: a second data erasure parameter, a second data encryption parameter, and a second data deduplication parameter; The first processing unit in the storage client is further configured to allocate a third address space for the third data stream in a cache unit in the first processing unit, and send a second identifier of a fourth address corresponding to the third address space to the storage client; The storage client is further used to construct a read control message according to the third address and the second identifier, and send the read control message to the second storage node corresponding to the second solid state drive, wherein the target storage node includes: the second storage node.

17. The data writing system according to claim 16, characterized in that: include: The second storage node is configured to obtain the third data stream in the second solid state drive according to the read control message, and send the third data stream to the third address space; The first processing unit is also used to perform a decoding calculation of a second calculation type on the third data stream stored in the third address space according to the calculation parameters, and update the data in the third address space to the third data stream after decoding and calculation, so that the storage client can read the third data stream after decoding and calculation according to the second identifier, wherein the second calculation type includes at least one of the following: data correction and erasure, data encryption, and data deduplication, and the second calculation type is the calculation type of the third data stream.

18. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the data writing method according to any one of claims 1 to 13 when executing the computer program.

19. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the data writing method according to any one of claims 1 to 13.

20. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the data writing method according to any one of claims 1 to 13 are implemented.

Citation Information

Patent Citations

  • Method for processing data using intermediate device, computer system, and intermediate device

    CN113961139A

  • Data stream processing method, storage control node and readable storage medium

    CN114201421A

  • Method for processing non-cache write data request, cache and node

    CN114731282A

  • Storage node and operating method thereof

    CN116414612A

  • Data processing method and device of storage system, storage system, equipment and medium

    CN116886719A