Data processing method, electronic equipment and computer readable storage medium

By using pattern index instead of CRC verification in the data processing method, the problem of slow read and write speed in PyNVMe is solved, and faster data read and write speed and smaller memory footprint are achieved, which is suitable for data processing on the host side and storage devices.

CN120255784APending Publication Date: 2025-07-04SHANGHAI LONGSYS DIGITAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410009906.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-03
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the prior art, the original reading and writing data methods in PyNVMe require frequent buffer filling and use a large amount of memory for CRC verification, resulting in slow reading and writing speed, low efficiency, and time-consuming to find logical block addresses.

Method used

Mode index is used instead of traditional CRC verification. By selecting the corresponding mode index for each logical block address, the first cache data is generated and stored in the write buffer, and the mode index is used for verification during reading, reducing memory usage and search time.

Benefits of technology

Speed up the overall speed of data reading and writing, especially for hosts with smaller memory, reduce memory overhead and save time in finding logical block addresses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120255784A_ABST
    Figure CN120255784A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method, electronic equipment and a computer readable storage medium, and the data processing method is applied to a host end and comprises the steps that in response to a data writing command, a corresponding mode index is selected for each logic block address in the data writing command; obtaining first cache data corresponding to each logic block address based on the mode index; in response to the fact that the first cache data corresponding to all the logic block addresses are written into the write buffer area, a data write-in command is sent to the storage device; and in response to completion of execution of the data writing command, caching all the mode indexes corresponding to the data writing command into the memory, the mode indexes being used for performing reading verification during data reading. By means of the mode, due to the fact that the data size of the mode index is small, occupied memory is small, only a small amount of time is needed to be saved in the memory, and the overall speed of data reading and writing can be increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, in particular to a data processing method, an electronic device, and a computer-readable storage medium. Background Art

[0002] In the field of data storage, the integrity of data is of utmost importance, and it is necessary to ensure the consistency between the read data and the corresponding data at the time of writing.

[0003] Currently, when using the original data reading and writing methods in PyNVMe to read and write data, each time a buffer needs to be filled, which is time-consuming. Moreover, the CRC (Cyclic Redundancy Check) verification method used occupies a large amount of memory on the host side, resulting in a slow reading and writing speed and low efficiency of data. Summary of the Invention

[0004] This application provides a data processing method, an electronic device, and a computer-readable storage medium. Since the amount of data of the pattern index is small, the memory occupancy is small, and only a small amount of time is required to store it in the memory, thereby accelerating the overall speed of data reading and writing.

[0005] To solve the above technical problems, a technical solution adopted by this application is: to provide a data processing method, which is applied to the host side and includes: in response to a data write command, selecting a corresponding pattern index for each logical block address in the data write command; obtaining first cache data corresponding to each logical block address based on the pattern index; in response to all the first cache data corresponding to the logical block addresses being written to the write buffer, sending the data write command to the storage device; in response to the completion of the execution of the data write command, caching all the pattern indexes corresponding to the data write command in the memory, where the pattern index is used for read verification during data reading.

[0006] In some embodiments, obtaining first cache data corresponding to each logical block address based on the pattern index includes: determining the data amount of the first cache data according to each logical block address; using the data corresponding to each logical block address and the pattern index to form the first cache data that meets the data amount.

[0007] In some embodiments, using the data corresponding to each logical block address and the pattern index to form the first cache data that meets the data amount includes: obtaining initial cache data with a corresponding data amount from the data buffer using the pattern index; replacing the data at both ends of the initial cache data with the logical block address and the namespace identifier corresponding to the logical block address to obtain the first cache data.

[0008] In some embodiments, in response to a data write command, selecting a corresponding pattern index for each logical block address in the data write command includes: in response to the data write command, randomly allocating a pattern index from a pattern index array for each logical block address; wherein each pattern index corresponds to an element in the pattern index array, and some of the elements in the pattern index array are different fixed values, and the other part of the elements are random values.

[0009] To solve the above technical problems, another technical solution adopted by this application is: to provide a data processing method, which is applied to the host side and includes: sending a data read command to a storage device to make the storage device feedback corresponding first target data; in response to the completion of the execution of the data read command, selecting a corresponding pattern index from the memory for each logical block address in the data read command, wherein the pattern index in the memory is obtained according to the data processing method as described above; obtaining second target data corresponding to each logical block address based on the pattern index; performing a read verification based on the first target data and the second target data.

[0010] In some embodiments, obtaining second target data corresponding to each logical block address based on the pattern index includes: using the data corresponding to each logical block address and the pattern index to form corresponding second target data.

[0011] In some embodiments, using the data corresponding to each logical block address and the pattern index to form corresponding second target data includes: using the pattern index to obtain initial cache data of a corresponding data volume from a data buffer; using the logical block address and the namespace identifier corresponding to the logical block address to replace the data at both ends of the initial cache data to obtain second target data.

[0012] In some embodiments, performing a read verification based on the first target data and the second target data includes: in response to the first target data and the second target data being consistent, determining that the data read is successful; in response to the first target data and the second target data being inconsistent, determining that the data read fails.

[0013] To solve the above technical problems, another technical solution adopted by this application is: to provide an electronic device, which includes a host side and a storage device, wherein the host side is provided with a memory, and the host side and the storage device are configured to implement the above data processing method.

[0014] To solve the above technical problems, another technical solution adopted by this application is: to provide a computer-readable storage medium, which is used to store a computer program, and the computer program is used to implement the above data processing method when executed.

[0015] In some embodiments of the present application, the data processing method provided responds to a data write command, selects a corresponding mode index for each logical block address in the data write command; obtains first cache data corresponding to each logical block address based on the mode index; sends the data write command to the storage device in response to all the first cache data corresponding to the logical block addresses being written to the write buffer; and caches all the mode indexes corresponding to the data write command in the memory in response to the completion of the execution of the data write command, where the mode index is used for read verification during data reading. In the above manner, since the amount of data of the mode index is small, the memory occupancy is small, and only a small amount of time is required to store it in the memory, thereby accelerating the overall speed of data reading and writing. Description of the Drawings

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. Among them:

[0017] Figure 1 is a schematic flowchart of an embodiment of the related art provided by the present application;

[0018] Figure 2 is a schematic flowchart of another embodiment of the related art provided by the present application;

[0019] Figure 3 is a schematic flowchart of the first embodiment of the data processing method provided by the present application;

[0020] Figure 4 is a schematic diagram of an embodiment of the mode index array provided by the present application;

[0021] Figure 5 is a schematic diagram of an embodiment of the first cache data provided by the present application;

[0022] Figure 6 is a schematic flowchart of the second embodiment of the data processing method provided by the present application;

[0023] Figure 7 is a schematic structural diagram of an embodiment of the electronic device provided by the present application;

[0024] Figure 8 is a schematic structural diagram of an embodiment of the computer-readable storage medium provided by the present application. Detailed Embodiments

[0025] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0026] In the related art, the original read and write data methods in PyNVMe are usually used to read and write data. At this time, reference can be made to Figure 1 and Figure 2 , Figure 1 is a schematic flow chart of writing data, Figure 2 is a schematic flow chart of reading data, both of which are applied to the host side.

[0027] See Figure 1 , Figure 1 The specific steps of writing data shown are as follows:

[0028] S11: Apply for the memory space of the cyclic redundancy check table.

[0029] Specifically, apply for a cyclic redundancy check table (CRC Table) according to the space size of the actual namespace. The memory size occupied by this check table is the total length of all CRC32 (that is, a CRC will generate a 32-bit / 8-bit hexadecimal check value) check sums, which is the number of logical block addresses (LBA) of the namespace multiplied by 4 bytes.

[0030] S12: Apply for a write buffer.

[0031] Among them, the write buffer corresponds to the Write Buffer, which is used to fill the write data, and the data in the Write Buffer is randomly filled.

[0032] S13: Send a data write command and wait for completion.

[0033] Among them, the data write command is sent to the storage device to be saved in the storage device.

[0034] S14: Determine whether the execution of the data write command is completed.

[0035] In response to the completion of the execution of the data write command, step S16 is executed; otherwise, step S15 is executed.

[0036] S15: Report an error.

[0037] It can be understood that if the execution of the data write command fails to be displayed, an error message is sent to the host, and the host can make adjustments according to the error message to send the data write command to the storage device again or execute other programs.

[0038] S16: Determine whether the written data corresponding to the verified data write command is correct.

[0039] If the written data corresponding to the data write command is incorrect, step S17 is executed; otherwise, step S18 is executed.

[0040] S17: Report an error.

[0041] It can be understood that when the data written according to the data write command is inconsistent with the data that actually needs to be written, it indicates that the data write fails. At this time, an error message is sent to the host so that the host can make adjustments according to the received error message to execute the above process again to rewrite the data, find the data that is inconsistent with the data that actually needs to be written, or detect the reason for the data write failure.

[0042] S18: Calculate the 32-bit cyclic redundancy checksum of the logical block address and store it in the cyclic redundancy check table.

[0043] It can be understood that to verify whether the data corresponding to the data write command is correct, if it is incorrect, an error will be reported; if it is correct, the CRC32 checksum of each LBA can be calculated and stored in the CRC Table.

[0044] The inventor of the present application found that based on the above-described content, it can be known that in the related method of writing data, the buffer needs to be refilled every time it is executed, and the CRC check method used occupies a large amount of memory, the reading and writing speed is slow and the efficiency is low, and since the entire process does not include LAB, it is time-consuming to search for LAB-related data every time.

[0045] See Figure 2 , Figure 2 The specific steps for reading data shown in

[0046] S21: Prepare the write buffer.

[0047] S22: Send a data read command and wait for the data read command to complete.

[0048] S23: In response to the completion of the data read command, determine whether the execution status of the data read command is successful.

[0049] If not, step S24 is executed; otherwise, step S25 is executed.

[0050] S24: Report an error.

[0051] S25: Calculate the actual 32-bit cyclic redundancy checksum for each logical block address.

[0052] Specifically, calculate the actual CRC32 checksum for each LBA according to the actual returned data of the data read command.

[0053] S26: Determine whether the actual 32-bit cyclic redundancy checksum is consistent with the 32-bit cyclic redundancy checksum in the cyclic redundancy check table.

[0054] Specifically, compare the calculated CRC32 checksum with the CRC32 checksum stored corresponding to each LBA in the CRC Table one by one. If all are consistent, execute step S27; otherwise, execute step S28.

[0055] S27: The verification passes and the data reading is successful.

[0056] S28: Report an error.

[0057] It can be understood that when the actual CRC Table is inconsistent with the CRC32 checksum in the CRC Table, it indicates that the verification fails, and the corresponding error is "Data Miscompare". According to the original data reading and writing method in PyNVMe used above, each time of reading and writing requires re-filling the data into the buffer, and the CRC verification method used occupies a large amount of memory, making the speed of reading and writing data slow in the overall process, the reading and writing efficiency low, and it takes a long time to search for the LBA address information.

[0058] Based on this, the present application provides a data processing method, an electronic device, and a computer-readable storage medium, which can solve any of the above technical defects.

[0059] In the embodiment of the present application, refer to Figure 3 , Figure 3 is a schematic flowchart of the first embodiment of the data processing method provided by the present application. This data processing method is applied to the host side and may include the following steps:

[0060] Step 31: In response to a data write command, select a corresponding pattern index for each logical block address in the data write command.

[0061] It can be understood that each data write command corresponds to a plurality of logical block addresses (Logical Block Address, LBA). Each logical block address can correspond to a pattern index (Pattern Index), and the pattern indexes corresponding to all the logical block addresses included in the data write command can be the same or different.

[0062] In an embodiment of the present application, each mode index occupies 4 bits. Compared with the original read and write data method in PyNVMe used in the related art, the memory consumed by the present application is 1 / 8 of the 4-byte CRC32 checksum in the original PyNVMe. Since the data volume of the mode index is small, the memory occupation is small, and only a small amount of time is required to store it in the memory, thereby accelerating the overall speed of data reading and writing.

[0063] In some embodiments, in response to a data write command, a mode index is randomly allocated from the mode index array for each logical block address; wherein, each mode index corresponds to an element in the mode index array, and some elements in the mode index array are different fixed values, and the other part of the elements are random values.

[0064] Among them, the mode index array may include a preset number of elements, such as including 18 groups, 24 groups or 36 groups of elements, which can be specifically determined according to the actual situation and are not limited here.

[0065] In an embodiment of the present application, refer to Figure 4 , the mode index array contains 18 groups of elements. Among them, the 1st to 12th, 16th to 18th groups of elements are different fixed values, and the 13th to 15th groups are random values. When selecting the corresponding mode index for each logical block address, the involved mode index should be determined among the 1st to 15th groups of elements.

[0066] Specifically, among the 1st to 12th groups of elements, each group of elements consists of a corresponding byte array. Each byte array contains four hexadecimal codes. For example, the 3rd group of elements is "bytearray([0xA5,0x5A,0x5A,0xA5])". The size of the logical block address involved in the present application is determined according to the actual situation, which can be 512 bytes, 4096 bytes or others, and is not limited here.

[0067] It should be noted that when the data write command is a TRIM command, the corresponding data is TRIM data. At this time, the mode index selected for each logical block address in the data write command is the element of the 2nd group in the mode index array, that is, the content of the corresponding mode index is "bytearray([0x00,0x00,0x00,0x00])".

[0068] Among the 13th to 15th groups of elements, each group of elements contains four 4-digit hexadecimal codes. Each 4-digit hexadecimal code contains 8 binary codes, and the first 8 binary codes (i.e., the first 4-digit hexadecimal code) and the last eight binary codes (i.e., the last 4-digit hexadecimal code) in each group of elements are the same. Moreover, when the pattern index corresponding to the logical block address is within the range of the 13th to 15th groups of elements, a random seed is generated while selecting the pattern index, and this random seed can be used to find the cached data to be read when reading data subsequently.

[0069] Among the 16th to 18th groups of elements, the 16th to 18th groups of elements are used to store invalid data or exception data fed back by the pattern index. These invalid data or exception data are generated by the content of the illegally configured pattern index. Among them, the 16th group of elements is used to store the logical block address for error checking and correction (Uncorrectable Error Checking and Correcting Logical Block Address, UECC LBA), and the 17th and 18th groups of elements are used to store the codes of invalid patterns, such as "LbaOutOfNsidRange" indicating "the logical block address exceeds the network number range" and "NsidNotInitialzed" indicating "the network number is not initialized".

[0070] Step 32: Obtain the first cached data corresponding to each logical block address based on the pattern index.

[0071] In some embodiments, the data volume of the first cached data is determined according to each logical block address; the data corresponding to each logical block address and the pattern index is used to form the first cached data that meets the data volume.

[0072] Specifically, the data volume of the first cached data is determined according to each logical block address, and then the initial data of the corresponding data volume is obtained from the data buffer area using the pattern index, and the first cached data is obtained by replacing the data at both ends of the initial cached data with the logical block address and the namespace identifier corresponding to the logical block address.

[0073] In the embodiments of the present application, the first cached data is obtained by adding the logical block address to both the head and the tail of the initial cached data, and the data volume of the first cached data is the data volume of the logical block address. At this time, reference can be made to Figure 5 , Figure 5 which is based on the namespace ID being 1, the data volume being 512 (16×32) bytes, the logical block address being "0x2AB49A8" and corresponding to the pattern index Figure 4The first cached data generated from the 5th group element "[0xFF, 0x00, 0x00, 0xFF]" in it. Among them, the 512-byte first cached data is filled with the ID of the namespace (corresponding to the namespace identifier), the logical block address "0x2AB49A8", and the 5th group element "[0xFF, 0x00, 0x00, 0xFF]". As Figure 5 shown, the initial buffered data is filled with the 5th group element "[0xFF, 0x00, 0x00, 0xFF]". Then, using the logical block address "0x2AB49A8" and the composition of the ID of the namespace "a8 49ab 02 0000 00 01" to replace the data at both ends of the initial buffered data, finally forming Figure 5 the data shown. That is, the first 8 bytes and the last 8 bytes of the final first cached data are the same, both representing the LBA and namespace information. Among the 8 bytes, the highest 1 byte represents the namespace ID, and the remaining 7 bytes represent the LBA.

[0074] By the way of obtaining the first cached data by replacing the data at both ends of the initial buffered data with the logical block address and the namespace identifier corresponding to the logical block address, the target data can be obtained more quickly by only modifying part of the data. Compared with the way of rewriting to find all data or modifying all data, the time to obtain the target data can be reduced, and the speed of obtaining the target data can be improved.

[0075] Step 33: In response to all the first cached data corresponding to the logical block addresses being written into the write buffer, send a data write command to the storage device.

[0076] Among them, the write buffer corresponds to the write buffer, and the storage device can be an SSD (Solid State Disk or Solid State Drive, solid state drive) device.

[0077] Step 34: In response to the completion of the execution of the data write command, cache all the mode indexes corresponding to the data write command in the memory, where the mode index is used for read verification during data reading.

[0078] Among them, the memory space corresponds to the memory zone.

[0079] Different from the prior art, the data processing method provided in this application is applied to the host side. By responding to a data write command, a corresponding pattern index is selected for each logical block address in the data write command; based on the pattern index, the first cache data corresponding to each logical block address is obtained; in response to all the first cache data corresponding to the logical block addresses being written into the write buffer, the data write command is sent to the storage device; in response to the completion of the execution of the data write command, all the pattern indexes corresponding to the data write command are cached in the memory, where the pattern index is used for read verification during data reading. In the above manner, since the amount of data of the pattern index is small, the memory occupancy is small, and only a small amount of time is required to store it in the memory, thereby accelerating the overall speed of data reading and writing. This is more friendly to some hosts with small memory, and can also reduce a part of the memory overhead for hosts with a large enough memory configuration.

[0080] In other embodiments, the first cache data can be mapped into a buffer (buffer). At this time, the data processing method corresponding to the embodiment shown above Figure 3 is applied to the host side and may include the following steps (not shown in the figure):

[0081] S1: Select any one LBA of the data write command.

[0082] It can be understood that the data write command contains several logical block addresses (Logical Block Address, LBA). Each logical block address can correspond to a pattern index, and the pattern indexes corresponding to all the logical block addresses included in the data write command can be the same or different.

[0083] In the embodiments of this application, each pattern index occupies 4 bits. Compared with the original read and write data method in PyNVMe used in the related art, the memory consumed by this application is 1 / 8 of the 4-byte CRC32 checksum in the original PyNVMe. Since the amount of data of the pattern index is small, the memory occupancy is small, and only a small amount of time is required to store it in the memory, thereby accelerating the overall speed of data reading and writing.

[0084] S2: Select a Pattern Index from a preset number of Pattern Indexes.

[0085] Among them, the preset number can be 16, 18 or others. In the embodiments of this application, the preset PatternIndex (pattern index) of the data is 15 groups of Pattern Indexes, that is, a Pattern Index corresponding to the LBA selected in step S1 is selected from 15 groups of Pattern Indexes.

[0086] S3: Obtain the corresponding Data Buffer from the Buffer Cache according to the Pattern Index, replace the address information of the first and last 8 bytes, copy all the data in the Data Buffer to the corresponding position in the Write Buffer, and temporarily cache the Pattern Index in the Pattern Cache.

[0087] Among them, the initial cached data is stored in the Data Buffer. By replacing the eight bytes of the data at both ends of the initial cached data stored in the Data Buffer, the first cached data can be obtained. The Buffer Cache is a buffer cache area that stores several Data Buffers. The Pattern Cache is a pattern cache area that can be used to store the Pattern Index.

[0088] Among them, the Write Buffer corresponds to the write buffer area.

[0089] S4: Repeat steps S1 to S3 to traverse all LBAs until all LBAs in the data write command are filled into the Write Buffer.

[0090] The Write Buffer also stores the data to be written corresponding to the data write command.

[0091] S5: In response to all LBAs being filled into the Write Buffer, send the data write command to the storage device.

[0092] S6: Determine whether the execution status of the data write command is successful.

[0093] If so, execute step S7; otherwise, execute step S8.

[0094] S7: Move the Pattern Index of all LBAs corresponding to the data write command from the Pattern Cache to the Pattern Buffer.

[0095] Among them, the Pattern Buffer corresponds to the memory.

[0096] S8: Report an error.

[0097] It can be understood that if the data write command fails, an error message is sent to the host side so that the host side can perform error verification or related operations for re-writing the data.

[0098] Different from the prior art, the data processing method provided by this application and Figure 3The same or similar to the illustrated embodiments can accelerate the data writing speed, and since the mode index (usually four bits) has a small memory, the content occupied during the data writing process is also small, which is more friendly to some host sides with small memory, and can also reduce a part of the memory overhead for host sides with large enough memory configurations. Moreover, compared with the existing verification scheme in PyNVMe that does not contain LBA address information, the Data Buffer in this application contains LBA address information, which can save the process of continuously searching for LBA to save time.

[0099] In an embodiment of this application, refer to Figure 6 , Figure 6 is a schematic flowchart of a second embodiment of the data processing method provided by this application. This data processing method is applied to the host side and includes:

[0100] Step 61: Send a data read command to the storage device to make the storage device feedback the corresponding first target data.

[0101] Among them, the first target data is the actual data read from the storage device according to the LBA in the data read command. This actual data is written in the manner of any of the above embodiments.

[0102] Step 62: In response to the completion of the execution of the data read command, select the corresponding mode index from the memory for each logical block address in the data read command.

[0103] Specifically, in response to the completion of the execution of the data read command, since the corresponding relationship between the defined mode index and the logical block address is defined during data writing, and the mode index is stored in the memory, therefore, when the logical block address is known, the mode index corresponding to the logical block address can be found from the memory. Among them, the mode index in the memory is obtained according to the data processing method as Figure 3 corresponds to, and Figure 4 the mode index array shown.

[0104] Each mode index corresponds to an element in the mode index array. Some elements in the mode index array are different fixed values, and the other part of the elements are random values. Among them, the mode index array can include a preset number of elements, such as including 18 groups, 24 groups or 36 groups of elements, which can be determined according to the actual situation specifically and are not limited here.

[0105] In an embodiment of this application, refer to Figure 4 , the mode index array contains 18 groups of elements. Among them, the 1st to 12th, 16th to 18th groups of elements are different fixed values, and the 13th to 15th groups are random values. When selecting the corresponding mode index for each logical block address, the involved mode index should be determined among the 1st to 15th groups of elements.

[0106] Specifically, among the 1st to 12th groups of elements, each group of elements consists of corresponding byte arrays. Each byte array contains four hexadecimal codes. For example, the 3rd group of elements is "bytearray([0xA5,0x5A,0x5A,0xA5])". The size of the logical block address involved in this application is determined according to the actual situation, which can be 512 bytes, 4096 bytes or others, and there is no limitation here.

[0107] It should be noted that when the data write command is a TRIM command, the corresponding data is TRIM data. At this time, the mode index selected for each logical block address in the data write command is the element of the 2nd group in the mode index array, that is, the content of the corresponding mode index is "bytearray([0x00,0x00,0x00,0x00])".

[0108] Among the 13th to 15th groups of elements, each group of elements contains four 4-bit hexadecimal codes. Each 4-bit hexadecimal code contains 8 binary codes, and the first 8 binary codes (i.e., the first 4-bit hexadecimal code) and the last eight binary codes (i.e., the last 4-bit hexadecimal code) in each group of elements are the same. Moreover, when the mode index corresponding to the logical block address is within the range of the 13th to 15th groups of elements, a random seed will be generated while selecting the mode index, and this random seed can be used to find the cached data of the data to be read when reading data subsequently.

[0109] Among the 16th to 18th groups of elements, the 16th to 18th groups of elements are used to store invalid data or abnormal data fed back through the mode index. These invalid data or abnormal data are generated by the content of the illegally configured mode index. Among them, the 16th group of elements is used to store the logical block address for error checking and correction (Uncorrectable Error Checking and Correcting Logical Block Address, UECC LBA), and the 17th and 18th groups of elements are used to store the codes of invalid patterns, such as "LbaOutOfNsidRange" means "the logical block address exceeds the network number range", and "NsidNotInitialzed" means "the network number is not initialized".

[0110] Step 63: Obtain the second target data corresponding to each logical block address based on the mode index.

[0111] Among them, the second target data is the data that actually needs to be written during the data writing process restored based on each logical block address and the corresponding pattern index. Due to hardware problems and software problems of the storage device, there may be differences between the data to be written and the actually written data.

[0112] In some embodiments, when the read command is executed and completed, the data corresponding to each logical block address and the pattern index can be used to form the corresponding second target data.

[0113] Specifically, when the read command is executed and completed, the initial cached data of the corresponding data volume is obtained from the data buffer using the pattern index, and then the data at both ends of the initial cached data is replaced with the logical block address and the namespace identifier corresponding to the logical block address to obtain the second target data. In the embodiments of the present application, the second target data is obtained by adding the logical block address to both the head and the tail of the initial cached data, and the data volume of the second target data is the data volume of the logical block address.

[0114] Step 64: Perform a read verification based on the first target data and the second target data.

[0115] It can be understood that since the first target data is the data actually read from the storage device and the second target data is the data that needs to be written during writing, these two data may not be the same. Therefore, it is necessary to use the first target data to verify the second target data to confirm whether they are the same.

[0116] Specifically, when the first target data and the second target data are the same, it is determined that the data reading is successful. On the contrary, when the first target data and the second target data are not the same, it is determined that the read data fails.

[0117] It should be noted that the first target data and the second target data are generated following the same set of rules.

[0118] Different from the prior art, the data processing method provided in the present application sends a data read command to the storage device to enable the storage device to feedback the corresponding first target data. In response to the completion of the execution of the data read command, the corresponding pattern index is selected from the memory for each logical block address in the data read command, and then the second target data corresponding to each logical block address is obtained based on the pattern index, and a read verification is performed based on the first target data and the second target data. In the above manner, data reading and data read verification can be performed, and the speed of data reading and writing can also be accelerated. In addition, since the pattern index (generally four bits) has a small memory, the content occupied during the data reading and writing process is also small, and only a small amount of time is required to store it in the memory, thereby accelerating the overall speed of data reading and writing. This is more friendly to some host sides with small memory, and can also reduce a part of the memory overhead for host sides with sufficient large memory configuration.

[0119] In other embodiments, the first target data and the like can be mapped into a buffer. At this time, the data processing method corresponding to the above Figure 6 illustrated embodiment is applied to the host side and may include the following steps (not shown in the figure):

[0120] S1: Prepare a Read Data Buffer and an applied Expected Data Buffer.

[0121] Among them, the Expected Data Buffer is generated by the host driver's active application and can be used to store expected data.

[0122] S2: Send a data read command to the storage device and wait for completion.

[0123] S3: Determine whether the execution status of the data read command is successful.

[0124] If not, execute step S4; if so, execute step S5.

[0125] S4: Report an error.

[0126] S5: Select any LBA of the data read command.

[0127] S6: Obtain the Pattern Index corresponding to the LBA from the Pattern Buffer.

[0128] S7: Obtain the corresponding Pattern Value according to the Pattern Index.

[0129] S8: Compose a temporary Expected Buffer according to the LBA and the Pattern Value.

[0130] The data stored in the Expected Buffer is generated by the LBA and the Pattern Value. Physically, the host side does not record the Pattern Value but records the Pattern Index. And since there is a one-to-one correspondence between the Pattern Value and the Pattern Index, the Pattern Value is also known when the Pattern Index is known. In addition, since the index occupies less memory, for several fixed patterns, only 4 bits are needed to record the Pattern Index.

[0131] S9: Copy the temporary Expected Buffer data to the corresponding position of the Expected Data Buffer in step S1.

[0132] S10: Repeat steps S5 - S9 to traverse all LBAs until all LBAs in the data write command are filled into the Expected Data Buffer.

[0133] It should be noted that the Expected Data Buffer at this time is the same as the Write Buffer in the above - mentioned embodiment.

[0134] S11: Determine whether the Expected Data Buffer and the Read Data Buffer are consistent.

[0135] If they are consistent, it indicates that the data read verification is successful; otherwise, it indicates that the data read verification fails.

[0136] Different from the prior art, the data processing method provided by this application is the same as or similar to the Figure 6 illustrated embodiment, which can accelerate the data reading speed. And because the mode index (usually four - bit) has a small memory, the content occupied during the data reading process is also small, which is more friendly to some host - sides with small memory, and can also reduce a part of the memory overhead for host - sides with large enough memory configuration. Moreover, compared with the existing verification scheme in PyNVMe that does not contain LBA address information, the Data Buffer in this application contains LBA address information, which can save the process of continuously searching for the LBA to save time.

[0137] Refer to Figure 7 , Figure 7 FIG. is a schematic structural diagram of an embodiment of an electronic device provided by this application. The electronic device 70 includes a host - side 701 and a storage device 702. The host - side 701 is provided with memory. The host - side 701 and the storage device 702 are configured to implement the data processing method of any of the above - mentioned embodiments, which will not be elaborated here.

[0138] Refer to Figure 8 , Figure 8 FIG. is a schematic structural diagram of an embodiment of a computer - readable storage medium provided by this application. The computer - readable storage medium 80 stores a computer program 801, and the computer program is used to implement the data processing method of any of the above - mentioned embodiments when executed, which will not be elaborated here.

[0139] The storage media used in this application include various media that can store program codes, such as USB flash drives, external hard drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), or optical discs.

[0140] In summary, the data processing method provided by this application can achieve data reading and writing. Specifically, when the host receives a data writing command, the host actively generates write data and a write buffer, and then, according to the communication protocol between the host and the storage device, sends the address of the write buffer of the host to the storage device. Then, the storage device can obtain the data to be written from the write buffer of the host according to this address and the data writing command; when the host generates a data reading command, the host generates a read buffer, and then the host sends the address of the read buffer to the storage device according to the communication protocol with the storage device, so that the storage device can send the read data to the read buffer of the host according to this address and the data reading command.

[0141] Since the read data may be incorrect, it is necessary to perform data verification on the read data, that is, compare the data read from the read buffer with the data written in the write buffer. When the two are consistent, it indicates that the data reading is successful, otherwise it indicates that the data reading fails.

[0142] That is, when writing data at the host side, the data index is recorded, and when reading data at the host side, the actual data is restored through the data index recorded at that time, and then compared with the actual data read from the NAND flash memory to ensure data integrity.

[0143] The data processing method provided by this application can accelerate the data reading speed, and since the mode index (usually four bits) has a small memory, the content occupied during the data reading process is also small, which is more friendly to some host sides with small memory, and can also reduce a part of the memory overhead for host sides with a large enough memory configuration. And, compared with the existing verification scheme in PyNVMe that does not contain LBA address information, the Data Buffer in this application contains LBA address information, which can save the process of continuously searching for LBA to save time.

[0144] Moreover, it can effectively reduce the memory space of the host side required to store the written data, accelerate the reading and writing speed, and share the written data with the PyNVMe high-speed input / output workload engine (IO Worker) through the memory space (Memory Zone), so that the reading and writing modes of the IO Worker and the upper-layer Python script can be freely switched according to the actual scenario.

[0145] In addition, it is worth noting that the pattern index is obtained from the Pattern Recorder, which is integrated with the IO Worker engine. This not only improves the speed of data reading and writing, but also, since the IO Worker runs as a separate process during operation, the Pattern Recorder can share the memory of the Pattern Buffer with the IO Worker using the Memory Zone, thus achieving data sharing.

[0146] The above are only the embodiments of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present application.

Claims

1. A data processing method, characterized in that, Applied to the host side, the method includes: In response to a data write command, select a corresponding pattern index for each logical block address in the data write command; Based on the pattern index, obtain first cached data corresponding to each logical block address; In response to all the first cached data corresponding to the logical block addresses being written to the write buffer, send the data write command to the storage device; In response to the completion of the execution of the data write command, cache all the pattern indexes corresponding to the data write command in the memory, where the pattern index is used for read verification during data reading.

2. The method according to claim 1, wherein The obtaining first cached data corresponding to each logical block address based on the pattern index includes: Determine the data volume of the first cached data according to each logical block address; Use the data corresponding to each logical block address and the pattern index to form the first cached data that meets the data volume.

3. The method according to claim 2, wherein The using the data corresponding to each logical block address and the pattern index to form the first cached data that meets the data volume includes: Use the pattern index to obtain initial cached data corresponding to the data volume from the data cache area; Use the logical block address and the namespace identifier corresponding to the logical block address to replace the data at both ends of the initial cached data to obtain the first cached data.

4. The method according to claim 1, wherein The selecting a corresponding pattern index for each logical block address in the data write command in response to the data write command includes: In response to the data write command, randomly assign a pattern index from the pattern index array for each logical block address; where each pattern index corresponds to an element in the pattern index array, and some elements in the pattern index array are different fixed values, and the other part of the elements are random values.

5. A data processing method, characterized in that, Applied to the host side, the method includes: Send a data read command to the storage device to cause the storage device to feedback corresponding first target data; In response to the completion of the execution of the data read command, select a corresponding pattern index for each logical block address in the data read command from the memory; where the pattern index in the memory is obtained according to the method described in any one of claims 1-4; Based on the pattern index, obtain second target data corresponding to each logical block address; Perform read verification based on the first target data and the second target data.

6. The method according to claim 5, wherein The obtaining second target data corresponding to each logical block address based on the pattern index includes: Use the data corresponding to each logical block address and the pattern index to form corresponding second target data.

7. The method according to claim 6, characterized in that, The using the data corresponding to each logical block address and the pattern index to form corresponding second target data includes: Use the pattern index to obtain initial cached data corresponding to the data volume from the data cache area; Use the logical block address and the namespace identifier corresponding to the logical block address to replace the data at both ends of the initial cached data to obtain the second target data.

8. The method according to claim 5, characterized in that The performing read verification based on the first target data and the second target data includes: Determine that the data reading is successful in response to the consistency between the first target data and the second target data; Determine that the data reading fails in response to the inconsistency between the first target data and the second target data.

9. An electronic device, characterized in that, The electronic device includes a host side and a storage device. The host side is provided with a memory, and the host side and the storage device are configured to implement the method according to any one of claims 1-8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program, and the computer program is used to implement the method according to any one of claims 1-8 when executed.