Data processing method and device, electronic equipment, medium and computer program product
By using files with different storage modes to store data separately in the block storage system and using index information to quickly locate and read data, the inefficiency problem caused by random modifications in the block storage system is solved, and the data reading and storage efficiency is improved.
Patent Information
- Application Number
- CN202510796788.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-12
AI Technical Summary
Random modifications in block storage systems lead to inefficient data storage and reading, especially the read and write amplification problem of small IO data.
The first storage file and the second storage file with different storage modes are used to store data respectively, and the sub-data with small data length are quickly located using the index information of the first storage file, and the sub-data with large data length are read through the second storage file to avoid logical coupling and complex operations.
It improves data reading efficiency, reduces complex operations caused by coupling of processing logic of different data lengths, and improves storage and reading efficiency.
Smart Images

Figure CN120631271A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technology, specifically to the field of data reading and data storage technology, and more particularly to a data processing method, device, electronic device, medium, and computer program product. Background Art
[0002] In the block storage scenario, user modifications to files are treated as new write operations, and file modifications are implemented by gradually writing the modified data to the block storage system.
[0003] However, block storage usually has a large number of random modifications. The above modification method will cause the modified data to be stored in a multi-layer structure, resulting in a decrease in both data storage efficiency and reading efficiency. Summary of the Invention
[0004] The present disclosure provides a data processing method, apparatus, electronic device, medium, and computer program product.
[0005] According to one aspect of the present disclosure, a data processing method is provided, comprising: in response to receiving a data read request, determining a first storage file and a second storage file according to an identifier of data to be read in the data read request, wherein the first storage file and the second storage file have different storage modes and the data length of the second storage file is greater than the data length of the first storage file; determining an index result according to position information of the data to be read and first index information of the first storage file; in a case where the index result indicates that the first storage file includes first sub-data, reading the first sub-data from the first storage file, and the data to be read includes the first sub-data and the second sub-data; and reading the second sub-data from the second storage file according to the position information and at least one index information of the second storage file.
[0006] According to another aspect of the present disclosure, a data processing device is provided, including: a first determination module, for determining, in response to receiving a data read request, a first storage file and a second storage file according to an identifier of data to be read in the data read request, wherein the storage modes of the first storage file and the second storage file are different, and the data length of the second storage file is greater than the data length of the first storage file; a second determination module, for determining an index result according to position information of the data to be read and first index information of the first storage file; a first reading module, for reading the first sub-data from the first storage file when the index result indicates that the first storage file includes the first sub-data, and the data to be read includes the first sub-data and the second sub-data; and a second reading module, for reading the second sub-data from the second storage file according to the position information and at least one index information of the second storage file.
[0007] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above method.
[0008] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the method described above.
[0009] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein the computer program implements the method described above when executed by a processor.
[0010] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0012] Figure 1 Schematically illustrates an exemplary system architecture that can be applied to a data processing method and apparatus according to an embodiment of the present disclosure;
[0013] Figure 2 The following schematically shows a flow chart of a data processing method according to an embodiment of the present disclosure;
[0014] Figure 3 A schematic diagram of a scenario for determining index results according to an embodiment of the present disclosure is schematically shown;
[0015] Figure 4 Schematically shows an application scenario diagram of reading first sub-data according to an embodiment of the present disclosure;
[0016] Figure 5 Schematically shows a scenario diagram of reading the second sub-data according to an embodiment of the present disclosure;
[0017] Figure 6 Schematically illustrates a scenario diagram of reading the second sub-data based on the second index information and the third index information according to an embodiment of the present disclosure;
[0018] Figure 7 A schematic diagram schematically illustrates a scenario of updating the second index information after merging the data in the second storage file according to an embodiment of the present disclosure;
[0019] Figure 8 Schematically illustrates a flow chart of migrating data in a first storage file to a second storage file according to another embodiment of the present disclosure;
[0020] Figure 9 A flowchart of storing data to be stored in a second storage file according to another embodiment of the present disclosure is schematically shown;
[0021] Figure 10 A block diagram schematically shows a data processing device according to an embodiment of the present disclosure;
[0022] Figure 11 A block diagram schematically shows an electronic device suitable for implementing a data processing method according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0023] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0024] Storage based on erasure coding (EC) technology (EC storage) is typically used for block storage or file storage using block storage technology. For example, EC storage with 8 data blocks and 12 parity blocks, equivalent to 1.5 replicas, is suitable for read (input) and store (output) operations with relatively large data lengths (IOs), but not for small IOs. For example, EC storage with 8 data blocks and 12 parity blocks requires IO operations with a data length of at least 32KB. Block storage and file storage using EC storage often experience a large number of random modifications, especially for small data sizes (small IO data). Random modifications require data to be read, modified, and then written again, resulting in significant read / write amplification. EC storage, on the other hand, typically utilizes hierarchical data structures such as LSM-tree (Log Structured Merge Tree) to convert random modifications into append-only writes, aggregating small IO data into large IO data for writing. However, this append-write method not only couples the logic of small IO data and large IO data together, but also causes the read operation to require multiple searches along the layers of multiple append-writes to find the corresponding data, resulting in low reading efficiency.
[0025] Therefore, an embodiment of the present disclosure provides a data processing method, including: in response to receiving a data read request, determining a first storage file and a second storage file according to an identifier of the data to be read in the data read request, wherein the storage modes of the first storage file and the second storage file are different, and the data length of the second storage file is greater than the data length of the first storage file; determining an index result according to the position information of the data to be read and the first index information of the first storage file; when the index result indicates that the first storage file includes the first sub-data, reading the first sub-data from the first storage file, and the data to be read includes the first sub-data and the second sub-data; and reading the second sub-data from the second storage file according to the position information and at least one index information of the second storage file.
[0026] Figure 1 An exemplary system architecture that can be applied to a data processing method and apparatus according to an embodiment of the present disclosure is schematically shown.
[0027] It should be noted that Figure 1 The examples shown are merely examples of system architectures to which the embodiments of the present disclosure may be applied, to help those skilled in the art understand the technical content of the present disclosure, but do not mean that the embodiments of the present disclosure may not be used in other devices, systems, environments or scenarios.
[0028] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used as a medium for providing communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0029] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software (for example only).
[0030] The terminal devices 101 , 102 , and 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.
[0031] Server 105 may be a server that provides various services, such as a background management server (for example only) that supports content browsed by users using terminal devices 101, 102, and 103. The background management server may analyze and process received data such as user requests, and feed back processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the terminal device.
[0032] The server can be a cloud server, also known as a cloud computing server or cloud host. It is a hosting product within the cloud computing service system that addresses the management difficulties and poor scalability of traditional physical hosts and VPS services ("Virtual Private Servers" or "VPS"). The server can also be a distributed system server or a server integrated with blockchain.
[0033] It should be noted that the data processing method provided in the embodiment of the present disclosure can generally be executed by the server 105. Accordingly, the data processing apparatus provided in the embodiment of the present disclosure can also be set in the terminal device 101, 102, or 103. The data processing method provided in the embodiment of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105. Accordingly, the data processing apparatus provided in the embodiment of the present disclosure can also be set in a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105.
[0034] For example, a user may interact with terminal devices 101, 102, and 103, and terminal devices 101, 102, and 103 may generate and send a data read request to server 105. In response to receiving the data read request, server 105 may determine a first storage file and a second storage file based on an identifier of the data to be read in the data read request, wherein the first storage file and the second storage file have different storage modes and the data length of the second storage file is greater than the data length of the first storage file; determine an index result based on the location information of the data to be read and the first index information of the first storage file; if the index result indicates that the first storage file includes the first sub-data, read the first sub-data from the first storage file, and the data to be read includes the first sub-data and the second sub-data; and read the second sub-data from the second storage file based on the location information and at least one index information of the second storage file. Server 105 may return the read data to be read to terminal devices 101, 102, and 103.
[0035] It should be understood that Figure 1The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0036] In the embodiments of the present disclosure, the collection, storage, use, processing, transmission, provision, disclosure and application of user personal information involved comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good morals.
[0037] In the technical solution disclosed herein, the user's authorization or consent is obtained before obtaining or collecting the user's personal information.
[0038] Figure 2 The flowchart of the data processing method according to the embodiment of the present disclosure is schematically shown. Figure 2 As shown, embodiment 200 includes operations S210 to S240.
[0039] In operation S210 , in response to receiving a data read request, a first storage file and a second storage file are determined according to an identifier of the data to be read in the data read request.
[0040] The data read request includes an identifier of the data to be read, for example, the identifier may be an identifier of the file to which the data to be read belongs. In addition, the data read request may also include location information of the data to be read in the file, so that the data to be read can be read from the file based on the identifier and location information.
[0041] The data to be stored may include multiple data in multiple locations, and is not limited to the entire data stored in the same location, such as multiple data dispersedly stored in multiple locations in the first storage file and / or the second storage file.
[0042] Since the data to be read may be stored through multiple storage operations, and the lengths of the data stored in the multiple storage operations may be different, the data to be read may be stored in both the first storage file and the second storage file. For example, a portion of the data to be read may be stored in the first storage file, and another portion may be stored in the second storage file.
[0043] The first storage file and the second storage file have different storage modes, and the data length of the second storage file is greater than that of the first storage file. For example, within a certain length threshold, data with a length greater than the threshold is stored in the second storage file; otherwise, it is stored in the first storage file. In one embodiment, the storage mode of the first storage file can be a storage mode suitable for storing small data lengths, such as triple-copy storage in multi-copy storage; the storage mode of the second storage file can be a storage mode suitable for storing large data lengths, such as EC storage.
[0044] For example, the first storage file and the second storage file can be the upper layer (overlay layer) and lower layer (underlay layer) of the storage layer. For the first storage file stored in multi-copy storage mode, the storage layer can have a physical storage file for storage; for the second storage file stored in EC storage mode, it can be the general term for the physical space (value) mapped to the logical space (key) in the key-value database.
[0045] In operation S220 , an index result is determined according to the location information of the to-be-read data and the first index information of the first storage file.
[0046] The first index information is used to indicate the validity status of data stored at various locations in the first storage file. For example, if data A1 exists at location A in the first storage file, the first index information can indicate that the data stored at location A is valid.
[0047] For the data to be read, an index result can be determined based on the location information of the data to be read and the first index information of the first storage file. For example, if the first index information indicates that the data corresponding to the location information is in an invalid state, the index result indicates that the first storage file does not include the data to be read; if the first index information indicates that the portion of the data corresponding to the location information is in a valid state, the index result indicates that the first storage file includes the portion of the data to be read, such as the first sub-data.
[0048] In operation S230, in a case where the index result indicates that the first storage file includes the first sub-data, the first sub-data is read from the first storage file.
[0049] The data to be read includes first sub-data and second sub-data. Since the data to be read can be stored in the first storage file and the second storage file, for ease of understanding, the part or all of the data to be stored stored in the first storage file is referred to as the first sub-data, and the part or all of the data to be stored stored in the second storage file is referred to as the second sub-data.
[0050] If the data of the position of the first sub-data in the first index information is in a valid state, the first sub-data may be read from the first storage file according to the first index information.
[0051] In actual applications, most random modifications in block storage and file storage involve data with a small data length of less than 16KB. Therefore, when reading data, the first storage file storing the smaller data length is first searched based on the first index information to obtain an index result. For the second sub-data not included in the first storage file, the data is then read from the second storage file storing the larger data length in the following operation S240.
[0052] In operation S240, second sub data is read from the second storage file based on the location information and at least one index information of the second storage file.
[0053] Similar to the first storage file having a separate first index information, the second storage file also has at least one separate index information. The at least one index information is used to indicate the validity status of the data stored at each location in the second storage file. In some embodiments, the second storage file can directly locate the physical space of the data in the second storage file using a single index information; in other embodiments, the second storage file can use multiple index information to first locate the logical space of the data and then locate the physical space of the data.
[0054] For the second sub-data in the data to be read, the location information of the second sub-data can be determined based on the location information of the data to be read and the location information of the first sub-data; the second sub-data can be read from the second storage file based on the location information of the second sub-data. Alternatively, after determining the location information of the second sub-data, it can be determined based on at least one index information whether the second storage file includes the second sub-data; and if the second storage file includes the second sub-data, the second sub-data can be read from the second storage file.
[0055] In an embodiment of the present disclosure, by storing data of different lengths in a first storage file and a second storage file in different storage modes, and with the first storage file having a separate first index information and the second storage file having a separate at least one index information, data of different data lengths are decoupled in terms of storage and reading, avoiding the complex and inefficient operations caused by coupling the processing logic of different data lengths together, thereby improving reading efficiency. In addition, based on the fact that small data lengths are subject to more random modifications in block storage, the first storage file is used to first read part of the first sub-data of the target data from the first storage file, and then read the second sub-data from the second storage file, thereby avoiding the search and reading of data of large data lengths as much as possible, further improving reading efficiency.
[0056] According to an embodiment of the present disclosure, an index result is determined based on the location information of the data to be read and the first index information of the first storage file, including: determining a plurality of first numerical values that match the location information from the first index information, wherein the first numerical value is used to represent the valid status of the data in the first storage file; and determining the index result based on the plurality of first numerical values.
[0057] The data structure of the first index information may be a plurality of first numerical values arranged in order according to the position of the data in the physical space, and the first numerical value at a certain position indicates the valid state of the data stored at the position.
[0058] In one embodiment, each position in the first index information may be a data block with a data length of 4KB as the unit. The first value of each position indicates whether data occupies the 4KB data block at that position. If data occupies the data block, the data stored at that position is valid; conversely, if no data occupies the data block, the data stored at that position is invalid. Furthermore, since data length is relative, the unit of data blocks for each position in the first index information may vary. The unit may be determined by aligning the data block size in physical space, such as establishing the first index information in units of 512KB.
[0059] The location information of the data to be read may include an offset and a data length. The location information of the data to be read may be in the form of (offset, length). Based on the offset + the data length, the location of the data block that may be occupied by the data to be read can be located from the first index information.
[0060] For example, the location information of the data to be read may be (0K, 8K), (12K, 4K), or (16K, 4K). Through the above location information, the first and second data blocks (corresponding to data with an offset of 0K and a data length of 8K), the fourth data block (corresponding to data with an offset of 12K and a data length of 4K), and the first value of the fifth data block (corresponding to data with an offset of 16K and a data length of 4K) can be determined from the first index information.
[0061] For multiple first values determined from the first index information, the validity status of the data at each first value position can be determined based on the specific value of each first value. In a specific embodiment, the first index information can be a bitmap, and the first value can be 0 or 1, where 0 indicates that the data is invalid and 1 indicates that the data is valid.
[0062] In an embodiment of the present disclosure, by first determining multiple first numerical values that match the position information from the first index information, and determining the index result based on the multiple first numerical values, it is possible to determine whether the first storage file stores all or part of the data to be read based on the first index information alone of the first storage file, so that the reading operation of the data to be read can be performed based on the first storage file with a small data length as much as possible, avoiding the search and reading of data with a large data length, and further improving the reading efficiency.
[0063] Figure 3 The following schematically shows a scenario diagram for determining index results according to an embodiment of the present disclosure. Figure 3As shown, in embodiment 300, an identifier 302 and location information 303 of the to-be-read data 301 can be parsed from a data read request, and a first storage file 304 and a second storage file 305 can be determined based on the identifier 302. An index result 307 can be determined based on the first index information 306 and location information 303 of the first storage file 304.
[0064] According to an embodiment of the present disclosure, an index result is determined based on multiple first numerical values, including: when multiple first numerical values are all numerical values indicating a valid state, a first index result is determined, the first index result indicating that the first storage file includes data to be read; when multiple first numerical values are all numerical values indicating an invalid state, a second index result is determined, the second index result indicating that the second storage file includes data to be read; when some of the multiple first numerical values are numerical values indicating a valid state, a third index result is determined, the third index result indicating that the first storage file includes first sub-data, and the data at a position matching some of the first numerical values is the first sub-data.
[0065] The first storage file and the second storage file may be pre-allocated for data to be stored, but the first storage file or the second storage file does not currently store data to be read. Therefore, based on the multiple first values determined from the first index information, three different index results can be obtained, which are distinguished below by the first, second, and third index results.
[0066] When multiple first values are all values representing valid states, a first index result is obtained, and all the data to be read are stored in the first storage file. The above data to be read can be read directly from the first storage file without having to operate on at least one index information of the second storage file again or perform a read operation on the second storage file, thereby improving reading efficiency in random modification scenarios where data of small data length accounts for a large proportion.
[0067] When multiple first values are all values representing valid states, a second index result is obtained, and all the data to be read are stored in the second storage file. At this time, the data to be read is read from the second storage file, supporting random modification scenarios of large data lengths.
[0068] If some of the multiple first values are valid, the data to be stored is divided into first sub-data and second sub-data, which are stored in the first storage file and the second storage file, respectively. Since the embodiment of the present disclosure first searches the first index information of the first storage file, the embodiment of the present disclosure can read all the first sub-data in the first storage file before reading the second sub-data from the second storage file, thereby reducing the search and reading operations for large data lengths and improving reading efficiency.
[0069] In a specific embodiment, for a first storage file stored in a multi-copy storage mode and a second storage file stored in an EC storage mode, since the way in which EC storage randomly modifies data of small data length will cause a large read-write amplification problem, the embodiment of the present disclosure uses the EC storage mode to store data of larger data length, and when reading data, the first sub-data of small data length stored in the multi-copy storage model should be read first, and then the second sub-data is read from the second storage file. This can reduce the read-write amplification problem caused by the EC storage mode itself and further improve reading efficiency and storage efficiency.
[0070] Figure 4 The following schematically shows an application scenario diagram of reading the first sub-data according to an embodiment of the present disclosure. Figure 4 As shown, embodiment 400 includes a first storage file 403 and first index information 401 of the first storage file 403. The first index information 401 may be a bitmap, wherein the first values 0 / 1 represent the invalid state and the valid state of the data respectively.
[0071] According to the location information of the data to be read, multiple first values that match the location information can be determined from the first index information 401, such as three {1, 0, 1}. According to the specific data of the multiple first values, the index result 402 can be determined, such as the data at the two 1 positions in the multiple first values is valid, and the data at the 0 position is invalid. Since there is a fixed mapping relationship between the first index information 401 and the position of the first storage file 403, the data at the 32K~36K and 40K~44K positions can be directly read from the first storage file 403 based on the positions of the two "1"s in the first index information 401. The data at the 32K~36K and 40K~44K positions constitute the first sub-data 404. For the second sub-data missing at the 36K~40K position, it can be read from the second storage file.
[0072] According to an embodiment of the present disclosure, for operation S240, the second sub-data is read from the second storage file according to the position information and at least one index information of the second storage file, including: determining the starting physical address of the third sub-data in the second storage file according to the position information and the second index information of the second storage file, the third sub-data being data other than the first sub-data in the data to be read; and reading the second sub-data from the second storage file according to the starting physical address and the third index information of the second storage file, the second sub-data being valid data in the third sub-data.
[0073] The second storage file may include two index information, such as second index information and third index information. The second index information is used to indicate the starting physical address of each data item in the second storage file, and the third index information is used to indicate the physical location of each data item in the third storage file. The second index information can be considered as index information that aligns the logical space of the data with the physical space.
[0074] Since the first sub-data of the data to be stored can be stored in the first storage file, the remaining third sub-data of the data to be read and its position information can be determined based on the position information of the data to be read and the position information of the first sub-data. Therefore, based on the position information of the third sub-data, such as the offset, the starting physical address of the third sub-data in the second storage file can be determined in the second index information; then, based on the starting physical address, the length in the position information, and the third index information, the second sub-data can be located and read from the second storage file.
[0075] In one embodiment, if all the third sub-data in the second storage file are valid data, the second sub-data and the third sub-data are identical. In another embodiment, if only a portion of the valid data in the third sub-data is stored in the second storage file, in this case, the second sub-data read from the second storage file are different from the third sub-data, and the second sub-data is only a portion of the third sub-data.
[0076] It is understood that if the starting physical address of the third sub-data in the second storage file cannot be determined based on the position information and the second index information, this indicates that all third sub-data in the second storage file are invalid. In this case, the second sub-data cannot be read from the second storage file, and further processing of the third index information can be omitted. A direct feedback of 0 can be provided to indicate that the second sub-data was not read, thereby improving reading efficiency. Alternatively, if the third index information determines that all third sub-data in the second storage file are invalid, a feedback of 0 can be provided to indicate that the second sub-data was not read, thereby improving reading efficiency.
[0077] In an embodiment of the present disclosure, the starting physical address of the third sub-data in the second storage file is determined based on the position information and the second index information of the second storage file; the second sub-data is read from the second storage file based on the starting physical address and the third index information of the second storage file, and the starting physical address of the third sub-data can be quickly located based on the second index information to improve the efficiency of subsequent reading of the second sub-data in a valid state in the third sub-data.
[0078] Figure 5 The schematic diagram of the scene of reading the second sub-data according to the embodiment of the present disclosure is shown schematically. Figure 5As shown, in embodiment 500, second storage file 501 includes second index information 503 and third index information 502. Location information 504 of the third sub-data can be determined based on the location information of the data to be read and the location information of the first sub-data. Furthermore, a starting physical address 505 of the third sub-data in the second storage file can be determined based on the location information 504 of the third sub-data and the second index information 503. Based on the starting physical address 505 and the third index information 502, second sub-data 506 can be read from the second storage file.
[0079] According to an embodiment of the present disclosure, the second sub-data is read from the second storage file according to the starting physical address and the third index information of the second storage file, including: obtaining a target number of second values located after the starting physical address from the third index information, the target number being determined based on the data length of the third sub-data; and reading the second sub-data from the second storage file according to the target number of second values, at least one second value among the target number of second values being a value representing a valid state, and the data at a position matching the at least one second value is the second sub-data.
[0080] For example, the data structure of the third index information may be a plurality of second values arranged in order of the data's location in physical space, and the data structure of the second index information may be a plurality of third values arranged in order of the data's starting physical address. For example, if the third value of the starting physical address B is a predetermined value, then the starting physical address of data B1 is B.
[0081] In a specific embodiment, both the second index information and the third index information can be bitmaps. When the third value in the second index information is 1 (that is, the predetermined value is 1), the starting physical address of the data B1 is B; when the second value in the third index information is 1, it indicates that the data at this position is in a valid state.
[0082] It should be noted that, according to the order of appearance of the second index information and the third index information, the value in the third index information is referred to as the second value, and the value in the second index information is referred to as the third value.
[0083] In one embodiment, each position in the third index information represents a data block with a data length of 4KB. The second value of each position indicates whether data occupies the 4KB data block at that position. If data occupies that data block, the data stored at that position is valid. For the third sub-data, if the position information of the third sub-data is (0KB, 1024KB), (1536KB, 512KB), or (2048KB, 512KB), the target number can be 1024KB / 4KB = 256, 1536KB / 4KB = 384, or 2048KB / 4KB = 512.
[0084] Similar to determining multiple first values from the first index information, after determining the starting physical address and data length of the third sub-data, a target number of second values in the third index information that match the third sub-data can be determined. Based on whether the target number of second values are valid values, it can be determined whether the second storage file contains valid data. If all of the target number of second values in the third index information are valid values, the second sub-data is the third sub-data. If only at least one of the target number of second values is valid, the data at the location matching the at least one second value is the second sub-data, and only the second sub-data can be read from the second storage file.
[0085] In an embodiment of the present disclosure, a target number of second values located after the starting physical address are obtained from the third index information, and second sub-data in a valid state are read from the second storage file based on the target number of second values. When a read operation is performed, only data in a valid state is read based on the third index information, and there is no need to read data in an invalid state, thereby further improving the reading efficiency.
[0086] Figure 6 The following schematically shows a scenario diagram of reading the second sub-data based on the second index information and the third index information according to an embodiment of the present disclosure. Figure 6 As shown, in embodiment 600, the physical space of the second storage file 601 stores three pieces of data, and the physical addresses are (0K, 1024K), (1536K, 512K), and (2048K, 512K) respectively.
[0087] The offset in the location information of the third sub-data can be understood as a logical address in the logical space. The logical address is 0K. From the second index information 602 of the second storage file 601, it can be determined that the starting physical address of the third sub-data is also 0K. Therefore, based on the data length of the third sub-data (1024K) and the starting physical address (0K), a target number of second values is determined from the third index information 603. For example, the second values at positions 0 through 255 in the third index information 603 total 256. In this embodiment, all 256 second values are 1, indicating that the stored data is valid. Therefore, the data at (0K, 1024K) can be read from the first storage file. This data is the second sub-data. In this embodiment, the second sub-data and the third sub-data are identical.
[0088] According to an embodiment of the present disclosure, the data processing method also includes: determining, based on the second index information, a first index interval in the second storage file whose fragmentation rate exceeds a predetermined threshold; and merging the data in the second storage file based on the third index information and the first index interval.
[0089] In an embodiment of the present disclosure, for a second storage file storing data of a larger length, considering that randomly modified data of a small length may cause data at some locations in the second storage file to become invalid data, and further cause data of a large length stored in the second storage file to become fragmented valid data, therefore, data merging (compaction) of the data in the second storage file may be performed periodically to improve the locality of data access.
[0090] Since the second index information is used to represent the starting physical address of each data in the second storage file, the number of data included in the second storage file can be determined based on the second index information, thereby measuring the degree of data fragmentation in the second storage file.
[0091] For example, the second index information can be divided into multiple index intervals, and the fragmentation rate of each index interval is determined based on the ratio of the number of third values in the index interval that are predetermined values to the total number of third values. The predetermined threshold can be determined based on actual needs.
[0092] For the first index interval whose fragmentation rate exceeds the predetermined threshold, it indicates that the degree of fragmentation of the data in the first index interval is high. Therefore, the data in the first index interval in the second storage file can be merged according to the third index information and the first index interval.
[0093] In an embodiment of the present disclosure, by determining the first index interval in the second storage file where the fragmentation rate exceeds a predetermined threshold based on the second index information, the degree of data fragmentation in the second storage file can be quickly determined without reading the data, and the data in the second storage file can be merged based on the third index information and the first index interval. As a result, the embodiment of the present disclosure can accurately and efficiently identify severely fragmented data in the second storage file (such as data stored in EC storage mode) and merge this data, while avoiding the serious read-write amplification problem caused by the sorting and merging.
[0094] According to an embodiment of the present disclosure, based on the second index information, determining the first index interval in the second storage file whose fragmentation rate exceeds a predetermined threshold includes: obtaining multiple third numerical values of each second index interval in the second index information; determining the fragmentation rate of each second index interval based on the multiple third numerical values; and determining the second index interval whose fragmentation rate exceeds the preset threshold as the first index interval.
[0095] Multiple second index intervals of preset lengths can be determined as needed. For example, a sliding window of preset length can be set, and using this sliding window, starting from the head of the second stored file, multiple second index intervals of preset lengths can be obtained. Alternatively, starting from the head of the second stored file, the second stored file can be divided into multiple second index intervals of preset lengths. The length of the last second index interval can be less than or greater than the preset length, as long as the fragmentation rate of the index interval can be determined.
[0096] For each second index interval, a fragmentation rate can be determined based on the ratio of the number of third values that are predetermined values to the total number of third values in the second index interval. A second index interval with a fragmentation rate exceeding a preset threshold can be determined as a first index interval.
[0097] For example, the second index information can be a bitmap, where each 256 bits represents 1MB of address space, and the second index interval is divided into 256-bit intervals with a preset length, and the preset value is 1. Thus, based on the number of bits with a third value of 1 in the 256 bits and the total number of third values, such as 256, the fragmentation rate of each second index interval can be determined. For example, the fragmentation rate x = n / 256, where n is the number of bits with a value of 1. The more bits with a value of 1 in the second index interval, the more data in the second index interval, and the higher the fragmentation rate of the second index interval.
[0098] For example, the second index interval with a fragmentation rate greater than 50% may be determined as the first index interval that needs to be merged.
[0099] In an embodiment of the present disclosure, by dividing the second index information into multiple second index intervals, it is possible to further evaluate the degree of fragmentation of local data in the second index information, and accurately identify intervals with high local fragmentation in the second index information, thereby achieving finer-grained data merging, thereby improving the efficiency of local data reading while reducing the problem of data read amplification.
[0100] According to an embodiment of the present disclosure, data in the second storage file is merged according to the third index information and the first index interval, including: obtaining multiple second numerical values located in the first index interval in the third index information; reading valid data from the second storage file according to the multiple second numerical values; merging the valid data; and re-storing the merged valid data to the second storage file and updating the second index information.
[0101] There may be multiple valid data scattered within the first index interval, and the third index information can indicate the valid status of the data at a certain position through the second value. Therefore, after determining the first index interval, multiple second values located in the first index interval in the third index information are obtained, and it is possible to determine which data in the first index interval is valid data and perform a read operation.
[0102] For example, both the second index information and the third index information may be bitmaps, and the first index interval determined based on the second index information may be a 256-bit interval ranging from 0 to 255. A second value of the 256 bits is obtained from the third index information. The second value may be 0 or 1, where 1 indicates that the data at that location is valid.
[0103] For the valid data that has been determined, it can be directly read from the second storage file according to the third index information, without having to read other invalid data from the second storage file, so as to reduce the amount of data read. For the valid data read out, two or more valid data can be merged to obtain the merged valid data. It is understandable that the amount of valid data after the merger is less than the amount of valid data before the merger. Therefore, after the merged valid data is stored in the second storage file, the degree of data fragmentation in the second storage file is reduced.
[0104] Since the amount of valid data changes before and after the merge operation, the starting physical address of each data also changes. Therefore, when the merged valid data is stored again in the second storage file, the second index information needs to be updated synchronously.
[0105] In the embodiment of the present disclosure, since the location of valid data in the second storage file can be determined based on multiple second numerical values located in the first index interval in the third index information, the embodiment of the present disclosure can read only valid data from the second storage file, without having to read all the data in the second storage file and then perform validity judgment and merging, thereby avoiding the read-write amplification problem caused by compacting the second storage file.
[0106] Figure 7 The following schematically illustrates a scenario in which the second index information is updated after the data in the second storage file is merged according to an embodiment of the present disclosure. Figure 7 As shown, embodiment 700 includes second index information 701 before merging and second index information 702 after merging.
[0107] In the second index information 701 before the merge, the fragmentation rate of the first index interval is 3 / 5=60%, which is greater than 50%. Therefore, the three valid data in the first index interval can be merged and re-stored to the 0k position. In the second index information 702 after the merge, the three valid data become one valid data.
[0108] Figure 8 The flowchart of migrating data in a first storage file to a second storage file according to another embodiment of the present disclosure is schematically shown.
[0109] like Figure 8 As shown, embodiment 800 includes operations S810 to S830.
[0110] In operation S810, location information of data to be migrated is determined from a first storage file according to first index information.
[0111] The data to be migrated includes multiple valid data stored at consecutive addresses in the first storage file. Compared with the second storage file, the first storage file is used to store data of smaller data length. However, during multiple random modifications, the data of smaller data length may accumulate into data of larger data length. In this case, the multiple data of smaller data length may be stored in the storage mode of larger data length.
[0112] For example, taking the data length of the second storage file as the threshold, if the total data length of multiple valid data at consecutive addresses in the first storage file is greater than or equal to the data length of the second storage file, then the multiple valid data at the consecutive addresses are determined as data to be migrated.
[0113] Since the first index information is used to indicate the validity status of the data stored at each location in the first storage file, the location information of the data to be migrated can be determined through the first index information.
[0114] In operation S820, the data to be migrated is read and deleted from the first storage file according to the location information of the data to be migrated, and the data to be migrated is stored in the second storage file in the storage mode of the second storage file.
[0115] Based on the location information of the data to be migrated, multiple valid data items at consecutive addresses can be read from the first storage file. These multiple valid data items at consecutive addresses are then treated as a complete set of data to be migrated and stored using the storage mode of the second storage file. To avoid duplicate data to be migrated between the first and second storage files, the data to be migrated is deleted from the first storage file after it is read from the first storage file.
[0116] In operation S830, the first index information and at least one index information of the second storage file are updated.
[0117] After migrating the data to be migrated from the first storage file to the second storage file, in order to facilitate subsequent accurate reading of the data to be migrated, the first index information of the first storage file and at least one index information of the second storage file need to be updated.
[0118] For example, if the first index information is a bitmap, and there are 150 consecutive ones starting from 0K in the first index information, then the total data length of the valid data at the corresponding positions of the multiple ones is 4K*150=600K, which is greater than the data length of the second storage file of 500K. In this case, the 150 consecutive data from 0K to 600K can be determined as the data to be migrated, and their location information is (0K, 4K), (4K, 4K), (8K, 4K), (12K, 4K), etc. After each piece of data to be migrated is read from the first storage file one by one according to the determined location information, the 150 data are concatenated into data of large data length and stored in the storage mode corresponding to the second storage file, such as EC storage. To avoid duplicate data to be migrated in the first and second storage files, after reading the data to be migrated from the first storage file, the data to be migrated is synchronously deleted from the first storage file and the first value of the corresponding position in the first index information is updated to 0.
[0119] Similarly, when storing the data to be migrated in the second storage file, at least one index information of the second storage file is updated so that the data to be migrated can be subsequently read from the second storage file. For example, if the second storage file includes the second index information and third index information in the bitmap format described above, the third value of the starting physical address where the data to be migrated is stored in the second index information can be updated to 1, and the second values of the corresponding positions of the data to be migrated in the third index information can all be updated to 1.
[0120] In an embodiment of the present disclosure, for the first storage file and the second storage file that are decoupled from each other, for data with small data length, by determining the location information of the data to be migrated from the first storage file based on the first index information, reading and deleting the data to be migrated from the first storage file based on the location information of the data to be migrated, and storing the data to be migrated in the second storage file in the storage mode of the second storage file, the small data length and scattered valid data can be merged into large data length data and flushed to the second storage file, thereby reducing the number of small data lengths, thereby improving the processing efficiency of small data lengths, and improving the locality of data reading.
[0121] Figure 9The flowchart of storing the data to be stored into the second storage file according to another embodiment of the present disclosure is schematically shown. Figure 9 As shown, embodiment 900 includes operations S910 to S930. Operations S910 to S930 can be used as an embodiment of storing data in the first storage file or the second storage file during random modification and / or data writing. The data to be read and the data to be migrated described above can be stored in the first storage file or the second storage file through operations S910 to S930.
[0122] In operation S910 , in response to receiving a data storage request, a data length of data to be stored is determined according to the data storage request.
[0123] In operation S920 , when the data length of the data to be stored is greater than the length threshold, the data to be stored is stored in the second storage file.
[0124] In operation S930, at least one index information of the second storage file is updated.
[0125] The data storage request may include data length, such as length. By parsing the data storage request, the data length of the data to be stored may be obtained.
[0126] The first storage file and the second storage file can be divided according to a predetermined length threshold. For example, the length threshold can be 4K, 500K, etc., and the length threshold can be determined according to actual needs. When the length of the data to be stored is less than or equal to the length threshold, the data to be stored is stored in the first storage file in the storage mode of the first storage file, and the first index information of the first storage file is updated. Otherwise, the data is stored in the second storage file, and at least one index information of the second storage file is updated.
[0127] For example, when the data length of the data to be stored is greater than the length threshold, the data to be stored is stored in the second storage file in the ec storage mode, and the second index information and the third index information of the first storage file are updated; when the data length of the data to be stored is less than or equal to the length threshold, the data to be stored is stored in the first storage file in a multi-copy storage mode, such as a 3-copy storage mode, and the first index information is updated.
[0128] In the embodiment of the present disclosure, for data to be stored with a small data length, it is directly written into the first storage file, and there is no need to wait until the data of the small data length is accumulated to the large data length before it can be stored. Therefore, the embodiment of the present disclosure can avoid the write delay of accumulating the small data length to the large data length. In addition, for data to be stored with a small data length, only the data in the first storage file and the index of the first storage file are modified, and the processing operation of the small data length data is decoupled from the second storage file with the large data length, avoiding the processing operation of the small data length affecting the processing operation of the large data length, thereby improving the processing efficiency of the small data length. When the data length of the data to be stored is greater than the length threshold, the storage mode suitable for the large data length is used for storage, and the read-write amplification problem caused by storing the small data length in the storage mode corresponding to the second storage file is avoided as much as possible.
[0129] For example, for the embodiment of using EC storage for the second storage file and multi-copy storage for the first storage file, by storing data of small data length through multiple copies and storing data of large data length through EC, the problems of EC storage for random modification, read and write amplification of small data length, low IO efficiency, etc. can be solved, thereby improving data reading and storage efficiency.
[0130] In a specific embodiment, considering that the data length is unevenly distributed in actual applications, the storage and reading (IO) of data less than or equal to 16K (small data length) and greater than or equal to 500K (large data length) are mainly performed.
[0131] For example, the length threshold can be 16KB, and data with a length less than or equal to 16KB is stored in the first storage file. Data in the range of 16KB to 500KB and data greater than or equal to 500KB are both considered as data with a larger data length and stored in the second storage file to reduce the amount of data with a smaller data length.
[0132] Alternatively, the length threshold can be 500KB, and data with a length less than or equal to 500KB is stored in the first storage file. Data in the range of 16KB to 500KB and data less than or equal to 16KB are both considered data with a smaller length and stored in the first storage file. Because data is read from the first storage file first, storing data in the range of 16KB to 500KB in the first storage file can improve the reading efficiency of this portion of data.
[0133] In a specific embodiment, for data to be stored whose length is greater than a length threshold, since the data to be stored may be modified data (including random modifications and data modified by user operations), the data to be stored may be involved in the first storage file. Therefore, in addition to operations S910 to S930, the data processing method also includes: performing historical data detection on the first storage file based on the location information of the data to be stored and the first index information to obtain a historical detection result; when the historical detection result indicates that the first storage file includes historical data associated with the data to be stored, deleting the historical data from the first storage file; and updating the first index information.
[0134] In the case where the data to be stored is modified data, the data storage request includes the location information of the data to be stored. Historical data detection is performed based on the location information of the data to be stored and the first index information of the first storage file to obtain a historical detection result. For example, multiple fourth values that match the location information of the data to be stored are determined from the first index information. In the case where the multiple fourth values are all values indicating an invalid state, the historical detection result indicates that the first storage file does not include historical data associated with the data to be stored; conversely, in the case where at least one fourth value among the multiple fourth values is a value indicating a valid state, the historical detection result indicates that the first storage file includes historical data associated with the data to be stored. At this time, the data at the position corresponding to the at least one fourth value is historical data, and the historical data can be deleted from the first storage file, and the at least one fourth value in the first index information is updated to a value indicating an invalid state. The fourth value in the first index information is similar to the first value and is only used here to distinguish between the reading and storage processes.
[0135] In an embodiment of the present disclosure, if the historical data in the first storage file is not deleted when storing data, when data needs to be read, since the data reading operation starts from the first storage file first, the reading operation will read the historical data, resulting in the read data not being the latest data, affecting the accuracy of the read data.
[0136] Therefore, in an embodiment of the present disclosure, if the length of the data to be read exceeds a length threshold, a historical data check is performed on the first storage file based on the location information of the data to be stored and the first index information to obtain a historical data check result. If the historical data check result indicates that the first storage file includes historical data associated with the data to be stored, the historical data is deleted from the first storage file, and the first index information is updated. Thus, when the data to be stored is stored in the second storage file, the embodiment of the present disclosure can avoid duplicate historical data in the first storage file, thereby preventing the accuracy of the read data from being affected.
[0137] Figure 10The block diagram schematically shows a data processing device according to an embodiment of the present disclosure.
[0138] like Figure 10 As shown, the data processing device 1000 includes a first determining module 1010 , a second determining module 1020 , a first reading module 1030 and a second reading module 1040 .
[0139] The first determination module 1010 is used to determine a first storage file and a second storage file in response to receiving a data read request and according to an identifier of the data to be read in the data read request, wherein the storage modes of the first storage file and the second storage file are different, and the data length of the second storage file is greater than the data length of the first storage file.
[0140] The second determining module 1020 is configured to determine an indexing result according to the location information of the to-be-read data and the first index information of the first storage file.
[0141] The first reading module 1030 is configured to read the first sub-data from the first storage file when the index result indicates that the first storage file includes the first sub-data, and the data to be read includes the first sub-data and the second sub-data.
[0142] The second reading module 1040 is configured to read the second sub-data from the second storage file according to the position information and at least one index information of the second storage file.
[0143] According to an embodiment of the present disclosure, the second determining module 1020 includes a first determining submodule and a second determining submodule.
[0144] The first determining submodule is configured to determine a plurality of first numerical values matching the position information from the first index information, wherein the first numerical values are used to represent a valid state of the data in the first storage file.
[0145] The second determining submodule is configured to determine an index result according to the plurality of first numerical values.
[0146] According to an embodiment of the present disclosure, the second determining submodule includes: a first determining unit, a second determining unit, and a third determining unit.
[0147] The first determining unit is configured to determine a first index result when all of the plurality of first values represent valid states, the first index result indicating that the first storage file includes data to be read.
[0148] The second determining unit is configured to determine a second index result when all of the plurality of first values represent an invalid state, the second index result indicating that the second storage file includes data to be read.
[0149] The third determination unit is used to determine a third index result when some of the first values among the multiple first values are values representing a valid state, and the third index result indicates that the first storage file includes first sub-data, and the data at the position matching the part of the first values is the first sub-data.
[0150] According to an embodiment of the present disclosure, the second reading module 1040 includes: a first reading submodule and a second reading submodule.
[0151] The first reading submodule is used to determine the starting physical address of the third sub-data in the second storage file according to the position information and the second index information of the second storage file, where the third sub-data is the data to be read except the first sub-data.
[0152] The second reading submodule is configured to read second sub-data from the second storage file according to the starting physical address and the third index information of the second storage file, where the second sub-data is valid data in the third sub-data.
[0153] According to an embodiment of the present disclosure, the second reading submodule includes:
[0154] The acquiring unit is configured to acquire a target number of second values located after the starting physical address from the third index information, where the target number is determined according to the data length of the third sub-data.
[0155] The first reading unit is used to read second sub-data from the second storage file according to a target number of second values, at least one second value among the target number of second values is a value representing a valid state, and the data at a position matching the at least one second value is the second sub-data.
[0156] According to an embodiment of the present disclosure, the data processing device 1000 further includes a third determining module, configured to determine, based on the second index information, a first index interval in the second storage file whose fragmentation rate exceeds a predetermined threshold.
[0157] The merging module is used to merge the data in the second storage file according to the third index information and the first index interval.
[0158] According to an embodiment of the present disclosure, the third determination module includes:
[0159] The acquisition module is configured to acquire a plurality of third values of each second index interval in the second index information.
[0160] The third determining submodule is configured to determine a fragmentation rate of each second index interval according to a plurality of third values.
[0161] The fourth determining submodule is configured to determine the second index interval whose fragmentation rate exceeds a preset threshold as the first index interval.
[0162] According to an embodiment of the present disclosure, the merging module includes:
[0163] The data acquisition submodule is configured to acquire a plurality of second values located within the first index interval in the third index information.
[0164] The third reading submodule is used to read valid data from the second storage file according to the multiple second numerical values.
[0165] The merging submodule is used to merge valid data.
[0166] The storage submodule is used to store the merged valid data back into the second storage file and update the second index information.
[0167] According to an embodiment of the present disclosure, the data processing device 1000 further includes:
[0168] The fourth determining module is configured to determine location information of the data to be migrated from the first storage file according to the first index information.
[0169] The third reading module is configured to read and delete the data to be migrated from the first storage file according to the location information of the data to be migrated, and store the data to be migrated in the second storage file in the storage mode of the second storage file.
[0170] The first updating module is used to update the first index information and at least one index information of the second storage file.
[0171] According to an embodiment of the present disclosure, the data processing device 1000 further includes:
[0172] The fifth determining module is configured to, in response to receiving the data storage request, determine the data length of the data to be stored according to the data storage request.
[0173] The storage module is used to store the data to be stored in the second storage file when the data length of the data to be stored is greater than the length threshold.
[0174] The second updating module is used to update at least one index information of the second storage file.
[0175] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0176] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described above.
[0177] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the method described above.
[0178] According to an embodiment of the present disclosure, a computer program product includes a computer program, and when the computer program is executed by a processor, the computer program implements the method described above.
[0179] Figure 11 A block diagram of an electronic device suitable for implementing a data processing method according to an embodiment of the present disclosure is schematically shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0180] like Figure 11 As shown, electronic device 1100 includes a computing unit 1101, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1102 or a computer program loaded from a storage unit 1108 into a random access memory (RAM) 1103. RAM 1103 may also store various programs and data required for the operation of electronic device 1100. Computing unit 1101, ROM 1102, and RAM 1103 are connected to each other via a bus 1104. An input / output (I / O) interface 1105 is also connected to bus 1104.
[0181] Multiple components in electronic device 1100 are connected to an input / output (I / O) interface 1105, including an input unit 1106, such as a keyboard and mouse; an output unit 1107, such as various types of displays and speakers; a storage unit 1108, such as a magnetic disk and optical disk; and a communication unit 1109, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1109 allows electronic device 1100 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0182] The computing unit 1101 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 1101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1101 performs the various methods and processes described above, such as the data processing method. For example, in some embodiments, the data processing method may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 1108. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 1100 via the ROM 1102 and / or the communication unit 1109. When the computer program is loaded into the RAM 1103 and executed by the computing unit 1101, one or more steps of the data processing method described above may be performed. Alternatively, in other embodiments, the computing unit 1101 may be configured to perform the data processing method by any other suitable means (e.g., via firmware).
[0183] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0184] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0185] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0186] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0187] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0188] A computer system may include clients and servers. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0189] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0190] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A data processing method, comprising: In response to receiving a data read request, determining a first storage file and a second storage file according to an identifier of the data to be read in the data read request, wherein the first storage file and the second storage file have different storage modes, and a data length of the second storage file is greater than a data length of the first storage file; determining an index result according to the location information of the data to be read and the first index information of the first storage file; If the index result indicates that the first storage file includes the first sub-data, reading the first sub-data from the first storage file, the data to be read includes the first sub-data and the second sub-data; and The second sub-data is read from the second storage file according to the position information and at least one index information of the second storage file.
2. The method according to claim 1, wherein The determining the index result according to the location information of the to-be-read data and the first index information of the first storage file includes: Determining a plurality of first values matching the position information from the first index information, wherein the first values are used to represent a valid state of data in the first storage file; and The index result is determined according to the multiple first values.
3. The method according to claim 2, wherein: The determining the index result according to the multiple first values includes: In a case where the plurality of first values are all values indicating a valid state, determining a first index result, where the first index result indicates that the first storage file includes the data to be read; In a case where the plurality of first values are all values indicating an invalid state, determining a second index result, wherein the second index result indicates that the second storage file includes the data to be read; When some of the multiple first values are values representing a valid state, a third index result is determined, and the third index result indicates that the first storage file includes the first sub-data, and the data at the position matching the some of the first values is the first sub-data.
4. The method according to claim 1, wherein The reading the second sub-data from the second storage file according to the position information and at least one index information of the second storage file includes: determining, based on the position information and the second index information of the second storage file, a starting physical address of third sub-data in the second storage file, the third sub-data being data in the data to be read excluding the first sub-data; and The second sub-data is read from the second storage file according to the starting physical address and the third index information of the second storage file, where the second sub-data is valid data in the third sub-data.
5. The method according to claim 4, wherein The reading the second sub-data from the second storage file according to the starting physical address and the third index information of the second storage file includes: Obtaining a target number of second values located after the start physical address from the third index information, wherein the target number is determined according to the data length of the third sub-data; and According to the target number of second values, the second sub-data is read from the second storage file, at least one second value among the target number of second values is a value representing a valid state, and the data at the position matching the at least one second value is the second sub-data.
6. The method according to any one of claims 2 to 5, further comprising: Determining, based on the second index information, a first index interval in the second storage file whose fragmentation rate exceeds a predetermined threshold; as well as The data in the second storage file is merged according to the third index information and the first index interval.
7. The method according to claim 6, wherein: The determining, according to the second index information, a first index interval in the second storage file in which the fragmentation rate exceeds a predetermined threshold includes: Obtaining multiple third values of each second index interval in the second index information; determining a fragmentation rate of each second index interval according to the plurality of third values; and The second index interval whose fragmentation rate exceeds the preset threshold is determined as the first index interval.
8. The method according to claim 6, wherein: The merging of the data in the second storage file according to the third index information and the first index interval includes: Obtaining a plurality of second values in the third index information that are within the first index interval; Reading valid data from the second storage file according to the plurality of second values; Merging the valid data; and The merged valid data is stored again in the second storage file, and the second index information is updated.
9. The method according to any one of claims 1 to 8, further comprising: determining location information of the data to be migrated from the first storage file according to the first index information; Reading and deleting the data to be migrated from the first storage file according to the location information of the data to be migrated, and storing the data to be migrated in the second storage file in the storage mode of the second storage file; The first index information and at least one index information of the second storage file are updated.
10. The method according to any one of claims 1 to 9, further comprising: In response to receiving a data storage request, determining a data length of the data to be stored according to the data storage request; When the length of the data to be stored is greater than a length threshold, storing the data to be stored in the second storage file; as well as At least one index information of the second storage file is updated.
11. A data processing device comprising: a first determining module, configured to, in response to receiving a data read request, determine a first storage file and a second storage file according to an identifier of data to be read in the data read request, wherein the first storage file and the second storage file have different storage modes, and a data length of the second storage file is greater than a data length of the first storage file; a second determining module, configured to determine an index result according to location information of the to-be-read data and first index information of the first storage file; a first reading module, configured to read the first sub-data from the first storage file if the index result indicates that the first storage file includes the first sub-data, and the data to be read includes the first sub-data and the second sub-data; The second reading module is configured to read the second sub-data from the second storage file according to the position information and at least one index information of the second storage file.
12. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 10.
13. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable the computer to execute the method according to any one of claims 1 to 10.
14. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method according to any one of claims 1 to 10.