Data processing method and device, equipment and storage medium

By clearing junk data related to the target data during the write process on the storage platform, the problems of IOPS decline and low garbage data recycling caused by limited CPU resources are solved, and more efficient data processing and junk data recycling are achieved.

CN120104056APending Publication Date: 2025-06-06CHINA ELECTRONICS CLOUD DIGITAL INTELLIGENCE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510164357.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

With limited CPU resources, existing storage platforms adopt ROW to improve data coverage and write speed, resulting in a decrease in front-end IOPS and low efficiency of garbage data recycling.

Method used

It provides a data processing method, which receives a user's write request, writes the target data to the storage space, determines the garbage data related to the target data, and clears the garbage data during the writing process, reducing the execution time of the background garbage collection task.

Benefits of technology

By timely releasing data space, improving the number of input and output operations per second in the front desk, improving the efficiency of garbage data recycling, solving the problems of IOPS decline and garbage data recycling caused by limited CPU resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104056A_ABST
    Figure CN120104056A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and device, equipment and a storage medium. The method comprises the steps that a write-in request of a user for target data is received, the target data is written into a corresponding storage space, junk data related to the target data is determined, the junk data is cleared, and a write-in state for the target data is returned. By the adoption of the technical scheme, the data processing device deletes the junk data related to the target data in the writing process of the target data by the user, the space of the data can be released in time, and therefore the execution time of a background junk recycling task is shortened, the number of times of input and output operation per second of a foreground is increased, and user experience is improved. And the recovery efficiency of the junk data is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing, and in particular to a data processing method, device, equipment and storage medium. Background Art

[0002] In the current storage field, in order to increase the speed of data overwriting, storage platforms use the ROW (Redirect-on-write) method for data, and at the same time start a background task to regularly clear useless data and recycle unused data space. However, in the case of limited CPU (Central Processing Unit) resources, this method affects the front-end IOPS (Input / Output Operations Per Second) and leads to low efficiency in garbage data recovery. Summary of the invention

[0003] In order to solve the above technical problems, the embodiments of the present disclosure provide a data processing method, an apparatus, a device and a storage medium.

[0004] In a first aspect, the present disclosure provides a data processing method, the method comprising:

[0005] Receiving a write request from a user for target data, and writing the target data into a corresponding storage space;

[0006] determining junk data associated with the target data;

[0007] The garbage data is cleared and the write status of the target data is returned.

[0008] In an optional implementation manner, the determining junk data related to the target data includes:

[0009] Based on the metadata segments containing the target data and the metadata segments contained by the target data, a first metadata segment set is constructed, and unwritten metadata segments are removed from the first metadata segment set to obtain a second metadata segment set;

[0010] In the second metadata segment set, searching for a metadata segment other than the target data that is closest to the current time and contains the target data as the target metadata segment;

[0011] The metadata segments that are earlier in time than the target metadata segment and included in the target metadata segment are used as a third metadata segment set;

[0012] The data at the physical storage locations respectively corresponding to the metadata segments in the third metadata segment set are used as junk data related to the target data.

[0013] In an optional implementation manner, before treating the data at the physical storage locations respectively corresponding to the metadata segments in the third metadata segment set as garbage data related to the target data, the method further includes:

[0014] Based on the correspondence between the logical addresses and physical storage locations respectively corresponding to the metadata segments in the third metadata segment set, the physical storage locations respectively corresponding to the metadata segments are determined.

[0015] In an optional implementation manner, the method further includes:

[0016] According to the preset rules, the space utilization of the data is queried regularly to determine the current space utilization;

[0017] Based on a preset correspondence between space utilization and garbage collection time, determining the garbage collection time corresponding to the current space utilization;

[0018] Garbage data is recycled based on the garbage collection time.

[0019] In an optional implementation manner, before determining the garbage collection time corresponding to the current space utilization based on the correspondence between the preset space utilization and the garbage collection time, the method further includes:

[0020] The space utilization is divided into multiple levels;

[0021] The corresponding garbage task collection time is configured for each level to establish a corresponding relationship between the preset space utilization and the garbage task collection time.

[0022] In an optional implementation, the space utilization is positively correlated with the garbage task recovery time.

[0023] In an optional implementation, the data processing method is applied to a distributed storage system.

[0024] In a second aspect, the present disclosure provides a data processing device, the device comprising:

[0025] A receiving module, used to receive a user's write request for target data, and write the target data into a corresponding storage space;

[0026] A first determining module, used to determine junk data related to the target data;

[0027] The clearing module is used to clear the garbage data and return the writing status of the target data.

[0028] In a third aspect, the present disclosure provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, and when the instructions are executed on a terminal device, the terminal device implements the above method.

[0029] In a fourth aspect, the present disclosure provides a data processing device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above method when executing the computer program.

[0030] In a fifth aspect, the present disclosure provides a computer program product, wherein the computer program product comprises a computer program / instructions, and the computer program / instructions implement the above method when executed by a processor.

[0031] Compared with the prior art, the technical solution provided by the embodiments of the present disclosure has at least the following advantages:

[0032] The disclosed embodiment provides a data processing method, which receives a user's write request for target data, writes the target data into a corresponding storage space, determines garbage data related to the target data, clears the garbage data, and returns the write status of the target data. By adopting the above technical solution, the data processing device deletes the garbage data related to the target data during the user's writing process for the target data, and can release the data space in time, thereby reducing the execution time of the background garbage collection task, increasing the number of input and output operations per second in the foreground, and further improving the garbage data recovery efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0034] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0035] Figure 1 A flowchart of a data processing method provided by an embodiment of the present disclosure;

[0036] Figure 2 A schematic diagram of data distribution provided in an embodiment of the present disclosure;

[0037] Figure 3 A schematic diagram of another data distribution provided in an embodiment of the present disclosure;

[0038] Figure 4 A schematic diagram of another data distribution provided for an embodiment of the present disclosure;

[0039] Figure 5 A schematic diagram of another data distribution provided in an embodiment of the present disclosure;

[0040] Figure 6 A flowchart of another data processing method provided by an embodiment of the present disclosure;

[0041] Figure 7 A schematic diagram of the structure of a data processing device provided in an embodiment of the present disclosure;

[0042] Figure 8 A schematic diagram of the structure of a data processing device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0043] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.

[0044] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.

[0045] In the current storage field, in order to improve the data overwrite writing speed, the storage platform adopts the ROW method for data, and at the same time starts a background task to regularly clear useless data and recycle unused data space. However, in the case of limited CPU resources, if the background garbage collection is executed too frequently, it will affect the IOPS of the foreground IO; if the background garbage collection execution time is too short, the writing speed exceeds the background garbage data recovery speed, and as time goes by, garbage data will continue to accumulate. When the storage space is full, the user's valid data is relatively low in the entire disk space, resulting in the foreground IO falling to 0. Therefore, this method not only affects the performance of the foreground IO, but also leads to low efficiency in the recovery of garbage data.

[0046] To this end, the disclosed embodiment provides a data processing method, which receives a user's write request for target data, writes the target data into a corresponding storage space, determines garbage data related to the target data, clears the garbage data, and returns the write status of the target data. With the above technical solution, the data processing device deletes the garbage data related to the target data during the user's writing process for the target data, and can release the data space in time, thereby reducing the execution time of the background garbage collection task, increasing the number of input and output operations per second in the foreground, and further improving the garbage data recovery efficiency.

[0047] Based on this, the present disclosure provides a data processing method, referring to Figure 1 , is a flow chart of a data processing method provided by an embodiment of the present disclosure, the method comprising:

[0048] S101: receiving a write request from a user for target data, and writing the target data into a corresponding storage space.

[0049] The data processing method provided in the embodiments of the present disclosure can be applied to a distributed storage system. A distributed storage system refers to a system that stores data in a dispersed manner on multiple nodes or devices, interconnects and works together through a network to provide data storage and access services.

[0050] Target data refers to any data that the user needs to write. A write request refers to an operation instruction issued by a client to a distributed storage system to store specific data in a specified location according to specified requirements. Specifically, a write request can be a request to write data incrementally. Storage space refers to the physical space used to store data.

[0051] In the disclosed embodiment, after receiving a write request for target data from a user, the distributed storage system parses the write request, determines the target data and the storage location, and writes the target data into the storage space indicated by the storage location.

[0052] S102: Determine junk data related to target data.

[0053] The garbage data refers to useless data. The garbage data related to the target data refers to the data formed by the original existing data losing its validity or being replaced by the newly written target data due to the writing of the target data.

[0054] In the disclosed embodiment, junk data related to target data can be determined by checking metadata segments in a metadata tree in a storage system, and junk data related to target data can also be determined by comparing actually stored data with currently valid target data.

[0055] Exemplarily, the above step S102 may include step S201, step S202, step S203 and step S204, which specifically include:

[0056] S201: construct a first metadata segment set based on the metadata segments containing the target data and the metadata segments contained in the target data, remove the unwritten metadata segments from the first metadata segment set, and obtain a second metadata segment set.

[0057] A metadata segment refers to a metadata block in a distributed storage system, which is responsible for describing and managing data in a specific logical address range. Specifically, the metadata segment is used to record the logical address, physical storage location, size, timestamp, and other related attributes of the data. A metadata segment containing target data refers to a metadata segment whose logical address range is greater than or equal to the logical address range of the target data. A metadata segment contained by the target data refers to a metadata segment whose logical address is less than or equal to the logical address range of the target data.

[0058] The first metadata segment set refers to a set consisting of metadata segments whose logical address range is greater than or equal to the logical address range of the target data, and metadata segments whose logical address is less than or equal to the logical address range of the target data. An unwritten metadata segment refers to a metadata segment that has not been completely written to the distributed storage device, is in an incomplete or temporary state, and cannot be used for normal data access or management. The second metadata segment set refers to a set consisting of the remaining metadata segments after removing the unwritten metadata segments from the first metadata segment set.

[0059] In the disclosed embodiment, after the target data is written into the corresponding storage space, a first metadata segment set is constructed from the metadata tree based on metadata segments whose logical address range is greater than or equal to the logical address range of the target data, and metadata segments whose logical addresses are less than or equal to the logical address range of the target data. Then, the unwritten metadata segments are removed from the first metadata segment set, and the set consisting of the remaining metadata segments is used as the second metadata segment set.

[0060] For example, Figure 2 A schematic diagram of data distribution provided in an embodiment of the present disclosure shows the data distribution when the target data is written. Figure 2 As shown, the horizontal axis represents the position, and the vertical axis represents the time. The figure shows an unwritten metadata segment 202, a metadata segment 201 of target data, and completed metadata segments 203-209.

[0061] First, from the metadata tree, a first metadata segment set is constructed based on metadata segments whose logical address range is greater than or equal to the logical address range of the target data (metadata segment 201, metadata segment 203, metadata segment 207), and metadata segments whose logical addresses are less than or equal to the logical address range of the target data (i.e., metadata segment 201, metadata segment 202, metadata segment 205, metadata segment 206, metadata segment 208). Figure 3 As shown, Figure 3 This is a schematic diagram of another data distribution provided by an embodiment of the present disclosure. In the figure, only metadata segment 201, metadata segment 202, metadata segment 203, metadata segment 205, metadata segment 206, metadata segment 207, and metadata segment 208 (i.e., the first metadata segment set) remain. Then, the unwritten metadata segment 202 is removed from the first metadata segment set, and the set consisting of the remaining metadata segments (i.e., data segment 201, metadata segment 203, metadata segment 205, metadata segment 206, metadata segment 207, and metadata segment 208) is used as the second metadata segment set. Figure 4 As shown, Figure 4 A schematic diagram of another data distribution provided for an embodiment of the present disclosure, in which only data segment 201, metadata segment 203, metadata segment 205, metadata segment 206, metadata segment 207, and metadata segment 208 (i.e., a second metadata segment set) remain.

[0062] S202: Searching, in the second metadata segment set, for a metadata segment other than the target data, which is closest to the current time and contains the target data, as the target metadata segment.

[0063] The target metadata segment refers to the metadata segment other than the target data, which is closest to the current time and has a logical address range greater than or equal to the target data range. Figure 4 As shown, metadata segment 203 is the target metadata segment.

[0064] Because in a distributed storage system, the writing operation of the target data can only be considered completed when the target data on each node is successfully written. In the case where the writing status has not yet been determined, in order to avoid data inconsistency or errors caused by reading an older version of the data, it is necessary to retain the metadata segment that is closest to the current time and has a logical address range greater than or equal to the target data range in addition to the target data, so as to ensure that access to the data is accurate and consistent. To this end, it is necessary to search in the second metadata segment set for the metadata segment that is closest to the current time and has a logical address range greater than or equal to the target data range in addition to the target data, as the target metadata segment, so as to facilitate the subsequent clearing of the data at the physical storage location corresponding to the metadata segment that is earlier than the target metadata segment and is included in the target metadata segment.

[0065] S203: The metadata segments that are earlier in time than the target metadata segment and included in the target metadata segment are taken as the third metadata segment set.

[0066] The third metadata segment set refers to a set of metadata segments whose time is earlier than the target metadata segment and whose logical address range is less than or equal to the target metadata segment, that is, Figure 4 The set composed of metadata segment 205, metadata segment 206, and metadata segment 208 shown.

[0067] In the disclosed embodiment, a set consisting of metadata segments whose time is earlier than the target metadata segment and whose logical address range is less than or equal to the target metadata segment is constructed as a third metadata segment set.

[0068] S204: Use the data at the physical storage locations corresponding to the metadata segments in the third metadata segment set as junk data related to the target data.

[0069] Among them, the physical storage location refers to the specific storage address or location of data on the physical storage medium (such as a hard disk, solid-state drive), which describes the actual storage location of the data on the underlying hardware.

[0070] In the disclosed embodiment, each metadata segment in the third metadata segment set corresponds to a specific physical storage location, and the data stored in the specific physical storage location is treated as junk data related to the target data.

[0071] In an optional implementation manner, before treating the data at the physical storage locations respectively corresponding to the metadata segments in the third metadata segment set as garbage data related to the target data, the method further includes:

[0072] Based on the correspondence between the logical addresses and physical storage locations respectively corresponding to the metadata segments in the third metadata segment set, the physical storage locations respectively corresponding to the metadata segments are determined.

[0073] In the embodiment of the present disclosure, each metadata segment in the third metadata segment set describes the correspondence between the logical address and the physical storage location. Based on this relationship, the physical storage location specifically corresponding to each metadata segment can be determined, that is, the actual location of each logical address on the physical storage medium can be found.

[0074] It should be noted that the data processing method provided by the embodiment of the present disclosure can also be applied to a centralized storage system or a stand-alone storage system. In the centralized storage system or the stand-alone storage system, the write operation of the target data is completed, which means that the write status of the target data is successful. Correspondingly, the junk data related to the target data is the metadata segment that is earlier than the metadata segment corresponding to the target data and is included in the metadata segment corresponding to the target data.

[0075] S103: clear the garbage data and return to the write status of the target data.

[0076] The write status refers to the result information after the data write operation is completed, including success, failure or other feedback conditions, which is used to reflect the execution status of the write process.

[0077] In the embodiment of the present disclosure, after determining the junk data related to the target data, the junk data is cleared and the write status of the target data is returned to the user. It can be seen that in the process of writing the target data by the user, the embodiment of the present disclosure recycles certain junk data, further reducing the space occupied by the junk data.

[0078] like Figure 5 As shown, Figure 5 A schematic diagram of another data distribution provided in an embodiment of the present disclosure is shown. Figure 5 The metadata segments 201, 202, 203, 204, 207 and 209 shown in FIG. 1 are the results after the junk data is cleared.

[0079] In the data processing method provided by the embodiment of the present disclosure, a user's write request for target data is received, the target data is written into the corresponding storage space, the garbage data related to the target data is determined, the garbage data is cleared, and the write status of the target data is returned. With the above technical solution, the data processing device deletes the garbage data related to the target data during the user's writing process for the target data, and can release the data space in time, thereby reducing the execution time of the background garbage collection task, increasing the number of input and output operations per second in the foreground, and further improving the garbage data recovery efficiency.

[0080] In some embodiments, the data processing method may also include: regularly querying the space utilization of the data according to preset rules, determining the current space utilization, determining the garbage collection time corresponding to the current space utilization based on the correspondence between the preset space utilization and the garbage task collection time, and recycling garbage data based on the garbage collection time.

[0081] Among them, the preset rules refer to pre-set operating conditions, which are used to guide the distributed storage system to perform task operations in a specified manner or frequency. Space utilization refers to the percentage of used space in the data storage space to the total space. Garbage task recovery time refers to the time spent by the distributed storage system in the process of recovering garbage data. The correspondence between the preset space utilization and the garbage task recovery time refers to a rule or mapping set by the distributed storage system based on the relationship between the usage of the storage space (i.e., space utilization) and the time required for garbage recovery (i.e., garbage recovery time). This relationship can usually be set in advance through empirical data to guide the distributed storage system on how to reasonably arrange garbage data recovery operations under different space utilization conditions.

[0082] In the disclosed embodiment, the distributed storage system queries the space utilization of the data according to preset rules, such as at 14:00 every day, to determine the current space utilization, and then determines the garbage collection time corresponding to the current space utilization based on the correspondence between the preset space utilization and the garbage task collection time, so as to recycle garbage data according to the garbage task time, thereby ensuring effective management and optimized utilization of the storage space.

[0083] In an optional implementation, based on the preset correspondence between the space utilization and the garbage task recovery time, before determining the garbage recovery time corresponding to the current space utilization, the method further includes: establishing the preset correspondence between the space utilization and the garbage task recovery time. Specifically, the space utilization is divided into levels to obtain multiple levels, and the corresponding garbage task recovery time is configured for each level to establish the preset correspondence between the space utilization and the garbage task recovery time.

[0084] The disclosed embodiments can divide the space utilization into several different levels according to the usage of the storage space, such as defining 0%-50% as level 1, 50%-70% as level 2, 70%-90% as level 3, and 90%-100% as level 4, and then configure the corresponding garbage task recovery time for each level to establish a corresponding relationship between the preset space utilization and the garbage task recovery time.

[0085] Among them, space utilization is positively correlated with garbage task recovery time.

[0086] Specifically, the positive correlation between space utilization and garbage task recovery time means that when space utilization increases, the garbage task recovery time will increase accordingly, and vice versa, when space utilization decreases, the garbage task recovery time will decrease. That is, there is a trend of synchronous change between space utilization and garbage task recovery time. For example, when the space utilization of the storage space is high, the distributed storage system needs more garbage task recovery time to clean up garbage data to free up additional storage space, thereby ensuring normal space management and data storage.

[0087] Exemplarily, Table 1 is a table showing the correspondence between space utilization and garbage task recovery time.

[0088] Space Utilization level Garbage collection time 0%~50% Level 1 x 50%~70% Level 2 x+step 70%~90% Level 3 x+2*step 90%~100% Level 4 x+3*step

[0089] Table 1

[0090] Among them, x represents the garbage collection time corresponding to the space utilization level 1, and step represents the step size. Specifically, x and step can be set according to experience. Specifically, by dividing the space utilization level, the corresponding garbage collection time is set for different levels. As the space utilization increases, the garbage collection time gradually increases, and the foreground IOPS will gradually decrease, but it will not suddenly drop to 0.

[0091] It can be seen that the embodiment of the present disclosure dynamically adjusts the execution time of the background garbage collection task (ie, the garbage task collection time) through space utilization, thereby accelerating the garbage data collection rate without having a significant impact on the foreground IOPS.

[0092] Based on the above embodiments, the present disclosure also provides a data processing method, such as Figure 6 As shown, Figure 6 A flowchart of another data processing method provided by an embodiment of the present disclosure, such as Figure 6 As shown, a user's write request for target data is received, the target data is written into the corresponding storage space, the metadata segment containing the target data and the metadata segment contained by the target data are searched, and it is determined whether the result is empty. If the result is not empty, the metadata segment that has not been written is removed from the above-screened metadata segments, and it is determined whether the result is empty. If the result is not empty, the metadata segment closest to the current time and containing the target data is searched as the target metadata segment, and then the data of the physical storage locations corresponding to the metadata segments that are earlier than the target metadata segment and contained by the target metadata segment are deleted. It can be seen that in the process of the user writing the target data, the embodiment of the present disclosure recycles certain garbage data, further reduces the space occupied by the garbage data, thereby reducing the execution time of the background garbage collection task, increasing the number of input and output operations per second in the foreground, and further improving the garbage data recovery efficiency.

[0093] Based on the above method embodiment, the present disclosure also provides a data processing device, such as Figure 7 As shown, Figure 7 A schematic diagram of a data processing device provided in an embodiment of the present disclosure, the device comprising:

[0094] The receiving module 701 is used to receive a write request from a user for target data and write the target data into a corresponding storage space;

[0095] A first determination module 702, configured to determine junk data related to the target data;

[0096] The clearing module 703 is used to clear the garbage data and return the write status of the target data.

[0097] In an optional implementation manner, the first determining module 702 includes:

[0098] A first determining submodule is used to construct a first metadata segment set based on the metadata segment containing the target data and the metadata segment contained by the target data, and remove the unwritten metadata segment from the first metadata segment set to obtain a second metadata segment set;

[0099] A second determination submodule is used to search, in the second metadata segment set, for a metadata segment which is closest to the current time and contains the target data, except for the target data, as a target metadata segment;

[0100] A third determining submodule, configured to take metadata segments that are earlier in time than the target metadata segment and included in the target metadata segment as a third metadata segment set;

[0101] The fourth determining submodule is used to use the data of the physical storage locations respectively corresponding to the metadata segments in the third metadata segment set as junk data related to the target data.

[0102] In an optional implementation, the device further includes:

[0103] The second determination module is used to determine the physical storage locations corresponding to each metadata segment in the third metadata segment set based on the correspondence between the logical addresses corresponding to each metadata segment and the physical storage locations.

[0104] In an optional implementation, the device further includes:

[0105] A query module is used to query the space utilization of data regularly according to preset rules to determine the current space utilization;

[0106] A third determination module is used to determine the garbage collection time corresponding to the current space utilization based on a preset correspondence between the space utilization and the garbage collection time;

[0107] The recycling module is used to recycle garbage data based on the garbage collection time.

[0108] In an optional implementation, the device further includes:

[0109] A level division module is used to divide the space utilization rate into multiple levels;

[0110] The configuration module is used to configure the corresponding garbage task recovery time for each level to establish a corresponding relationship between the preset space utilization and the garbage task recovery time.

[0111] In an optional implementation, the space utilization is positively correlated with the garbage task recovery time.

[0112] In an optional implementation, the data processing method is applied to a distributed storage system.

[0113] In the data processing device provided by the embodiment of the present disclosure, a user's write request for target data is received, the target data is written into the corresponding storage space, the garbage data related to the target data is determined, the garbage data is cleared, and the write status of the target data is returned. With the above technical solution, the data processing device deletes the garbage data related to the target data during the user's writing process for the target data, and the data space can be released in time, thereby reducing the execution time of the background garbage collection task, increasing the number of input and output operations per second in the foreground, and further improving the garbage data recovery efficiency.

[0114] In addition to the above-mentioned method and apparatus, the embodiments of the present disclosure further provide a computer-readable storage medium, in which instructions are stored. When the instructions are executed on a terminal device, the terminal device implements the data processing method described in the embodiments of the present disclosure.

[0115] The embodiments of the present disclosure further provide a computer program product, which includes a computer program / instructions. When the computer program / instructions are executed by a processor, the data processing method described in the embodiments of the present disclosure is implemented.

[0116] In addition, the present disclosure also provides a data processing device, see Figure 8 As shown, this may include:

[0117] Processor 801, memory 802, input device 803 and output device 804. The number of processors 801 in the data processing device can be one or more. Figure 8 In some embodiments of the present disclosure, the processor 801, the memory 802, the input device 803 and the output device 804 may be connected via a bus or other means, wherein: Figure 8 The example of connecting through bus is taken in the following.

[0118] The memory 802 can be used to store software programs and modules. The processor 801 executes various functional applications and data processing of the data processing device by running the software programs and modules stored in the memory 802. The memory 802 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc. In addition, the memory 802 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices. The input device 803 can be used to receive input digital or character information, and generate signal input related to user settings and function control of the data processing device.

[0119] Specifically in this embodiment, the processor 801 will load the executable files corresponding to the processes of one or more applications into the memory 802 according to the following instructions, and the processor 801 will run the applications stored in the memory 802, thereby realizing various functions of the above-mentioned data processing device.

[0120] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0121] The above description is only a specific embodiment of the present disclosure, so that those skilled in the art can understand or implement the present disclosure. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to the embodiments described herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A data processing method, characterized in that: The method comprises: Receiving a write request from a user for target data, and writing the target data into a corresponding storage space; determining junk data associated with the target data; The garbage data is cleared and the write status of the target data is returned.

2. The method according to claim 1, characterized in that The determining of junk data related to the target data includes: Based on the metadata segments containing the target data and the metadata segments contained by the target data, a first metadata segment set is constructed, and unwritten metadata segments are removed from the first metadata segment set to obtain a second metadata segment set; In the second metadata segment set, searching for a metadata segment other than the target data that is closest to the current time and contains the target data as the target metadata segment; The metadata segments that are earlier in time than the target metadata segment and included in the target metadata segment are used as a third metadata segment set; The data at the physical storage locations respectively corresponding to the metadata segments in the third metadata segment set are used as junk data related to the target data.

3. The method according to claim 2, characterized in that Before treating the data at the physical storage locations respectively corresponding to the metadata segments in the third metadata segment set as garbage data related to the target data, the method further includes: Based on the correspondence between the logical addresses and physical storage locations respectively corresponding to the metadata segments in the third metadata segment set, the physical storage locations respectively corresponding to the metadata segments are determined.

4. The method according to claim 1, characterized in that: The method further comprises: Regularly query the space utilization of data according to preset rules to determine the current space utilization; Based on a preset correspondence between space utilization and garbage collection time, determining the garbage collection time corresponding to the current space utilization; Garbage data is recycled based on the garbage collection time.

5. The method according to claim 4, characterized in that Before determining the garbage collection time corresponding to the current space utilization rate based on the corresponding relationship between the preset space utilization rate and the garbage collection time, the method further includes: The space utilization is divided into multiple levels; The corresponding garbage task collection time is configured for each level to establish a corresponding relationship between the preset space utilization and the garbage task collection time.

6. The method according to claim 5, characterized in that The space utilization rate is positively correlated with the garbage task recovery time.

7. The method according to claim 1, characterized in that The data processing method is applied to a distributed storage system.

8. A data processing device, characterized in that: The device comprises: A receiving module, used to receive a user's write request for target data, and write the target data into a corresponding storage space; A first determining module, used to determine junk data related to the target data; The clearing module is used to clear the garbage data and return the writing status of the target data.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed on a terminal device, the terminal device implements the method according to any one of claims 1 to 7.

10. A data processing device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.