Data reduction processing method and electronic device

CN122837754APending Publication Date: 2026-09-29INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611339531.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-31
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

相关技术中,数据缩减方法大多基于数据块的相似度得到保留数据块和差分补丁,保留的原始数据块、引用索引、差分补丁等统一采用固定策略进行存储保护,然而这种方式存在基础数据块误删的风险

Benefits of technology

[0014]本申请还提供了一种计算机程序产品,包括计算机程序,计算机程序被处理器执行时实现上述任一种数据的缩减处理方法的步骤。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122837754A_ABST
    Figure CN122837754A_ABST
Patent Text Reader

Abstract

This application discloses a data reduction processing method and electronic device, relating to the field of data processing technology, comprising: acquiring a data block set; wherein the data block set includes multiple data blocks; determining the dependencies between data blocks to obtain the dependencies between data blocks; determining at least one retained data block and at least one data block to be reduced from the data block set according to the dependencies; wherein the data block to be reduced can be recovered based on the retained data block; verifying the recovery of the data block to be reduced, and generating reconstruction information for recovering the data block to be reduced after successful verification; calculating the reduction risk parameter corresponding to the data block to be reduced; determining the corresponding redundancy protection strength according to the reduction risk parameter; and storing the retained data block and reconstruction information according to the redundancy protection strength, which can reduce the risk of accidental deletion of basic data blocks and improve the reliability of data storage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to data reduction processing methods and electronic devices. Background Technology

[0002] As the scale of database backups, log archiving, file synchronization, object storage, and distributed disaster recovery storage continues to increase, storage systems typically require data reduction. Specifically, this involves reducing the physical write volume of duplicate, similar, and incremental data while ensuring data recoverability. Most data reduction methods rely on the similarity of data blocks to determine the retained data blocks and differential patches. The retained original data blocks, reference indexes, and differential patches are protected using a fixed strategy. However, this approach carries the risk of accidental deletion of underlying data blocks. Therefore, achieving accurate data reduction and improving data storage reliability are pressing technical challenges. Summary of the Invention

[0003] This application provides a data reduction processing method and electronic device to at least solve the problems of how to achieve accurate data reduction and improve the reliability of data storage in related technologies.

[0004] This application provides a data reduction processing method, including:

[0005] Retrieve a set of data blocks; where the set of data blocks includes multiple data blocks.

[0006] Dependency determination is performed on data blocks to determine the dependencies between them;

[0007] Based on dependencies, at least one reserved data block and at least one data block to be reduced are determined from the data block set; wherein the data block to be reduced can be recovered based on the reserved data block;

[0008] Perform recovery verification on the data block to be reduced, and generate reconstruction information for restoring the data block to be reduced after the verification is successful;

[0009] Calculate the reduction risk parameters corresponding to the data block to be reduced;

[0010] Determine the corresponding redundancy protection strength based on the reduced risk parameters;

[0011] Based on the redundancy protection strength, store the reserved data blocks and reconstruction information.

[0012] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for implementing the steps of any of the above-described data reduction processing methods when executing the computer program.

[0013] This application also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the steps of any of the above-described data reduction processing methods.

[0014] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described data reduction processing methods.

[0015] This application provides an accurate data reduction processing scheme. For multiple data blocks to be reduced, it first determines the dependencies between them, and then rationally determines the data blocks to be retained and the data blocks to be reduced based on these dependencies. Compared to methods that directly reduce data based on similarity, this scheme achieves more accurate identification of the data blocks to be reduced. Furthermore, by restoring and verifying the data blocks to be reduced, it ensures that the data blocks to be reduced can indeed be losslessly restored from the retained data blocks before generating reconstruction information, avoiding unreliable reduction operations. Different redundancy protection strengths are configured for different data blocks to be reduced based on their reduction risk coefficients, achieving differentiated storage protection. The accurate identification and restoration verification of the data blocks to be reduced reduces the risk of accidental deletion of basic data blocks, and the adaptive storage protection corresponding to the reduction risk improves the reliability of the stored data after reduction. Attached Figure Description

[0016] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 A schematic diagram of a data reduction processing system architecture provided in this application embodiment;

[0018] Figure 2 Flowchart of the data reduction processing method provided in the embodiments of this application Figure 1 ;

[0019] Figure 3 A flowchart illustrating a dependency determination method provided in an embodiment of this application;

[0020] Figure 4 Flowchart of the data reduction processing method provided in the embodiments of this application Figure 2 ;

[0021] Figure 5 A schematic diagram of a lossless data block reduction and collaborative storage system provided in this application embodiment;

[0022] Figure 6 A schematic diagram of the data reduction processing apparatus provided in the embodiments of this application;

[0023] Figure 7 A schematic diagram of the structure of the electronic device provided in this application. Detailed Implementation

[0024] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the embodiments of this application.

[0025] It should be noted that, in the description of the embodiments of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in the embodiments of this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0026] First, the terms used in the embodiments of this application will be explained:

[0027] Erasure coding: a channel coding method for forward error correction that achieves error correction by generating redundant data through polynomial overdetermined sampling.

[0028] Hash verification: By calculating the hash value of a file, it is used to verify data integrity and confirm whether the file has been tampered with during transmission or storage.

[0029] Data reduction schemes typically involve first dividing the input file or data stream into multiple data blocks, then calculating hash values, fingerprint values, or checksums for each block. Through methods such as complete deduplication, differential compression, erasure coding protection, and encrypted storage, the actual storage space is reduced while improving data security. A typical reduction scheme includes a data block splitting module, a hash deduplication module, a similar block matching module, a differential patch generation module, a reference index table, an encryption module, and an erasure coding module. The basic processing method is as follows: when two data blocks have the same hash value, only one data block is retained, and a reference index is created for the other data block; when two data blocks have a high degree of similarity, one data block is used as the base block to generate a differential patch for the other data block; then, the retained data block, reference index, and differential patch are written to the storage medium and protected using erasure coding or multiple copies. However, the above techniques still have significant structural flaws. First, existing solutions often only focus on whether data blocks are duplicated or similar, lacking a judgment on whether the underlying data block should be prioritized for retention. This can lead to a basic data block that is depended on by multiple differential blocks being misjudged as a reducible object, causing subsequent recovery chain breaks. Second, the reference indexes or differential indexes in these solutions typically only record the correspondence between the reduced block and the underlying block, lacking dependency graph loop prevention checks. Multiple reduced data blocks can easily form indirect circular dependencies or multi-level recovery chains, increasing the risk of recovery failure and delay. Third, the erasure coding protection in these solutions often uses fixed redundancy strength, failing to dynamically adjust the protection strength based on the proportion of reduced data, the length of the reconstructed witness index, and the degree of concentrated reference to the underlying block. This means that in partitions with a high degree of reduction, if the reconstructed index or critical underlying block is damaged, it will significantly affect the recovery reliability of a large number of reduced data blocks. Therefore, how to achieve accurate data reduction and improve the reliability of data storage is an urgent technical problem to be solved.

[0030] To address the aforementioned technical problems, embodiments of this application provide a data reduction processing method and electronic device. For multiple data blocks to be reduced, the method first determines the dependencies between the multiple data blocks, reasonably determines the data blocks to be retained and the data blocks to be reduced based on the dependencies, performs recovery verification on the data blocks to be reduced, and after the verification is passed, configures different redundancy protection strengths for different data blocks to be reduced based on their reduction risk coefficients to achieve differentiated storage protection.

[0031] Alternatively, a data block reduction and collaborative storage method is needed that can simultaneously control the retention of base blocks, dependencies, and fault tolerance strength during lossless reduction.

[0032] The specific application environment architecture or hardware architecture upon which the data reduction processing method depends is described here. (References) Figure 1 , Figure 1 This is a schematic diagram of a data reduction processing system architecture provided in an embodiment of this application. The data reduction processing system is a computer device. Figure 1 As shown, the above architecture includes at least one of a data acquisition device 101, a processing device 102, and a display device 103.

[0033] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the architecture of the data reduction processing system. In other feasible implementations of the embodiments of this application, the above architecture may include more or fewer components than illustrated, or combine some components, or split some components, or arrange different components, which can be determined according to the actual application scenario and is not limited here. Figure 1 The components shown can be implemented in hardware, software, or a combination of both.

[0034] In the specific implementation process, the data acquisition device 101 may include an input / output interface or a communication interface, and the data acquisition device 101 can be connected to the processing device through the input / output interface or the communication interface.

[0035] The processing device 102 can determine the dependency relationship between multiple data blocks, reasonably determine the data blocks to be retained and the data blocks to be reduced based on the dependency relationship, perform recovery verification on the data blocks to be reduced, and after the verification is passed, configure different redundancy protection strengths for different data blocks to be reduced based on their reduction risk coefficients to achieve differentiated storage protection.

[0036] The display device 103 can also be a touch screen or the screen of a terminal device, used to receive user commands while displaying the above-mentioned content, so as to realize interaction with the user.

[0037] It should be understood that the aforementioned processing device can be implemented by a processor reading instructions from memory and executing those instructions, or it can be implemented by a chip circuit.

[0038] Furthermore, the network architecture and business scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of network architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0039] Figure 2 Flowchart of the data reduction processing method provided in the embodiments of this application Figure 1 ,like Figure 2 As shown, an embodiment of this application provides a data reduction processing method, the execution subject of which can be the aforementioned processing device 102. The method is described in detail below:

[0040] S201: Get the data block set.

[0041] The data block set includes multiple data blocks.

[0042] Here, the data block set is used to carry multiple data blocks and serves as the input basis for subsequent dependency determination, retention decisions, recovery verification, risk calculation, and redundancy protection configuration. The data block is the basic processing unit in this method; each data block corresponds to a segment of independently processable data content and participates in dependency identification, retention decisions, and reduction processing.

[0043] Optionally, the data to be processed can be received and organized into multiple data blocks to form a data block set.

[0044] Optionally, after obtaining the data block set, relevant information for determining dependencies and performing recovery processing can be obtained based on subsequent processing needs.

[0045] In one possible implementation, an input data stream is received and data block metadata is generated. Specifically, the input data stream is received and divided into multiple data blocks according to a preset segmentation rule; a data block record is generated for each data block.

[0046] Optionally, the data block record includes at least one of the following: data block identifier, file identifier or data stream identifier, data block offset address, original data block length, content hash value, fast checksum, rolling fingerprint set, data block version number, and current status identifier.

[0047] Among them, the content hash value is used to determine whether the data block is completely duplicated, the fast check code is used to assist in verifying the data block recovery result, and the rolling fingerprint set is used to recall candidate basic data blocks.

[0048] S202: Determine the dependencies between data blocks to obtain the dependencies between them.

[0049] The dependency determination function is used to identify whether there are any relationships between data blocks that can be used for recovery, and outputs the inter-block relationship results that can be used for retention decisions.

[0050] Optionally, the dependency relationship includes the base data block as the dependent party and the data block to be reduced as the dependent party.

[0051] Optionally, data blocks and their related information are read one by one from the data block set, and dependencies are determined based on the recoverable associations between data blocks.

[0052] Optionally, for data blocks with recovery relationships, the correspondence between related data blocks can be recorded in the dependency results. If necessary, consistency or recoverability checks can be performed on the determined dependencies to prevent subsequent reduction processing from being based on unrecoverable relationships. Through the above processing, the resulting dependencies between data blocks not only reflect the relationships between data blocks but also directly serve subsequent retention decisions and recovery processing, avoiding unrecoverable situations caused by superficial similarity alone.

[0053] S203: Based on dependencies, determine at least one reserved data block and at least one data block to be reduced from the data block set.

[0054] Among them, the data blocks to be reduced can be recovered based on the retained data blocks.

[0055] Here, the reserved data blocks are data blocks that are retained after the dependency relationship is determined and will not be reduced, serving as the basis for the recovery of other data blocks; the data blocks to be reduced are data blocks that meet the reduction conditions and can be recovered based on the reserved data blocks.

[0056] Optionally, based on dependencies, determine which data blocks within each data block are designated as reserved data blocks and which are designated as data blocks to be reduced. Data blocks with independent content that cannot be recovered from other data blocks can be designated as reserved data blocks. Data blocks that can be recovered based on a specific reserved data block can be designated as data blocks to be reduced.

[0057] Optionally, if a data block corresponds to multiple optional recovery bases, one of them can be selected as a reserved data block for subsequent recovery processing to reduce the complexity of subsequent processing.

[0058] Optionally, to ensure that the data blocks to be reduced can be recovered based on the retained data blocks, each data block to be reduced must be associated with at least one retained data block during the determination process, and the retained data block must be able to support the recovery of the current data block to be reduced. For the data content of the data blocks determined to be reduced, their physical cache space can be released after subsequent verification, or their original physical write task can be canceled, retaining only the necessary reconstruction information.

[0059] It should be noted that, in this embodiment, the basic data block is a concept defined in the dependency determination stage, referring to the dependent data block that serves as the basis for restoring other data blocks; while the reserved data block is a concept defined in the storage decision stage, referring to the data block whose original data needs to be actually stored. All basic data blocks belong to the reserved data block category because they must be stored to serve as the basis for recovery. However, reserved data blocks are not limited to basic data blocks; data blocks that are not depended on by any other data block and do not meet the reduction conditions also belong to the reserved data block category. The set of basic data blocks is a subset of the set of reserved data blocks.

[0060] S204: Perform recovery verification on the data block to be reduced, and generate reconstruction information for restoring the data block to be reduced after the verification is successful.

[0061] Here, recovery verification is used to check whether the data block to be reduced can be correctly recovered based on the corresponding retained data block, while reconstruction information is the recovery basis retained after the verification is passed.

[0062] Optionally, for each identified data block to be reduced, its corresponding reserved data block and related recovery basis are obtained, and recovery verification is performed to check whether the data block to be reduced can be correctly recovered. If the verification passes, reconstruction information for recovering the data block to be reduced is generated. This reconstruction information can be saved in a structured record format to support subsequent read recovery. If the recovery verification fails, the data block's reduction status is canceled, and it is treated as a reserved data block, ensuring that there are no unrecoverable reduction results in the data block set. This step is crucial for resolving the issues of missing recovery basis and unreliable recovery chain, because reconstruction information is generated and the reduction of the original data block is allowed only if actual recovery is feasible. Therefore, a strict verification closed loop is established between the reduction action and the recovery capability.

[0063] It should be noted that, in the embodiments of this application, reconstruction information and reconstruction witness index refer to the same concept, both referring to information records used to recover the data block to be reduced. Reconstruction information may include one or more of the following: the identifier of the data block to be reduced, the file identifier to which the data block belongs, the offset address of the data block in the original data stream, the identifier of the base data block, the version number of the base data block, the reconstruction type, the differential patch data, the length of the original data block, the hash value of the original data block, the fast checksum of the original data block, the reconstruction algorithm identifier, the reconstruction information version number, the dependency depth, the information integrity check value, and the information generation time.

[0064] S205: Calculate the reduction risk parameters corresponding to the data block to be reduced.

[0065] Among them, the reduction risk parameter is used to characterize the degree of recovery risk faced by the data block to be reduced after reduction, and serves as the basis for subsequently determining the strength of redundancy protection.

[0066] Optionally, risk values ​​can be calculated for a single or multiple data blocks to be reduced based on the reduction information corresponding to those data blocks. The reduction information reflects the dependence of data recovery after reduction on the retained data blocks and reconstructed information, as well as the associated recovery burden. Through the above quantitative processing, the system can unify recovery uncertainties that were originally difficult to compare directly into measurable reduction risk parameters, providing a direct basis for subsequent risk-based matching of protection levels.

[0067] Optionally, the reduction risk parameters corresponding to the data blocks to be reduced are calculated, including: dividing the data block set into multiple processing partitions; and obtaining the reduction information of the data blocks to be reduced within the processing partitions.

[0068] The reduction information includes at least one of the following: the proportion of data already reduced, the proportion of reconstructed information, and the concentration of basic blocks; based on the reduction information, the reduction risk parameters corresponding to the data blocks to be reduced within the processing partition are calculated.

[0069] Optionally, the data block set can be divided into multiple processing partitions based on the data block ownership information, business domain identifier, or storage pool identifier, and the amount of data that has been reduced, the number of reconstructed information, and the distribution of basic data block references can be counted within each processing partition.

[0070] Optionally, the amount of reduced data within the processing partition can be obtained by summing the number of bytes of the data blocks that have been reduced, the reconstruction information can be obtained by statistically analyzing the patch, index, or verification information required to restore the data blocks to be reduced, and the concentration of the basic block can be determined by the ratio of the number of data blocks to be reduced that are referenced by the same basic data block to the total number of data blocks to be reduced within the partition.

[0071] Optionally, the aforementioned reduction information can be input into a preset mapping function, a weighted summation model, or a hierarchical rule to obtain the corresponding reduction risk parameter, and the parameter can be output for differentiated redundancy protection.

[0072] In terms of working principle, partitioning limits risk assessment to a local data block. The reduction information reflects the reduction scale, the additional information load for recovery, and the degree of dependency concentration. These three factors together characterize the recovery vulnerability of the data blocks to be reduced within the partition. The reduction risk parameter can increase with the proportion of reduced data, the proportion of reconstructed information, or the concentration of the base blocks, thus providing a quantitative basis for determining the subsequent redundancy protection strength.

[0073] Optionally, if there is a dependency between a data block and its underlying data block, the data block and the underlying data block will be preferentially assigned to the same processing partition.

[0074] Optionally, if the data block and the underlying data block cannot be assigned to the same processing partition, the underlying data block is set to a solidified reserved block state, or the underlying data block is written and solidified preferentially in the current processing batch.

[0075] Optionally, the risk reduction parameter can be calculated based on the global data block set or on the processing partition.

[0076] In one possible implementation, reduction risk parameters are calculated based on a global data block set: the proportion of reduced data, the proportion of reconstructed information, and the concentration of base blocks are statistically analyzed in the global data block set; the reduction risk parameters are then calculated based on this information. This implementation is suitable for scenarios with small data scales or where partitioning is not required.

[0077] In another possible implementation, reduction risk parameters are calculated based on processing partitions: the data block set is divided into multiple processing partitions; reduction information of the data blocks to be reduced within each processing partition is obtained; and based on the reduction information, the reduction risk parameters corresponding to the data blocks to be reduced within each processing partition are calculated. This implementation is suitable for large-scale data scenarios and can achieve fine-grained risk measurement at the partition level.

[0078] According to the risk reduction parameter calculation method in this application, the data block set is divided into multiple processing partitions and statistics are performed on a partition-by-partition basis. This ensures that the granularity of the risk reduction parameter measurement is appropriate and meaningful, avoiding the problem of risk averaging caused by global statistical algorithms, which fails to accurately identify high-risk areas. By obtaining multi-dimensional reduction information such as the proportion of reduced data volume, the proportion of reconstructed information, and the concentration of basic blocks, the recovery dependency risk after data reduction can be comprehensively quantified. This provides an accurate data foundation for the subsequent dynamic determination of redundancy protection strength, which is conducive to achieving precise matching between risk and protection.

[0079] Optionally, the data block set is partitioned to obtain multiple processing partitions, including: obtaining the ownership information of the data blocks; using the ownership information as a constraint, the data block set is partitioned to obtain multiple processing partitions.

[0080] The attribution information includes at least one of the file identifier, object identifier, or time window identifier to which the data block belongs.

[0081] Optionally, the ownership information of data blocks can be synchronously written to metadata when data blocks are generated or accessed, or it can be read from existing indexes, lists, or storage directories and parsed when entering the reduction processing chain. Based on the parsed file identifier, object identifier, or time window identifier, grouping mapping is performed on the data block set, so that data blocks belonging to the same constraint are aggregated into the corresponding partitions. The number, order, and boundaries of data blocks within a partition can be determined by preset partitioning rules. For data blocks that have multiple ownership information, their partition ownership can be determined according to preset priority, composite key, or intersection relationship, thereby forming a more granular processing partition.

[0082] Here, by acquiring ownership information such as file identifiers, object identifiers, or time window identifiers of data blocks, and using this ownership information as a constraint for partitioning, data blocks with the same ownership information are preferentially assigned to the same processing partition. This data ownership-based partitioning strategy helps maintain data locality and the integrity of dependencies, ensuring that dependencies between data blocks within the same file, object, or time window can be fully identified within the partition. This improves the accuracy of reduction decisions and also facilitates independent fault-tolerant coding and management on a partition-by-partition basis.

[0083] Optionally, calculating the reduction risk parameters corresponding to the data block to be reduced includes: obtaining the data block distribution information of the processing partition corresponding to the data block to be reduced; and calculating the reduction risk parameters corresponding to the data block to be reduced based on the data block distribution information.

[0084] The data block distribution information includes the proportion of reduced data volume, the proportion of reconstructed information, and the concentration of basic blocks. The concentration of basic blocks is used to characterize the degree to which multiple data blocks to be reduced depend on the same basic data block.

[0085] Optionally, the processing partition can consist of a set of data blocks constrained by the same file identifier, object identifier, or time window identifier. First, the data blocks that have been reduced, the data blocks that need to generate reconstruction information, and their corresponding basic data blocks within the processing partition are statistically analyzed. Then, the proportion of reduced data, the proportion of reconstruction information, and the concentration of basic blocks are calculated respectively.

[0086] Optionally, the concentration of the base block can be quantified based on the number of times the same base data block is referenced by multiple data blocks to be reduced, the scope of reference coverage, or the reference density, and can be normalized to a base block concentration value within a preset interval. Subsequently, the distribution information is input into the risk calculation model. The risk calculation model can output the reduction risk parameters using a weighted summation method, a segmented mapping method, or a lookup table method, and the weight of each indicator can be preset according to the reduction scale and recovery requirements of the processing partition.

[0087] As an optional implementation, the base block concentration is calculated as follows: the base block concentration is the quotient of the reference count of the most referenced base data block within the processing partition and the number of reduced data blocks within the processing partition. A higher base block concentration indicates that more data blocks to be reduced rely on the same base data block. If this base data block is damaged, it will affect the recovery of a large number of data blocks, thus increasing the risk of reduction.

[0088] Optionally, after completing the recovery verification of the data block to be reduced and generating reconstruction information, the distribution statistics of the processing partition where it is located are read, and the proportion of the reduced data volume, the proportion of reconstruction information and the concentration of the basic block are used as risk assessment inputs to obtain the reduction risk parameter. This parameter is then passed to the redundancy protection strength determination unit to configure the corresponding storage protection level accordingly.

[0089] In one possible implementation, the processing partitions are generated by dividing the current processing batch (a set of data blocks) into multiple processing partitions. These partitions are used to limit and reduce the processing scope and facilitate subsequent encryption and erasure coding.

[0090] Optionally, when dividing processing partitions, data blocks from the same file, the same object, or the same time window are preferentially assigned to the same processing partition; at the same time, the number of data blocks in a single processing partition does not exceed a preset upper limit, and the total length of the original data in a single processing partition does not exceed a preset upper limit.

[0091] Optionally, if there is a valid dependency between a data block and its candidate underlying data block, the two are preferentially assigned to the same processing partition; if they cannot be assigned to the same processing partition, the candidate underlying data block is placed in a solidified reserved block state, or it is preferentially written and solidified in the current processing batch.

[0092] This application embodiment acquires data block distribution information of the processing partition and comprehensively quantifies the reduction risk from three dimensions: the proportion of reduced data volume, the proportion of reconstructed information, and the concentration of basic blocks. The concentration of basic blocks characterizes the degree to which multiple data blocks to be reduced depend on the same basic data block, thereby accurately identifying dependency hotspots and potential single points of failure. Based on the comprehensive quantification of these three dimensions, the reduction risk parameter can fully reflect the recovery reliability risk caused by changes in dependency relationships after data reduction, providing a precise and quantitative decision-making basis for subsequent adaptive redundancy protection.

[0093] S206: Determine the corresponding redundancy protection strength based on the reduced risk parameters.

[0094] Among them, the redundancy protection strength is used to indicate what level of protection is taken for the reserved data blocks and the reconstructed information, and is further mapped to specific redundancy parameters.

[0095] Optionally, after obtaining the risk reduction parameters, these parameters are matched with a preset risk range to obtain the corresponding protection level. Subsequently, redundancy protection strength parameters are generated based on the protection level. By mapping the risk reduction parameters to graded protection strengths, a correspondence between risk reduction and storage protection is achieved, rather than applying a uniform strength protection to all retained data blocks and reconstruction information.

[0096] Based on the above analysis, it can be seen that in high-risk scaling-up scenarios, the responsibility for recovery of retained data blocks and reconstructed information carriers is greater. Therefore, by increasing the corresponding protection strength, it is possible to recover even when the media is damaged, the node fails, or some data is lost, thereby achieving a controlled balance between scaling-up efficiency and recovery reliability.

[0097] Optionally, based on the redundancy protection strength, the retained data blocks and reconstruction information are stored, including: determining the number of check symbols for erasure coding based on the risk reduction parameter; performing erasure coding on the retained data blocks and reconstruction information based on the number of check symbols to generate check symbols; and writing the retained data blocks, reconstruction information, and check symbols to the storage medium.

[0098] Among them, the number of verification symbols is positively correlated with the risk reduction parameter.

[0099] In the specific implementation, the risk parameter mapping module first receives the reduced risk parameter and determines the number of check symbols based on a preset mapping relationship or calculation formula. When the reduced risk parameter increases, the corresponding number of check symbols increases synchronously to enable subsequent encoding to form a stronger redundancy protection. Subsequently, the encoding module performs erasure encoding on the retained data blocks and reconstruction information according to the determined number of check symbols, generates check symbols consistent with the number, and combines the retained data blocks, reconstruction information, and check symbols into a data unit to be written. The data unit to be written can be distributed and written to multiple storage nodes by the fragmented writing unit, or written to a single storage medium in order of preset block size. In practical applications, the storage medium can also be a disk, solid-state drive, or object storage medium, which is not limited in this application.

[0100] Here, this embodiment establishes a positive correlation between the reduction risk parameter and the number of erasure code verification symbols, thereby realizing an adaptive redundancy protection strategy where higher reduction risk results in stronger fault tolerance protection. This allows limited storage resources to be precisely allocated to high-risk data partitions. By generating verification symbols through erasure coding of retained data blocks and reconstructed information, unified redundancy protection is provided for the reduced data and recovery-dependent information. By writing the encoded information into the storage medium, secure data persistence is achieved, ensuring the reliability and recoverability of the data throughout its entire storage lifecycle.

[0101] In one possible implementation, the reduction risk parameter is calculated based on the reduction results for each processing partition. The reduction risk parameter is determined at least based on the deletion ratio, index ratio, and base block concentration.

[0102] Among them, the deletion ratio represents the ratio between the total original length of the reduced blocks and the total original data length of the processing partition; the index ratio represents the proportion of the total length of the reconstructed witness index in the total length of the retained data and the reconstructed witness index; and the base block concentration represents the degree to which multiple reduced blocks within the processing partition are concentrated on the same base data block.

[0103] Specifically, the risk parameter reduction can be calculated using the following formula:

[0104]

[0105] in, This represents the reduction risk parameter for the i-th processing partition. Indicates the deletion ratio. Indicates the index ratio. Indicates the concentration of basic blocks. , , Let be the weight parameters, and satisfy: .

[0106] As one specific implementation method, if the data importance level is ordinary, then , , If the data importance level is important, then , , If the data importance level is critical, then , , .

[0107] S207: Store reserved data blocks and reconstruction information according to the redundancy protection strength.

[0108] Optionally, the reserved data blocks and reconstruction information are first processed according to the redundancy protection strength, and then written to the storage medium.

[0109] Optionally, after the write operation is completed, update the relevant records for the retained data blocks and reconstruction information that have been successfully persisted, and release the original physical content of the data blocks to be reduced that have passed verification from the cache, or mark them in the disk write plan as no longer to be written, and only retain their logical reference relationship.

[0110] Optionally, if an error occurs during the writing process, adjustments or retrying can be made according to the corresponding redundancy protection strength to meet the corresponding protection requirements.

[0111] Based on the above processing, both the retained data blocks and the reconstruction information are persistently stored in a form that matches their risk. During subsequent recovery, the retained data blocks are read first, and then the reconstruction information is combined to perform reconstruction, thus obtaining the original content of the data blocks to be reduced. This step ultimately implements the aforementioned dependency determination, recovery verification, risk calculation, and protection mapping at the physical storage layer, forming a complete data reduction closed loop.

[0112] The following section introduces a specific differentiated storage protection method to achieve reliable data recovery.

[0113] In one possible implementation, the processing partition includes a core reserved block, a regular reserved block, a reconstruction witness index, a dependency graph digest, and a partition recovery verification digest. The core reserved block, regular reserved block, reconstruction witness index, dependency graph digest, and partition recovery verification digest of the processing partition are encapsulated into a data packet to be protected, and the data packet to be protected is then encrypted.

[0114] After encryption is complete, the number of erasure code check symbols for the processed partition is determined based on the risk reduction parameters:

[0115]

[0116] in, This represents the number of parity symbols in the i-th processing partition. Indicates the minimum number of check symbols. This represents the redundancy strength adjustment coefficient. If the calculated number of check symbols is greater than the preset maximum value, then the preset maximum value is taken. This preset maximum value can be determined based on actual circumstances, and this application embodiment does not impose specific limitations on it. Indicates to The calculation result is rounded up.

[0117] The encrypted data is divided into multiple original symbols, and the corresponding number of verification symbols are generated by erasure coding.

[0118] For the witness index area containing the reconstructed witness index, a protection strength no less than that of the reserved data area should be adopted. As a preferred approach, the number of check symbols in the witness index area should be greater than or equal to the number of check symbols in the corresponding processing partition plus one, or the witness index area should be written to at least two different physical nodes.

[0119] If the witness index area has not completed write verification, the original cache of the corresponding reduced block will not be released, nor will the corresponding data block be marked as a reduced block.

[0120] Next, a physical write operation is performed to solidify the data blocks. The encrypted original symbol and check symbol are then written to different disks, storage nodes, or availability zones. After the write is complete, the write receipt is read and write verification is performed.

[0121] Optionally, for core reserved blocks and ordinary reserved blocks, their status is only updated to solidified reserved blocks after the data packet containing them has been encrypted, erasure encoded, physically written, and written verified.

[0122] Optionally, for a reduced block, its original data block cache is only released when its corresponding underlying data block has become a fixed reserved block, the corresponding reconstructed witness index has completed the witness index area protection, the corresponding processing partition has passed the review, the dependency graph verification has not found any loops, and the physical write verification has passed; otherwise, the corresponding data block is converted into a normal reserved block and written to the physical medium.

[0123] Next, perform data reading and recovery.

[0124] Optionally, when a data read request is received, the data packet corresponding to the processing partition is located according to the partition number and the data block address table; if some physical symbols are lost or damaged, the missing symbols are recovered by erasure coding decoding; then the data packet is decrypted and the authentication tag is verified.

[0125] Optionally, for core reserved blocks and ordinary reserved blocks, their data content is directly output. For reduced blocks, the corresponding reconstruction witness index is read, the integrity of the reconstruction witness index is verified, the base data block is located according to the reconstruction witness index, and a full repeat reconstruction or differential reconstruction is performed according to the reconstruction type to obtain the recovered data block.

[0126] Optionally, when the data length, content hash value, and fast checksum of the recovered data block are consistent with the information recorded in the reconstructed witness index, the recovered data block is output; if the checksum fails, a redundant copy of the witness index area is read or the erasure coding decoding and decryption process of the corresponding processing partition is re-executed.

[0127] This application provides an accurate data reduction processing scheme. For multiple data blocks to be reduced, the dependency relationships between them are first determined. Based on these dependencies, the data blocks to be retained and the data blocks to be reduced are reasonably determined. Compared to data reduction based directly on similarity, this scheme achieves accurate identification of the data blocks to be reduced. Furthermore, by restoring and verifying the data blocks to be reduced, it is ensured that the data blocks to be reduced can indeed be losslessly restored from the retained data blocks before generating reconstruction information, avoiding unreliable reduction operations. Different redundancy protection strengths are configured for different data blocks to be reduced based on their reduction risk coefficients, achieving differentiated storage protection. The accurate identification and restoration verification of the data blocks to be reduced reduces the risk of accidental deletion of basic data blocks, and the adaptive storage protection corresponding to the reduction risk improves the reliability of the stored data after reduction.

[0128] The steps S202 and S203 described above will be explained below with reference to specific embodiments.

[0129] Optionally, Figure 3 This is a flowchart illustrating a dependency determination method provided in an embodiment of this application. Figure 3 As shown, the steps for determining dependencies are as follows:

[0130] S2021: Obtain record information for the data block.

[0131] S2022: Based on the recorded information, perform a full duplication dependency determination on the data blocks to obtain the full duplication dependency relationships between the data blocks.

[0132] Among them, a completely duplicated dependency relationship includes the first base data block as the dependent party and the first data block to be reduced as the dependent party.

[0133] Optionally, the recorded information includes the content hash value, data length, and fast check code; based on the recorded information, a complete duplicate dependency determination is performed on the data blocks to obtain the complete duplicate dependency relationship between the data blocks, including: searching for candidate data blocks from the data block set based on the content hash value of the data block; if the found candidate data block has the same content hash value, the same data length, and the same fast check code as the data block, then it is determined that there is a complete duplicate dependency relationship between the data block and the candidate data block.

[0134] Specifically, when searching for candidate data blocks, the data block itself is excluded; that is, the candidate data block found has a different data block identifier than the data block currently being evaluated. If only the current data block itself can be found, then there are no other available candidate data blocks, and a complete duplicate dependency relationship is not formed.

[0135] Optionally, when multiple candidate data blocks satisfying the condition of complete duplicate dependency exist, the first basic data block is determined according to the following priority order: First, data blocks already in the state of solidified reserved blocks are selected; second, data blocks already in the state of core reserved blocks are selected; third, data blocks with higher reliability levels on their respective physical nodes are selected; fourth, data blocks belonging to the same file or data stream as the current data block are selected; and finally, data blocks with smaller block identifiers are selected. Through this priority selection mechanism, the most stable and reliable data block can be selected from multiple available candidate blocks as the basic data block, thereby improving the reliability of subsequent recovery.

[0136] Among them, the candidate data block is the first basic data block, and the data block is the first data block to be reduced.

[0137] This application embodiment can accurately identify complete duplicate dependencies between data blocks by simultaneously verifying three dimensions: content hash value, data length, and fast check code. The triple joint judgment mechanism effectively avoids collision misjudgment that may be caused by a single hash value, ensuring that only data blocks with truly identical content will be judged as complete duplicate dependencies, thereby guaranteeing the accuracy and losslessness of data reduction and preventing data corruption caused by misjudgment.

[0138] In some embodiments, a hash index table can be built for each data block in the data block set based on its content hash value. The key value of the index table can be the content hash value, and the table entries record the storage location, data length, and fast checksum of the corresponding data block. After inputting the data block to be judged, the candidate is first located in the hash index table based on its content hash value. Then, the content hash value, data length, and fast checksum of the candidate are compared with the data block to be judged for consistency. If all three pieces of information are consistent, the judgment result that there is a complete duplicate dependency relationship between the two is output, and the result is written to the dependency table for subsequent retention data block determination and reconstruction information generation. The content hash value can be generated using a SecureHash Algorithm (SHA), and the fast checksum can be generated using a lightweight cumulative checksum or cyclic redundancy check.

[0139] Optionally, the content hash value first completes the candidate range convergence, the data length further excludes data blocks of inconsistent sizes, and the fast checksum supplements the confirmation in cases of hash collisions or partial consistency, thus forming a multi-feature joint determination for completely duplicate dependencies. This method enables the dependency relationship between the first basic data block and the first data block to be reduced to be accurately identified, and provides reliable basic dependency information for subsequent reduction processing.

[0140] By employing the aforementioned recording information and judgment methods, the dependencies between completely duplicated data blocks in the data block set can be identified more accurately, reducing misjudgments caused by relying solely on a single hash feature, and making the distinction between the dependent and dependent parties more stable. Therefore, subsequent reduction processing can determine the data blocks to be retained and the data blocks to be reduced based on accurate completely duplicated dependencies, thereby improving the accuracy of dependency identification and the reliability of association recovery.

[0141] S2023: Based on the recorded information, perform differential reconfigurable dependency determination on the data blocks to obtain the differential reconfigurable dependency relationships between the data blocks.

[0142] The differentially reconfigurable dependency includes a second basic data block as the dependent party and a second data block to be reduced as the dependent party.

[0143] Optionally, the recorded information also includes a rolling fingerprint set; based on the recorded information, differential reconfigurable dependency determination is performed on the data blocks to obtain differential reconfigurable dependency relationships between the data blocks, including: recalling candidate basic data blocks from the data block set based on the rolling fingerprint set of the data blocks; generating differential patches for the data blocks relative to the candidate basic data blocks; performing recovery verification based on the candidate basic data blocks and the differential patches; if the recovered data block is consistent with the data length, content hash value, and fast checksum of the data block, it is determined that there is a differential reconfigurable dependency relationship between the data block and the candidate basic data block; wherein, the candidate basic data block is the second basic data block, and the data block is the second data block to be reduced.

[0144] Rolling fingerprints refer to the sequence of hash values ​​obtained by calculating a sliding window across a data block. Specifically, the data block is divided into multiple fixed-length segments (e.g., each segment is 128 bytes long). A rolling hash algorithm is used to calculate the hash value of each segment, and the hash values ​​of all segments constitute the rolling fingerprint set of the data block. Rolling fingerprints have incremental computation characteristics; when the window slides one byte, the new fingerprint value can be quickly updated based on the previous fingerprint value without recalculating the hash value of the entire window. Therefore, they are computationally efficient and suitable for fast similar data detection and recall of candidate base data blocks.

[0145] As one possible implementation, generating a differential patch includes: establishing a fragment index table for small segments of the candidate base data block, the fragment index table including the fragment rolling fingerprint, strong check value, and fragment offset address; scanning from the beginning position of the data block to be judged, if a small segment with the same rolling fingerprint and strong check value is found in the fragment index table of the candidate base data block, the matching length is extended backward from this small segment until an inconsistent byte appears, and a copy instruction is generated; if the current position cannot form a valid match with any small segment in the candidate base data block, the consecutive unmatched bytes from the current position are treated as literal data, and an insertion instruction is generated; copy instructions and insertion instructions are generated sequentially according to the byte order of the data block to be judged until the scanning of the data block to be judged is completed; the copy instruction, insertion instruction, candidate base data block identifier, candidate base data block version number, original data block length, original data block hash value, and patch check value together constitute the differential patch.

[0146] Here, by recalling candidate base data blocks through a rolling fingerprint set, candidate blocks that may constitute differential dependencies can be quickly located, effectively narrowing the search range for similar block matching and improving processing efficiency. Generating differential patches enables a simplified representation of similar data, reducing storage space usage. Recovery verification ensures that the differential patches can losslessly restore the original data, effectively avoiding the risk of unrecoverable data due to inaccurate differential algorithms or corrupted patch data.

[0147] Both the first and second basic data blocks mentioned above are basic data blocks that cannot be reduced in size. The first basic data block is the basic data block in a completely duplicate dependency relationship. The current data block has the same content as the first basic data block, and the first basic data block can be directly copied during recovery. The second basic data block is the basic data block in a differentially reconfigurable dependency relationship. The current data block has partial similarities to the second basic data block, and reconstruction requires combining the second basic data block and differential patches during recovery. Both serve as the basis for recovery and are both retained data blocks; the difference lies in the recovery method—the former is direct copying, while the latter is differential reconstruction.

[0148] Optionally, the recorded information may include at least one of the following: content hash value, data length, fast checksum, and rolling fingerprint set. When determining a complete duplicate dependency, candidate blocks are first located in the data block set based on the content hash value and data length. Then, the fast checksum is compared for consistency. If all three are consistent, a complete duplicate dependency relationship exists between the candidate block and the current data block. The retained block is designated as the first basic data block, and the reducible block is designated as the first data block to be reduced. When determining a differentially reconfigurable dependency, candidate basic data blocks are first recalled based on the rolling fingerprint set. Then, a differential patch is calculated between the current data block and the candidate basic data blocks. Recovery verification is performed using the candidate basic data blocks and the differential patch. When the recovery result matches the current data block, a differentially reconfigurable dependency relationship exists between them, and they are respectively designated as the second basic data block and the second data block to be reduced.

[0149] In this process, fully repeatable dependencies and differentially reconfigurable dependencies correspond to different reduction paths. The former establishes dependencies by direct reference, while the latter establishes dependencies by adding differential information to the second basic data block, thereby outputting dependency results that can be retained, reduced, and reconfigured later.

[0150] In some embodiments, the second data block to be reduced can first be calculated by rolling according to a preset window to obtain multiple rolling fingerprints and write them into the record information. Then, based on the fingerprint index, a matching recall is performed in the data block set to form a candidate basic data block set. Subsequently, the second data block to be reduced and each candidate basic data block are compared for differences to generate corresponding differential patches. The candidate basic data blocks and differential patches are then input into the recovery module to generate the recovered data block. After recovery, the length, content hash value, and fast check code of the recovered data block are compared item by item with the record information of the second data block to be reduced. If the three are consistent, the differential reconstructable dependency relationship is confirmed, and the candidate basic data block is determined as the second basic data block; if any item is inconsistent, the dependency relationship is not established. In practical applications, the differential patch can adopt byte-level differential encoding, block-level offset encoding, or compressed representation. This application embodiment does not limit this.

[0151] According to the dependency determination method in this embodiment, dependency determination is divided into two types: complete duplication dependency determination and differentially reconfigurable dependency determination. This comprehensively covers possible dependency scenarios between data blocks, including both duplicate data with completely identical content and differentially similar data. By explicitly defining each dependency relationship as consisting of a base data block as the dependent party and a data block to be reduced as the dependent party, the expression of dependency relationships is clear and unified. This provides an accurate data foundation for the subsequent establishment of candidate dependency relationship tables and the determination of data block status identifiers, which is beneficial to improving the accuracy and processing efficiency of reduction decisions.

[0152] Based on the above-mentioned determination of dependencies, the embodiments of this application can achieve accurate filtering of retained data blocks and data blocks to be reduced.

[0153] In one possible implementation, determining at least one reserved data block and at least one data block to be reduced from the data block set based on dependencies includes: establishing a candidate dependency table to characterize the dependencies of data blocks; determining the status identifier of the data blocks based on the candidate dependency table; and determining at least one reserved data block and at least one data block to be reduced from the data block set based on the status identifier.

[0154] The initial state of the status identifier is configured as pending block. The status identifier also includes core reserved block, ordinary reserved block, candidate reduced block and reduced block.

[0155] Here, a pending block is the initial state identifier after a data block is generated, indicating that the data block has generated metadata but its state identifier has not yet been determined. After the state identifier determination process, a pending block is converted into one of three types based on the dependency determination result: a core reserved block, a regular reserved block, or a candidate reduction block. The pending block itself does not exist as a final state at the end of the processing flow; its role is to provide a unified initial entry point for state identifiers, ensuring that each data block does not participate in any reduction decisions before the dependency determination is completed.

[0156] Optionally, a candidate dependency table is first generated based on fully repeatable dependencies and differentially reconfigurable dependencies, and the dependent objects, number of dependencies, and recoverable attributes of each data block are written into the table. Then, using pending blocks as the initial state, the status of each data block is determined. When a data block is referenced by multiple data blocks to be recovered and plays a fundamental supporting role, it is marked as a core reserved block. When a data block no longer bears the dependency required for reduction, it is marked as a regular reserved block. When a data block still has dependencies but needs to be verified for recovery first, it is marked as a candidate reduction block, and after successful verification, it is converted into a reduced block. Based on the above status identifiers, the system can directly filter reserved data blocks and data blocks to be reduced from the data block set, where core reserved blocks and regular reserved blocks correspond to reserved data blocks, and candidate reduction blocks and reduced blocks correspond to data blocks to be reduced.

[0157] By centrally representing dependencies in a candidate dependency table and classifying them using a unified state identification system, the decision to retain or reduce data blocks can be based on clear criteria, avoiding erroneous reduction decisions for critical foundational data blocks. At the same time, the selection of data blocks to be reduced and subsequent recovery verification are seamlessly connected, thus providing a stable foundation for data recovery during the reduction process and ensuring that the storage results are consistent with the dependencies.

[0158] In one possible implementation, determining at least one reserved data block and at least one data block to be reduced from the data block set based on the status identifier includes: if the status identifier of the data block is a candidate reduction block or a reduced block, then the data block is determined to be a data block to be reduced; if the status identifier of the data block is a core reserved block or a normal reserved block, then the data block is determined to be a reserved data block.

[0159] Optionally, each data block in the data block set is associated with a status identifier field, which can be stored in a candidate dependency table or a block status table and maintained by the status management unit. When the dependency determination result indicates that a data block is a dependent party and its recovery link has not yet been verified, its status identifier is written as a candidate reduced block; when the data block has passed recovery verification and completed the reconstruction information generation, its status identifier is updated to a reduced block. For data blocks determined to be critical basic data blocks, their status identifiers are set as core reserved blocks; for data blocks that only need to be retained within the current storage cycle but do not bear critical recovery foundation, their status identifiers are set as ordinary reserved blocks. The update of the status identifier can be completed based on a hash mapping table, a flag array, or a database record field. In practical applications, this status management structure can also be selected in other models or implementation forms, which are not limited in this embodiment.

[0160] After the status identifier is written, the data block set is classified and output according to the current status identifier. Data blocks with the status identifier of candidate reduction block or already reduced block are assigned to the data block set to be reduced, while data blocks with the status identifier of core reserved block or ordinary reserved block are assigned to the reserved data block set. This classification result can be directly used for subsequent recovery verification, reconstruction information generation, and redundancy protection configuration calls, so that the reduction process only applies to data blocks that meet the reduction conditions, while the data blocks corresponding to the recovery basis are continuously kept in the reserved set.

[0161] It should be noted that the term "data block to be reduced" is a concept from the storage decision-making stage. It refers to data blocks that require reduction processing, including candidate reduction blocks that have not yet been reduced and already reduced blocks that have been reduced. Candidate reduction blocks are data blocks that have not yet undergone reduction operations, while already reduced blocks are data blocks that have completed reduction operations. They correspond to different stages of the reduction process, but they share the common attribute of needing to be restored based on retained data blocks; therefore, both are categorized as data blocks to be reduced.

[0162] By adopting the aforementioned state-based classification method, whether a data block participates in reduction no longer depends on additional repetitive judgments or manual configuration, but is directly determined by a unified state field, thus reducing the overhead of classification judgment links and state transitions. Since core reserved blocks and ordinary reserved blocks can be directly distinguished from the data blocks to be reduced at the state level, subsequent storage, verification, and reconstruction processes for reserved and reduced data blocks can maintain consistent input boundaries, thereby ensuring that the reduction process matches the recovery constraints and improving the determinism and consistency of data block classification results.

[0163] Based on the foregoing embodiments, further, the status identifier of the data block is determined according to the candidate dependency table, including: calculating the number of times the data block is referenced as a basic data block in the dependency relationship according to the candidate dependency table; determining the core reserved block from the data block according to the number of references; determining the data block that is not determined as a core reserved block and has no dependency relationship as a normal reserved block; determining the data block that is not determined as a core reserved block but has a dependency relationship as a candidate reduced block; and determining the data block that passes the recovery verification among the candidate reduced blocks as a reduced block.

[0164] Optionally, the granularity of citation counts is at the processing partition level, meaning that the number of times each data block in each processing partition is referenced as a base data block is counted independently. Citation counts between different processing partitions do not affect each other. By using partition-level statistics, the election results of core reserved blocks can be matched with the dependencies within each partition, avoiding the problem of highly referenced base blocks being concentrated in one partition while other partitions lack core reserved blocks due to global statistics.

[0165] The basic data block includes the first basic data block and the second basic data block.

[0166] Here, the candidate dependency table serves as the data foundation for determining the state identifier. The determination of the state identifier relies on the dependency information recorded in the candidate dependency table, specifically including: counting the number of times a data block is referenced as a base data block in the candidate dependency table to determine the core retained block; distinguishing between ordinary retained blocks and candidate reduced blocks based on the existence of dependencies in the candidate dependency table; and updating candidate reduced blocks to reduced blocks based on the recovery verification results. The candidate dependency table centrally stores all dependency information, and the state identifier determination process obtains the decision basis by querying the candidate dependency table; the two form a data flow from dependency records to state identifier decisions.

[0167] In some embodiments, the candidate dependency table can be constructed from the dependency determination results into a mapping table or an adjacency table. Each entry records at least the base data block identifier, the dependent identifier, and a reference count field. The reference count can be accumulated after traversing the table entries, or the counter can be updated synchronously when the dependency is established to reduce the overhead of repeated queries. For data blocks with high reference counts, they are directly designated as core reserved blocks and their subsequent reduction permissions are locked, ensuring their complete retention status at the storage layer. For data blocks without dependencies, after confirming they are not selected as core reserved blocks, the system designates them as ordinary reserved blocks and includes them in the basic storage set. For data blocks with dependencies, the system first designates them as candidate reduction blocks, then calls the recovery verification module to perform a recovery comparison based on the corresponding base data block and reconstruction information. When the recovery result is consistent with the original data block in terms of data length, content hash value, and fast checksum, the status identifier is updated to a reduced block. This process ensures that the status identifier is consistent with the dependency strength and recovery result, thereby achieving differentiated management of core data blocks, ordinary data blocks, and reducible data blocks, and maintaining the integrity and traceability of the dependency chain in subsequent storage and recovery. In terms of working principle, partitioning limits risk assessment to a local data block. The reduction information reflects the reduction scale, the additional information load for recovery, and the degree of dependency concentration. These three factors together characterize the recovery vulnerability of the data blocks to be reduced within the partition. The reduction risk parameter can increase with the proportion of reduced data, the proportion of reconstructed information, or the concentration of the base blocks, thus providing a quantitative basis for determining the subsequent redundancy protection strength.

[0168] Using the above method, the risk reduction parameters can reflect the recovery risk of the data block to be reduced at the partition granularity, so that the risk assessment results match the reduction status and dependencies within the partition, and provide consistent input for subsequent protection strength configuration, thereby improving the synergy between reduction processing and recovery reliability.

[0169] In one possible implementation, the steps for establishing dependencies and creating a candidate dependency table are as follows:

[0170] First, establish data block status identifiers. Set status identifiers for multiple data blocks, including at least pending blocks, core reserved blocks, fixed reserved blocks, candidate reduced blocks, reduced blocks, and ordinary reserved blocks.

[0171] Among them, the core reserved block is a data block that is suitable as the basis for the recovery of other data blocks; the solidified reserved block is a data block that has completed encryption, erasure coding, physical writing and write verification; the candidate reduction block is a data block that has a valid base data block and has the benefit of reduction; and the reduced block is a data block that has generated a reconstructed witness index and passed the pre-deletion verification.

[0172] It's important to note that core reserved blocks and solidified reserved blocks represent two distinct states in the status identifier: a core reserved block is the state of a high-value underlying data block elected during the reduction decision phase, indicating that this data block must not be deleted in the current processing batch, but may not yet have completed physical write operations; a solidified reserved block is the final stable state of a core reserved block or ordinary reserved block after encryption, erasure coding, physical write operations, and write verification, indicating that the data block has been securely written to disk. Core reserved blocks are prioritized for addition to the write queue to complete solidification as quickly as possible, and solidified reserved blocks represent the target state for core reserved blocks.

[0173] In this embodiment of the application, candidate reduced blocks may not be used as the base data blocks for other candidate reduced blocks. The reconstructed witness index of a reduced block is only allowed to reference core reserved blocks, solidified reserved blocks, or ordinary reserved blocks, so as to avoid circular dependencies between reduced data blocks.

[0174] Optionally, the constraint that a candidate reduction block cannot be the base data block for other candidate reduction blocks is implemented as follows: During the dependency graph anti-loop verification process, if the base data block y of a candidate reduction block x is detected to be in a candidate reduction block state or an already reduced block state, the verification is directly determined to fail, and the candidate reduction block x is converted into a regular reserved block. Through this constraint, the system forces all data blocks to be reduced to depend only on core reserved blocks, fixed reserved blocks, or regular reserved blocks, thereby limiting the maximum dependency depth to 1 and avoiding the generation of multi-level dependency chains and circular dependencies.

[0175] The difference between ordinary reserved blocks and persistent reserved blocks lies in the following: Ordinary reserved blocks represent the state of reserved data blocks determined during the reduction decision phase, indicating that the data block does not meet the reduction conditions or has been identified as a reserved object, but may not yet have completed physical writing; persistent reserved blocks represent the final stable state of ordinary or core reserved blocks after completing encryption, erasure coding, physical writing, and write verification, indicating that the data block has been securely persisted to the storage medium. Ordinary reserved blocks are converted into persistent reserved blocks after completing the above-mentioned persistence process.

[0176] Next, a complete duplicate dependency determination is performed. A content hash index table is created, and based on the content hash index table, it is determined whether any data block has a complete duplicate dependency with an existing data block.

[0177] Specifically, for data block x, if data block y is found, and data block x and data block y satisfy the same content hash value, the same original length of data block, and the same fast check code, then it is determined that data block x can be completely recovered by data block y, and a candidate dependency edge: xy is generated.

[0178] In this context, candidate dependency edges indicate that data block x can be recovered by referencing the underlying data block y.

[0179] When there are multiple base data blocks that meet the condition of complete duplication dependency, the data block that is already in the state of solidified reserved block or the data block that is already in the state of core reserved block shall be selected as the base data block.

[0180] Next, differential reconfigurable dependency determination is performed. For data blocks that do not form completely duplicate dependencies, candidate base data blocks are recalled based on the rolling fingerprint set, and differential patches are generated for the data block to be reduced relative to the candidate base data blocks.

[0181] Specifically, the data block is divided into multiple small segments, a rolling fingerprint is calculated for each segment, and a fingerprint inverted index is established; for the data block to be judged, the candidate basic data block set is recalled from the fingerprint inverted index according to its rolling fingerprint set; and the candidate basic data block is selected to enter the differential patch generation process according to the number of matching segments between the candidate basic data block and the data block to be judged.

[0182] For a candidate base data block y, copy and insert instructions are generated according to the byte order of the data block x to be judged. The copy instruction is used to copy data with a specified offset address and a specified length from the base data block y, and the insert instruction is used to write literal data. The copy and insert instructions together constitute the differential patch of data block x relative to the base data block y.

[0183] When the total length of the differential patch and its metadata meets the preset differential benefit condition, a recovered data block is obtained based on the base data block and the differential patch. It is then determined whether the recovered data block and the original data block are consistent in data length, content hash value, and fast checksum. If they are consistent, a valid differentially reconfigurable dependency is formed between data block x and base data block y, and a candidate dependency edge xy is generated. If they are inconsistent, the differentially reconfigurable dependency is deemed invalid.

[0184] Finally, a candidate dependency table is created. This table is based on fully duplicated dependencies and effectively differentially reconfigurable dependencies.

[0185] Specifically, the candidate dependency table includes at least one of the following: the identifier of the reconstructed data block, the identifier of the candidate base data block, the dependency type, the length of the differential patch, the length of the metadata, the expected number of bytes to be reduced, the recovery verification result, the current status of the base data block, the generation time of the dependency edge, and the valid identifier of the dependency edge.

[0186] For fully duplicated dependencies, the expected reduction in bytes is determined based on the original data block length and the reference index length. For differentially reconfigurable dependencies, the expected reduction in bytes is determined based on the original data block length, the differential patch length, the metadata length, and the reconfiguration witness index length. If the expected reduction in bytes is less than or equal to zero, the corresponding candidate dependency will not proceed to the subsequent reduction process.

[0187] Optionally, the expected reduction in bytes is calculated as follows: for fully duplicated dependencies, the expected reduction in bytes equals the original data block length minus the reference index length; for differentially reconfigurable dependencies, the expected reduction in bytes equals the original data block length minus the sum of the differential patch length, metadata length, and reconfiguration information length. If the calculated expected reduction in bytes is less than or equal to zero, it indicates that the dependency relationship cannot bring positive storage benefits, and the system marks the candidate dependency edge as invalid, preventing it from entering the subsequent reduction process. Through this economic judgment mechanism, the system can automatically filter out dependencies with no or negative benefits, avoiding invalid calculations for the sake of reduction, and improving the efficiency of storage resource utilization.

[0188] Based on the above candidate dependency table, the steps for determining the retained data blocks and candidate reduced blocks are as follows:

[0189] First, determine the core reserved blocks. For any processing partition, calculate the number of references and the expected total reduction bytes for each data block when used as a base data block based on the candidate dependency table, and determine the core reserved blocks accordingly.

[0190] Specifically, for data block y, the number of data blocks within the processing partition that use data block y as a candidate base data block is counted to obtain the reference count; and the basic savings are obtained based on the expected reduction in bytes for each reconstructed data block when data block y is used as the base data block. Further, a base block score is calculated based on the reference count, the basic savings, and the reliability level of the storage location where the data block resides.

[0191] Data blocks whose basic block score is greater than a preset core threshold, whose reference count is greater than a preset reference count threshold, whose basic savings are greater than a preset basic savings threshold, whose data blocks have been referenced by the reconstructed witness index of other processing partitions, or whose data blocks belong to critical data blocks are identified as core reserved blocks. It is understood that each threshold in this embodiment can be determined according to actual circumstances, and no specific restrictions are imposed here.

[0192] Data blocks identified as core reserved blocks (i.e., core data blocks in this article) must not be deleted in the current processing batch; if a core reserved block has not yet been physically written, it is added to the priority write queue.

[0193] Next, candidate reduction blocks are determined. For data blocks in the processing partition that have not been identified as core reserved blocks, their suitability as candidate reduction blocks is determined based on the candidate dependency table.

[0194] Specifically, when data block x has at least one valid candidate dependency edge, and its underlying data block y is a core reserved block, a fixed reserved block, or a regular reserved block, and the underlying data block y is not in the candidate reduction block state or the already reduced block state, and the expected reduction ratio of data block x after reconstructing the witness index is not lower than the preset minimum reduction ratio, data block x is determined as a candidate reduction block.

[0195] To ensure data recoverability, this application provides multiple verification mechanisms before data reduction.

[0196] Optionally, the recovery verification of the data block to be reduced includes: generating directed dependency edges based on the data block to be reduced and the corresponding basic data block; performing dependency graph anti-loop verification based on the directed dependency edges; and determining that the recovery verification fails if the dependency graph anti-loop verification fails.

[0197] Specifically, after identifying the data block to be reduced and its corresponding underlying data block, the dependency relationship between them is encoded as a directed dependency edge, where the starting point of the edge points to the dependent party and the ending point points to the dependent party. Then, a dependency graph is constructed based on all directed dependency edges, and loop prevention checks are performed on the dependency graph using graph traversal, depth-first search, topological sorting, or path backtracking. If, during the check, a node is detected to have returned to a previously visited node along the dependency path, a loop is determined to exist in the dependency graph, thus confirming that the recovery verification of the data block to be reduced fails. If the recovery path of the data block to be reduced also includes multiple levels of underlying data blocks, the dependencies at each level can be included in the dependency graph for unified verification to ensure the integrity and consistency of the recovery path.

[0198] This application's embodiments abstract the recovery relationship between the data block to be reduced and the basic data block into a directed graph structure, and detect loops in the graph during the recovery verification phase, ensuring that the recovery path is constrained to a circular dependency-free structure before entering subsequent reduction or reconstruction processing. Therefore, the recovery verification result directly reflects whether the data block to be reduced has an executable recovery basis, and immediately outputs "recovery verification failed" when a circular dependency occurs, providing a clear rejection basis for reduction decisions.

[0199] Optionally, the recovery verification of the data block to be reduced includes: obtaining the basic data block corresponding to the data block to be reduced; performing a basic data availability verification on the basic data block; and determining that the recovery verification fails if the basic data availability verification fails.

[0200] Optionally, when obtaining the base data block corresponding to the data block to be reduced, the corresponding record between the data block to be reduced and the base data block can be read according to the dependency table, and the base data block identifier in the corresponding record can be used as the retrieval condition to locate the target base data block from the data block set or storage index. When performing basic data availability verification on the base data block, the verification value, length information and access status of the base data block can be read and compared with the preset valid conditions; when the base data block is missing, the verification value does not match, the length is inconsistent, or the access status indicates that it is unrecoverable, it is determined that the basic data availability verification fails. If the verification fails, the result of recovery verification failure is directly output, and the subsequent recovery confirmation process based on the base data block is terminated.

[0201] In this process, input location is completed based on the association between the data block to be reduced and the base data block. Then, availability verification is performed on the base data block, and a verification conclusion is output. This method moves recovery verification to the base data block status verification stage, ensuring that the data block to be reduced only enters the subsequent reconstruction process when its base data block meets the recovery conditions. This allows the recovery verification result to directly reflect whether the recovery basis is valid, and ensures that recovery processing continues only when the base data block is valid, thus making the determination of the reduced data reconstruction more accurate.

[0202] Optionally, the recovery verification of the data block to be reduced includes: obtaining the reconstruction information corresponding to the data block to be reduced; performing reconstruction correctness verification on the data block to be reduced based on the reconstruction information; if the reconstruction correctness verification passes, allowing the release of the physical storage of the data block to be reduced or canceling the physical write task of the data block to be reduced; if the reconstruction correctness verification fails, converting the data block to be reduced into a reserved data block.

[0203] In this embodiment, when obtaining the reconstruction information corresponding to the data block to be reduced, the corresponding recovery basis information can be read from the candidate dependency table, the recovery verification result, or the reconstruction description associated with the data block to be reduced. When performing reconstruction correctness verification based on the reconstruction information, the recovery result of the data block to be reduced can be compared with its data length, content hash value, and fast check code. Alternatively, consistency judgment can be made by combining the recovery order, reference relationship, and differential patch matching results. When the recovery result is consistent with the data block to be reduced in terms of data length, content hash value, and fast check code, the reconstruction correctness verification is deemed to have passed. After the verification passes, the storage layer can reclaim the corresponding physical storage or directly cancel the unfinished physical write task to release the reserved storage resources. When the verification fails, the corresponding resources are not released, and the data block to be reduced is remarked as a reserved data block so that it can continue to participate in subsequent redundancy protection or recovery links.

[0204] This approach verifies the reconstruction information before releasing physical resources, confirming the recoverability of the data block to be reduced before performing resource reclamation or task cancellation. If the verification fails, it switches to a retained data block, thus ensuring that the reduction processing and recovery reliability of the data block to be reduced are consistent, reducing the risk of data unrecoverability caused by erroneous reduction.

[0205] In one possible implementation, the embodiments of this application can implement verification in the following three ways. It should be noted that the recovery verification includes the following three levels of verification: first, dependency graph loop prevention verification, used to verify whether the dependency relationship between the candidate reduced block and the basic data block constitutes a loop or exceeds the maximum dependency depth; second, basic data availability verification, used to verify whether the basic data block is in a safe and reliable state; and third, reconstruction correctness verification (i.e., pre-deletion reconstruction verification), used to verify whether the original data can be recovered without loss based on the reconstruction information. A candidate reduced block can only be determined as a reduced block after passing the above three verifications in sequence. The three are progressive; if any verification fails, the process terminates, and the candidate reduced block becomes a regular retained block.

[0206] First, perform dependency graph loop prevention verification and basic block solidification verification.

[0207] Specifically, a directed dependency graph is generated based on the candidate reduced blocks and their underlying data blocks. For a candidate reduced block x and its underlying data block y, a directed edge : xy is generated.

[0208] Before allowing candidate reduced block x to enter the reduced state, perform dependency graph loop prevention verification and basic block solidification verification.

[0209] Optionally, the dependency graph anti-loop check includes: determining whether the addition of a directed edge causes a loop in the dependency graph, and determining whether the depth of the recovery path starting from the candidate reduction block x exceeds the maximum dependency depth.

[0210] Optionally, the basic block solidification verification includes: determining whether the basic data block y is in the state of core reserved block, solidified reserved block or ordinary reserved block; determining whether the basic data block y has been added to the protection write plan; and determining whether the version number of the basic data block y is consistent with the version number in the candidate dependency table.

[0211] Optionally, if the dependency graph loop prevention check or the basic block solidification check fails, the candidate reduced block x is converted into a normal reserved block, without releasing its physical storage or canceling its physical write task.

[0212] Secondly, a refactoring witness index is generated. For candidate reduced blocks that pass the dependency graph anti-loop check and the basic block solidification check, a refactoring witness index is generated.

[0213] The refactoring witness index includes at least one of the following: the identifier of the reduced data block, the file identifier to which the reduced data block belongs, the offset address of the reduced data block in the original data stream, the identifier of the base data block, the version number of the base data block, the refactoring type, the differential patch data, the length of the original data block, the hash value of the original data block, the fast checksum of the original data block, the refactoring algorithm identifier, the version number of the refactoring witness index, the dependency depth, the index integrity check value, and the index generation time.

[0214] The reconstruction types include fully repetitive reconstruction and differential reconstruction. When the reconstruction type is fully repetitive reconstruction, the differential patch data is empty; when the reconstruction type is differential reconstruction, the differential patch data includes a patch header and a patch instruction sequence, and the patch instruction sequence includes copy instructions and insert instructions.

[0215] Finally, perform pre-deletion reconstruction verification. Before releasing the physical storage of the candidate shrink block or canceling the physical write task of the candidate shrink block, perform pre-deletion reconstruction verification based on the reconstruction witness index.

[0216] Specifically, the underlying data block is read based on the reconstruction witness index, and the version number of the underlying data block is verified; when the reconstruction type is a fully repeatable reconstruction, the underlying data block is copied to obtain the recovery data block; when the reconstruction type is a differential reconstruction, the recovery data block is generated based on the copy and insert instructions in the differential patch.

[0217] Next, it is determined whether the recovered data block is consistent with the original data block in terms of data length, content hash value, and fast checksum. If they are consistent, the pre-deletion reconstruction verification passes; if they are inconsistent, the pre-deletion reconstruction verification fails.

[0218] Candidate blocks to be reduced are marked as reduced blocks only if the pre-deletion reconstruction verification passes and the length of the reconstruction witness index meets the preset minimum reduction ratio requirement; otherwise, the candidate blocks to be reduced are converted into ordinary reserved blocks.

[0219] Optionally, in this embodiment of the application, a temporary review can be performed before storing the reduced data to further reduce the risk of data loss.

[0220] Optionally, the staging area is a temporary buffer zone used to store intermediate states of the partition. Its function is to: provide a temporary storage space for the reconstruction information, dependency graph summary, and partition recovery verification summary of the shrunk blocks before final confirmation that all reduction operations within the partition are correct and effective; before the partition-level review is completed, the data in the staging area is not considered to be officially persisted; if the partition-level review fails, the data in the staging area will be completely deleted, implementing a rollback operation; if the partition-level review passes, the data in the staging area will enter the subsequent encryption, erasure coding, and physical write processes. The existence of the staging area makes partition-level transactional rollback possible.

[0221] Specifically, the temporary storage review is performed as follows: After completing the pre-deletion reconstruction verification of candidate reduced blocks within the processing partition, the processing partition is written to the temporary storage area. The temporary storage area stores at least one of the following information: core reserved blocks, ordinary reserved blocks, reconstruction witness indexes corresponding to reduced blocks, dependency graph summary, partition recovery verification summary, write plan, and original data block cache release status.

[0222] Among them, the dependency graph summary refers to the summary information obtained after compressing and recording the directed dependencies between all candidate reduced blocks and basic data blocks in the processing partition. It includes at least one of the following: the number of nodes in the dependency graph, the number of edges, the verification result of whether there is a loop, and the dependency graph version number. It is used to quickly verify the consistency of the dependencies during subsequent partition-level review.

[0223] Optionally, a partition-level review is performed before releasing the original data block cache. For high-priority data streams, a recovery review is performed on all reduced blocks within the processing partition; for ordinary data streams, a sampling review is performed according to a preset ratio.

[0224] Optionally, if the partition-level review passes and the core reserved blocks, ordinary reserved blocks, and reconstructed witness indexes have all entered the subsequent encrypted erasure protection process, the original cache of the reduced blocks is released or the physical write task of the reduced blocks is canceled; if the partition-level review fails, the corresponding data reduction operation is rolled back.

[0225] In some embodiments, the rollback operation specifically includes at least one of the following: restoring all data blocks in the processing partition that are marked as reduced blocks to candidate reduced blocks or ordinary retained blocks; releasing the generated reconstruction information; restoring the corresponding original data block cache or physical write task; marking the dependency edges related to the partition in the candidate dependency table as invalid; and re-incorporating the relevant data blocks into the subsequent storage process. The rollback operation ensures the transactionality of the reduction operation, meaning that the reduction operations within a partition either all succeed or are all canceled, avoiding data inconsistency caused by some data blocks in the partition successfully reducing while others fail.

[0226] Optionally, after the data in the staging area passes partition-level review, it will enter the subsequent encryption and erasure coding process. Specifically, the reconstruction information corresponding to the core reserved blocks, ordinary reserved blocks, and reduced blocks is read from the staging area, encapsulated into a data packet to be protected, and encrypted. After encryption, the number of erasure coding check symbols is determined according to the reduction risk parameters, and erasure coding is performed to generate check symbols. Finally, the encoded data is written to the physical storage medium. If any step fails, the original state saved in the staging area is used for rollback. Through the collaboration between the staging area and the encryption and erasure processes, the reduction operation is ensured to have complete rollback capability before entering the final persistent storage.

[0227] Figure 4 Flowchart of the data reduction processing method provided in the embodiments of this application Figure 2 See also Figure 4 This embodiment provides a data reduction processing method, which can be applied to scenarios such as database backup, log archiving, file synchronization, model sample archiving, sensor data storage, and distributed object storage. This method can be executed by a data storage server, backup server, distributed storage node, or cloud storage platform.

[0228] The device executing this method may include a processor, memory, a data receiving interface, a data block segmentation module, a fingerprint generation module, a candidate base block recall module, a differential patch generation module, a dependency management module, a core reserved block determination module, a reconstruction information generation module, a pre-deletion verification module, a temporary storage review module, an encryption module, an erasure coding module, and a physical write module, etc. The memory stores program instructions executable by the processor. When the processor executes the program instructions, it implements the following steps.

[0229] In this embodiment, data reduction refers to the ability to recover the original data block based on the retained data block and the corresponding reconstruction information for data blocks that were not directly written to the physical medium or released from the physical storage. The recovered data block is consistent with the original data block in terms of data length, content hash value, and fast check code.

[0230] S401: Receive input data stream.

[0231] Optionally, an input data stream can be received via a data receiving interface. The input data stream can be database backup data, log archive data, file synchronization data, object storage data, model sample data, or sensor-collected data. Upon receiving the input data stream, the system writes it to the input buffer and sends a splitting trigger message to the data block splitting module. The splitting trigger message includes the input data stream identifier, data source identifier, data length, data type, and processing priority.

[0232] S402: Perform data block segmentation on the input data stream.

[0233] The input data stream is divided into multiple data blocks according to a preset segmentation rule. The preset segmentation rule can be determined according to the actual situation, and this application embodiment does not impose specific restrictions on it.

[0234] Optionally, the preset segmentation rule can be a fixed-length segmentation rule, a content-boundary-based rolling hash segmentation rule, or a file structure-based field-boundary segmentation rule. As a specific implementation, the system uses a fixed-length segmentation method to divide the input data stream into data blocks of a target length of 16KB; when the remaining data at the end of the input data stream is less than 16KB, the last data block is formed according to the actual length. The multiple data blocks obtained from the segmentation constitute the current processing batch.

[0235] S403: Generate metadata based on the split data blocks.

[0236] Generate a data block record for each data block.

[0237] The content hash value can be generated using any hash algorithm and is used to determine whether the data block is completely duplicated; the fast check code can be generated using any fast check code generation method and is used to assist in verifying the data block recovery result; the rolling fingerprint set is used for subsequent recall of data blocks that may be suitable as the basis for differential retrieval.

[0238] S404: Perform a full duplicate dependency determination based on metadata.

[0239] Create a content hash index table and perform full duplicate dependency determination on the data blocks in the current processing batch based on the content hash index table.

[0240] S405: Perform differential reconfigurable dependency determination.

[0241] For data blocks that do not form complete duplicate dependencies, differential reconfigurable dependency determination is further performed. Differential reconfigurable dependency determination includes candidate base block recall, differential patch generation, and differential recovery verification. Each data block is divided into multiple fixed-length small segments. As a specific implementation, the segment length is 128 bytes. The system calculates a rolling fingerprint for each segment and writes the rolling fingerprint into a fingerprint inverted index. The entries in the fingerprint inverted index include the rolling fingerprint value, data block identifier, segment offset address, and segment strong checksum value.

[0242] For the data block x to be judged, the system recalls a set of candidate basic blocks from the fingerprint inverted table based on the rolling fingerprint set of data block x. The system counts the number of matching segments between each candidate basic block and data block x, sorts them from high to low according to the number of matching segments, and selects the top K candidate basic blocks to enter the differential patch generation process. In this embodiment, K is 32.

[0243] If the candidate base block set is empty, or the number of matching fragments of the candidate base block is lower than the preset recall threshold, the system will not perform differential reduction on data block x, and will convert data block x into a normal reserved block.

[0244] Optionally, the system determines whether the differential patch has a reduction benefit. Specifically, the system calculates the total length of the differential patch and its metadata. If the differential benefit condition is met (i.e., the total length of the differential patch and its metadata is less than or equal to a preset ratio multiplied by the original data block length), then the differential patch is determined to have a reduction benefit. In this embodiment, the differential benefit threshold η is set to 0.5.

[0245] After determining that the differential patch has reduction benefits, the system performs differential recovery verification. Specifically, the system generates a recovery data block x' based on the base data block y and the differential patch, and determines whether the recovery data block x' is consistent with the original data block x in terms of data length, content hash value, and fast checksum. If they are consistent, it is determined that a valid differentially reconfigurable dependency is formed between data block x and base data block y; if any condition is not met, the differentially reconfigurable dependency is determined to be invalid.

[0246] S406: Establish a candidate dependency table based on the decision results.

[0247] S407: Generate processing partitions based on the candidate dependency table.

[0248] Optionally, the current processing batch can be divided into multiple processing partitions. Processing partitions are used to control memory usage, limit the processing scope, and facilitate subsequent erasure coding.

[0249] Optionally, processing partitions are generated according to the following rules: data blocks within the same processing partition preferentially originate from the same file, the same object, or the same time window; the number of data blocks in a single processing partition does not exceed a preset upper limit; the total length of the original data in a single processing partition does not exceed a preset upper limit; if a data block has a valid dependency relationship with a candidate basic data block, they are preferentially assigned to the same processing partition; if they cannot be assigned to the same processing partition, it must be ensured that the candidate basic data block is already in a solidified reserved block state, or it will be preferentially written and solidified in the current processing batch. As a specific implementation, the upper limit of the number of data blocks in a single processing partition is 8192, and the upper limit of the total length of the original data in a single processing partition is 128MB.

[0250] S408: Determine the core reserved block within the processing partition.

[0251] For any processing partition, the system uses the candidate dependency table to count the number of references and the expected total number of bytes to be reduced when each data block is used as a base data block, and determines the core reserved blocks accordingly.

[0252] Specifically, for data block y, the system counts the number of data blocks within the processing partition that use data block y as a candidate base data block, obtaining the reference count; and based on the expected reduction in bytes for each reconstructed data block when data block y is used as the base data block, it obtains the base savings. Further, based on the reference count, the base savings, and the reliability level of the storage location where the data block resides, a base block score is calculated.

[0253] Optionally, a data block that meets any of the following conditions will be identified as a core reserved block: the basic block score is greater than a preset core threshold; the number of references is greater than a preset reference count threshold; the basic savings are greater than a preset basic savings threshold; it has been referenced by reconstruction information of other processing partitions; or it belongs to a key data block or metadata block specified by the system.

[0254] Data blocks identified as core reserved blocks must not be deleted in the current processing batch; if a core reserved block has not yet been physically written, the system adds it to the priority write queue.

[0255] S409: Determine candidate reduction blocks.

[0256] For data blocks in the processing partition that are not identified as core reserved blocks, determine whether they are candidate reduction blocks based on the candidate dependency table.

[0257] Specifically, when data block x has at least one valid candidate dependency edge, and its underlying data block y is a core reserved block, a fixed reserved block, or a regular reserved block, and the underlying data block y is neither a candidate reduction block nor a reduced block, and the expected reduction ratio of data block x after using the reconstruction information is not lower than the preset minimum reduction ratio, data block x is determined as a candidate reduction block. In this embodiment, the minimum reduction ratio σ is taken as 0.2.

[0258] If data block x has multiple valid candidate dependency edges, then a fully duplicated dependency is selected first; if no fully duplicated dependency exists, then a differentially reconfigurable dependency with the smallest total length of differential patch and metadata is selected.

[0259] S410: Build a dependency graph and perform dependency graph loop prevention checks.

[0260] Optionally, dependency graph loop prevention verification and base block solidification verification are performed before allowing candidate reduced block x to enter the reduced state.

[0261] Optionally, the dependency graph loop prevention check includes: determining whether the newly added directed edge causes a loop in the dependency graph, and determining whether the depth of the recovery path starting from the candidate reduced block x exceeds the maximum dependency depth. In this embodiment, the maximum dependency depth is set to 1, that is, the reduced data block can only directly depend on the core retained block, the fixed retained block, or the ordinary retained block.

[0262] The basic block solidification verification includes: determining whether the basic data block y is in the state of core reserved block, solidified reserved block or ordinary reserved block; determining whether the basic data block y has been added to the protection write plan; and determining whether the version number of the basic data block y is consistent with the version number in the candidate dependency table.

[0263] If the dependency graph loop prevention check or the basic block solidification check fails, the candidate reduced block x will be converted into a normal reserved block, without releasing its physical storage or canceling its physical write task.

[0264] S411: Generate reconstruction information.

[0265] For candidate reduced blocks that pass the dependency graph anti-loop check and the basic block solidification check, refactoring information is generated.

[0266] S412: Perform a pre-deletion reconstruction check.

[0267] Before releasing the physical storage of the candidate shrink block or canceling the physical write task of the candidate shrink block, perform pre-deletion reconstruction verification based on the reconstruction information.

[0268] Candidate blocks are marked as reduced blocks only if the pre-deletion reconstruction verification passes and the length of the reconstruction information meets the preset minimum reduction ratio requirement; otherwise, the candidate blocks are converted into ordinary retained blocks.

[0269] S413: Calculate the risk reduction parameters.

[0270] The reduction risk parameter is calculated based on the reduction result of each processing partition. The reduction risk parameter is used to determine the required erasure coding protection strength for that processing partition.

[0271] For each processing partition, the system calculates the deletion ratio, index ratio, and base block concentration. The deletion ratio represents the ratio between the total original length of the reduced blocks and the total original data length of the processing partition; the index ratio represents the proportion of the total length of the reconstructed information in the total length of the retained data and the reconstructed information; and the base block concentration represents the degree to which multiple reduced blocks within the processing partition are concentratedly dependent on the same base data block.

[0272] S414: Perform encryption processing on the data.

[0273] The system encapsulates the core reserved blocks, ordinary reserved blocks, reconstruction information, dependency graph digest, and partition recovery verification digest in the processing partition into a data packet to be protected. The data packet to be protected includes a reserved data area, an information area, a verification information area, and a partition parameter area.

[0274] The system encrypts the data packets to be protected. The encryption method can be any block cipher with integrity verification capabilities. Encrypted additional authentication data may include one or more of the following: partition number, data block version number, reconstruction information version number, and partition parameter digest. If decryption authentication fails, the data packet is rejected for recovery.

[0275] S415: Perform erasure coding on the encrypted data.

[0276] S416: Physically write the encoded data into the storage medium and solidify it.

[0277] The system writes the encrypted original symbol and check symbol to different disks, different storage nodes, or different availability zones. After the writing is complete, the system reads the write receipt and performs write verification.

[0278] For core reserved blocks and ordinary reserved blocks, the system will only update their status to solidified reserved blocks after the data packet containing them has been encrypted, erasure encoded, physically written, and written verified.

[0279] For a reduced block, the system will only release its original data block cache when the corresponding basic data block has become a fixed reserved block, the corresponding reconstructed information has been protected in the information area, the corresponding processing partition has passed the review, no loops have been found in the dependency graph verification, and the physical write verification has passed; otherwise, the corresponding data block will be converted into an ordinary reserved block and written to the physical medium.

[0280] S417: Perform data read and restore operations.

[0281] The above steps S401-S417 provide a detailed description of the data reduction processing method provided in the embodiments of this application. The specific implementation methods can all refer to the content in the above embodiments.

[0282] Optionally, Figure 5 This application provides a schematic diagram of the structure of a lossless data block reduction and collaborative storage system, as shown in the embodiments. Figure 5 As shown, the system is used to execute the data reduction processing method in the above method embodiments. This system can be deployed in a data storage server, backup server, distributed storage node, edge storage device, or cloud storage platform.

[0283] in, Figure 5Dashed arrows indicate metadata or index streams, while solid arrows indicate data block streams (data streams) or recovery streams. The direction of a solid arrow determines whether it is a data block stream or a recovery stream; the target module of a dashed arrow determines whether it is a metadata stream or an index stream.

[0284] The system includes a data receiving module, a data block splitting module, a metadata generation module, a data block status management module, a fully duplicated dependency determination module, a differentially reconfigurable dependency determination module, a candidate dependency table generation module, a processing partition generation module, a core retained block determination module, a candidate reduced block determination module, a dependency graph verification module, a reconstruction information generation module, a pre-deletion reconstruction verification module, a temporary storage review module, a reduction benefit monitoring module, a reduction risk calculation module, an encryption module, an erasure coding module, a physical write module, and a data recovery module. It may also include N storage nodes, where N is any positive integer.

[0285] The aforementioned modules can be stored in memory as software program modules and executed by the processor; alternatively, they can be implemented using dedicated hardware circuits, programmable logic devices, or a combination of hardware and software. Modules interact with each other via a system bus, shared buffer, metadata table, or message queue. The specific functions of each module correspond to the steps described in the foregoing method embodiments; to avoid repetition, the detailed execution process of each module will not be elaborated here.

[0286] During system operation, the data receiving module first receives the input data stream, the data block segmentation module segments the input data stream into multiple data blocks, and the metadata generation module generates a data block record for each data block. Subsequently, the fully repeatable dependency determination module and the differentially reconfigurable dependency determination module identify fully repeatable dependencies and differentially reconfigurable dependencies, respectively, and the candidate dependency table generation module establishes a candidate dependency table.

[0287] After the processing partition generation module divides the data block into multiple processing partitions, the core reserved block determination module determines the core reserved blocks based on the number of references, basic savings, and reliability level. After the candidate reduction block determination module determines the candidate reduction blocks, the dependency graph verification module performs loop prevention verification and basic block solidification verification on the directed dependency relationship between the candidate reduction blocks and the basic data blocks.

[0288] For candidate reduced blocks that pass verification, the reconstruction information generation module generates reconstruction information, and the pre-deletion reconstruction verification module restores the candidate reduced blocks based on the reconstruction information and performs verification. After verification, the candidate reduced block is updated as a reduced block. Then, the reduction risk calculation module calculates reduction risk parameters based on the deletion ratio, index ratio, and base block concentration. The erasure coding module determines the number of verification symbols based on the reduction risk parameters and performs erasure coding on the encrypted data packet to be protected. After the physical write module completes write verification, the system releases the original cache of the reduced blocks that meet the conditions.

[0289] The data recovery module is used to locate the data packet of the corresponding processing partition according to the partition number and data block address table when a data read request is received. If some physical symbols are lost or damaged, they are recovered by erasure coding decoding. The data packet is also decrypted and the authentication tag is verified. For core reserved blocks and ordinary reserved blocks, the data content is directly output. For reduced blocks, the corresponding reconstruction information is read and full repeat reconstruction or differential reconstruction is performed. After the verification is passed, the recovered data block is output.

[0290] Through the above system structure, a continuous processing link is formed between the modules, from data block segmentation, dependency determination, core retention block determination, candidate reduction block determination, dependency graph loop prevention, reconstruction information generation, pre-deletion verification, risk adaptive erasure coding to data recovery. This can reduce the risk of data recovery failure caused by accidental deletion of basic blocks, circular dependencies, and damage to reconstruction information while achieving data block reduction.

[0291] This application's embodiments reduce the risk of unrecoverable data due to misjudgment caused by similarity by performing full duplicate dependency determination and differential reconfigurable dependency determination on data blocks, generating a reconstruction witness index and performing data block reduction only when the recovery verification passes. Simultaneously, by determining core reserved blocks, restricting candidate reduction blocks to depend only on core reserved blocks, fixed reserved blocks, or ordinary reserved blocks, and combining dependency graph anti-loop verification and maximum dependency depth control, the problem of recovery chain breakage caused by accidental deletion of basic blocks and circular dependencies can be reduced. Furthermore, by calculating reduction risk parameters based on deletion ratio, index ratio, and basic block concentration, and adjusting the number of erasure code verification symbols accordingly, while applying a protection strength to the reconstruction witness index area no less than that of the reserved data area, the recovery dependencies after data reduction can obtain corresponding storage protection, thereby reducing the physical write volume of some duplicate or differentially reconfigurable data blocks while ensuring lossless recovery.

[0292] Figure 6 This is a schematic diagram of the data reduction processing apparatus provided in an embodiment of this application. Figure 6As shown in the embodiments of this application, an embodiment of the present application also provides a data reduction processing device, which includes: an acquisition module 601, a determination module 602, a first determination module 603, a verification module 604, a calculation module 605, a second determination module 606, and a storage module 607.

[0293] The acquisition module is used to acquire a set of data blocks; the set of data blocks includes multiple data blocks.

[0294] The decision module is used to determine the dependencies between data blocks.

[0295] The first determining module is used to determine at least one reserved data block and at least one data block to be reduced from the data block set according to the dependency relationship; wherein the data block to be reduced can be recovered based on the reserved data block;

[0296] The verification module is used to perform recovery verification on the data block to be reduced, and generate reconstruction information for restoring the data block to be reduced after the verification is passed;

[0297] The calculation module is used to calculate the reduction risk parameters corresponding to the data block to be reduced;

[0298] The second determination module is used to determine the corresponding redundancy protection strength based on the reduced risk parameters;

[0299] The storage module is used to store reserved data blocks and reconstruction information according to the redundancy protection strength.

[0300] In one possible embodiment, the determination module is specifically used for:

[0301] Retrieve record information for data blocks;

[0302] Based on the recorded information, a complete duplicate dependency determination is performed on the data blocks to obtain the complete duplicate dependency relationship between the data blocks; wherein, the complete duplicate dependency relationship includes the first base data block as the dependent party, and the first data block to be reduced as the dependent party.

[0303] Based on the recorded information, differential reconfigurable dependency determination is performed on the data blocks to obtain the differential reconfigurable dependency relationship between the data blocks; wherein, the differential reconfigurable dependency relationship includes the second basic data block as the dependent party, and the second data block to be reduced as the dependent party.

[0304] In one possible embodiment, the recorded information includes a content hash value, data length, and a fast checksum; the determination module is specifically used for:

[0305] Based on the recorded information, perform a complete duplicate dependency determination on the data blocks to obtain the complete duplicate dependencies between the data blocks, including:

[0306] Based on the content hash value of the data block, search for candidate data blocks from the data block set;

[0307] If the candidate data block and the data block have the same content hash value, the same data length, and the same fast check code, it is determined that there is a complete duplicate dependency relationship between the data block and the candidate data block; where the candidate data block is the first basic data block and the data block is the first data block to be reduced.

[0308] In one possible embodiment, the recorded information further includes a rolling fingerprint set; the determination module is specifically used for:

[0309] Based on the rolling fingerprint set of data blocks, recall candidate basic data blocks from the data block set;

[0310] Generate differential patches for the data blocks relative to the candidate base data blocks;

[0311] Recovery verification is performed based on candidate basic data blocks and differential patches. If the data length, content hash value and fast check code of the recovered data block are consistent with those of the data block, it is determined that there is a differential reconstructable dependency relationship between the data block and the candidate basic data block. Among them, the candidate basic data block is the second basic data block, and the data block is the second data block to be reduced.

[0312] In one possible embodiment, the first determining module is specifically used for:

[0313] Based on the dependencies, a candidate dependency table is created, which is used to represent the dependencies of data blocks;

[0314] Based on the candidate dependency table, the status identifier of the data block is determined; the initial status of the status identifier is configured as pending block, and the status identifier also includes core reserved block, ordinary reserved block, candidate reduced block and reduced block;

[0315] Based on the status identifier, at least one reserved data block and at least one data block to be reduced are determined from the data block set.

[0316] In one possible embodiment, the first determining module is further specifically used for:

[0317] If the status of a data block is identified as a candidate reduction block or a reduced block, then the data block is determined to be a data block to be reduced.

[0318] If the status of a data block is identified as a core reserved block or a regular reserved block, then the data block is determined to be a reserved data block.

[0319] In one possible embodiment, the first determining module is further specifically used for:

[0320] Based on the candidate dependency table, calculate the number of times a data block is referenced as a base data block in the dependency; where the base data block includes the first base data block and the second base data block;

[0321] The core reserved blocks are determined from the data blocks based on the number of references.

[0322] Data blocks that are not identified as core reserved blocks and have no dependencies are identified as ordinary reserved blocks;

[0323] Data blocks that were not identified as core reserved blocks but have dependencies are identified as candidate reduction blocks;

[0324] The data blocks that pass the recovery verification in the candidate reduction blocks are identified as the reduced blocks.

[0325] In one possible embodiment, the computing module is specifically used for:

[0326] The data block set is divided into multiple processing partitions;

[0327] Obtain reduction information for the data blocks to be reduced within the processing partition; wherein, the reduction information includes at least one of the following: the proportion of data already reduced, the proportion of reconstructed information, and the concentration of the base blocks;

[0328] Based on the reduction information, calculate the reduction risk parameters corresponding to the data blocks to be reduced within the processing partition.

[0329] In one possible embodiment, the computing module is further specifically used for:

[0330] Obtain the ownership information of the data block; wherein, the ownership information includes at least one of the file identifier, object identifier, or time window identifier to which the data block belongs;

[0331] By using attribution information as a constraint, the data block set is divided into multiple processing partitions.

[0332] In one possible embodiment, the computing module is further specifically used for:

[0333] Obtain the data block distribution information of the processing partition corresponding to the data block to be reduced; wherein, the data block distribution information includes the proportion of the reduced data volume, the proportion of the reconstructed information, and the basic block concentration, and the basic block concentration is used to characterize the degree to which multiple data blocks to be reduced depend on the same basic data block;

[0334] Based on the data block distribution information, calculate the reduction risk parameters corresponding to the data blocks to be reduced.

[0335] In one possible embodiment, the storage module is specifically used for:

[0336] The number of check symbols in the erasure coding is determined based on the risk reduction parameter; the number of check symbols is positively correlated with the risk reduction parameter.

[0337] Based on the number of check symbols, erasure coding is performed on the retained data blocks and reconstructed information to generate check symbols;

[0338] The retained data blocks, reconstruction information, and verification symbols are written to the storage medium.

[0339] In one possible embodiment, the verification module is specifically used for:

[0340] Generate directed dependency edges based on the data block to be reduced and the corresponding base data block;

[0341] Perform dependency graph anti-loop verification based on directed dependency edges;

[0342] If the loop prevention check fails, then the recovery check fails.

[0343] In one possible embodiment, the verification module is specifically used for:

[0344] Obtain the base data block corresponding to the data block to be reduced;

[0345] Perform basic data availability verification on the basic data blocks;

[0346] If the basic data availability verification fails, then the recovery verification is deemed to have failed.

[0347] In one possible embodiment, the verification module is specifically used for:

[0348] Obtain the reconstruction information corresponding to the data block to be reduced;

[0349] Based on the reconstruction information, perform a reconstruction correctness check on the data block to be reduced;

[0350] If the reconstruction correctness check passes, the physical storage of the data block to be reduced can be released or the physical write task of the data block to be reduced can be canceled.

[0351] If the reconstruction correctness check fails, the data block to be reduced will be converted into a retained data block.

[0352] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 7 As shown, the electronic device 70 provided in this embodiment includes at least one processor 701 and a memory 702. In one possible implementation, the electronic device 70 further includes a communication component 703. The processor 701, memory 702, and communication component 703 are connected via a bus.

[0353] In a specific implementation, at least one processor 701 executes computer execution instructions stored in memory 702, causing at least one processor 701 to execute the above-described data reduction processing method embodiment.

[0354] The specific implementation process of processor 701 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.

[0355] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.

[0356] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.

[0357] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings of this application's embodiments are not limited to only one bus or one type of bus.

[0358] An embodiment of this application also provides a computer-readable storage medium storing a computer program configured to execute the steps in any of the above-described data reduction processing method embodiments at runtime.

[0359] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0360] An embodiment of this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described data reduction processing method embodiments.

[0361] The embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described data reduction processing method embodiments.

[0362] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of the embodiments of this application.

[0363] The data reduction processing method and electronic device provided in the embodiments of this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of the embodiments of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the methods and core ideas of the embodiments of this application. It should be noted that those skilled in the art can make several improvements and modifications to the embodiments of this application without departing from the principles of the embodiments of this application, and these improvements and modifications also fall within the protection scope of the embodiments of this application.

Claims

1. A data reduction processing method, characterized in that, include: Obtain a set of data blocks; wherein the set of data blocks includes multiple data blocks; Dependency determination is performed on the data blocks to obtain the dependencies between the data blocks; Based on the dependencies, at least one reserved data block and at least one data block to be reduced are determined from the set of data blocks; wherein the data block to be reduced can be recovered based on the reserved data block; The data block to be reduced is restored and verified, and reconstruction information for restoring the data block to be reduced is generated after the verification is passed; Calculate the reduction risk parameters corresponding to the data block to be reduced; Based on the risk reduction parameters, determine the corresponding redundancy protection strength; Based on the redundancy protection strength, store the reserved data block and the reconstruction information.

2. The method according to claim 1, characterized in that, The step of determining the dependencies between the data blocks includes: Obtain the record information of the data block; Based on the recorded information, a complete duplicate dependency determination is performed on the data blocks to obtain the complete duplicate dependency relationship between the data blocks; wherein, the complete duplicate dependency relationship includes a first base data block as the dependent party and a first data block to be reduced as the dependent party. Based on the recorded information, differential reconfigurable dependency determination is performed on the data blocks to obtain the differential reconfigurable dependency relationship between the data blocks; wherein, the differential reconfigurable dependency relationship includes a second basic data block as the dependent party and a second data block to be reduced as the dependent party.

3. The method according to claim 2, characterized in that, The recorded information includes the content hash value, data length, and fast checksum. The step of determining complete duplicate dependencies between data blocks based on the recorded information includes: Based on the content hash value of the data block, search for candidate data blocks from the data block set; If the candidate data block found has the same content hash value, the same data length, and the same fast check code as the data block, then it is determined that there is a complete duplicate dependency relationship between the data block and the candidate data block; wherein, the candidate data block is the first basic data block, and the data block is the first data block to be reduced.

4. The method according to claim 3, characterized in that, The recorded information also includes a rolling fingerprint set; The step of determining differentially reconfigurable dependencies between data blocks based on the recorded information includes: Based on the rolling fingerprint set of the data blocks, candidate basic data blocks are recalled from the data block set; Generate a differential patch for the data block relative to the candidate base data block; Recovery verification is performed based on the candidate basic data block and the differential patch. If the recovered data block is consistent with the data length, content hash value and fast check code of the data block, it is determined that there is a differential reconfigurable dependency relationship between the data block and the candidate basic data block; wherein, the candidate basic data block is the second basic data block and the data block is the second data block to be reduced.

5. The method according to claim 2, characterized in that, The step of determining at least one reserved data block and at least one data block to be reduced from the data block set based on the dependency relationship includes: Based on the dependencies, a candidate dependency table is established, which is used to characterize the dependencies of the data blocks; Based on the candidate dependency table, the status identifier of the data block is determined; wherein, the initial state of the status identifier is configured as a pending block, and the status identifier also includes core reserved block, ordinary reserved block, candidate reduced block, and reduced block; Based on the status identifier, at least one reserved data block and at least one data block to be reduced are determined from the data block set.

6. The method according to claim 5, characterized in that, The step of determining at least one reserved data block and at least one data block to be reduced from the data block set based on the status identifier includes: If the status identifier of the data block is the candidate reduction block or the already reduced block, then the data block is determined to be the data block to be reduced; If the status identifier of the data block is the core reserved block or the ordinary reserved block, then the data block is determined to be the reserved data block.

7. The method according to claim 5, characterized in that, Determining the status identifier of the data block based on the candidate dependency table includes: Based on the candidate dependency table, calculate the number of times the data block is referenced as a base data block in the dependency; wherein the base data block includes the first base data block and the second base data block; The core reserved block is determined from the data block based on the number of references. Data blocks that are not identified as core reserved blocks and do not have the aforementioned dependencies are identified as ordinary reserved blocks; Data blocks that were not identified as core retained blocks but have the aforementioned dependencies are identified as candidate reduction blocks; The data block that passes the recovery verification in the candidate reduction block is determined as the reduced block.

8. The method according to any one of claims 1 to 7, characterized in that, The calculation of the reduction risk parameters corresponding to the data block to be reduced includes: The data block set is divided into multiple processing partitions; Obtain the reduction information of the data block to be reduced within the processing partition; wherein the reduction information includes at least one of the following: the proportion of data already reduced, the proportion of reconstructed information, and the concentration of basic blocks; Based on the reduction information, calculate the reduction risk parameters corresponding to the data blocks to be reduced within the processing partition.

9. The method according to claim 8, characterized in that, The process of dividing the data block set to obtain multiple processing partitions includes: Obtain the ownership information of the data block; wherein, the ownership information includes at least one of the file identifier, object identifier, or time window identifier to which the data block belongs; Using the attribution information as a constraint, the data block set is divided into multiple processing partitions.

10. The method according to claim 9, characterized in that, The calculation of the reduction risk parameters corresponding to the data block to be reduced includes: Obtain the data block distribution information of the processing partition corresponding to the data block to be reduced; wherein, the data block distribution information includes the proportion of reduced data volume, the proportion of reconstructed information and the basic block concentration, and the basic block concentration is used to characterize the degree to which multiple data blocks to be reduced depend on the same basic data block; Based on the data block distribution information, calculate the reduction risk parameter corresponding to the data block to be reduced.

11. The method according to any one of claims 1 to 7, characterized in that, The step of storing the reserved data block and the reconstruction information according to the redundancy protection strength includes: The number of check symbols in the erasure coding is determined based on the risk reduction parameter; wherein the number of check symbols is positively correlated with the risk reduction parameter. Based on the number of check symbols, erasure coding is performed on the retained data block and the reconstructed information to generate check symbols; The reserved data block, the reconstruction information, and the verification symbol are written into the storage medium.

12. The method according to any one of claims 1 to 7, characterized in that, The recovery verification of the data block to be reduced includes: Based on the data block to be reduced and the corresponding basic data block, generate directed dependency edges; Based on the directed dependency edges, perform dependency graph anti-loop verification; If the dependency graph loop prevention check fails, then the recovery check fails.

13. The method according to any one of claims 1 to 7, characterized in that, The recovery verification of the data block to be reduced includes: Obtain the base data block corresponding to the data block to be reduced; Perform a basic data availability check on the aforementioned basic data block; If the basic data availability verification fails, then the recovery verification is determined to have failed.

14. The method according to any one of claims 1 to 7, characterized in that, The recovery verification of the data block to be reduced includes: Obtain the reconstruction information corresponding to the data block to be reduced; Based on the reconstruction information, a reconstruction correctness check is performed on the data block to be reduced; If the reconstruction correctness check passes, the physical storage of the data block to be reduced is allowed to be released or the physical write task of the data block to be reduced is cancelled. If the reconstruction correctness check fails, the data block to be reduced will be converted into a retained data block.

15. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the data reduction processing method as described in any one of claims 1 to 14.