Data reconstruction method, device and equipment and readable storage medium

By dynamically switching storage units in parallel to perform data reconstruction and writing operations in a distributed storage system, the performance bottlenecks and consistency risks caused by traditional data reconstruction mechanisms are solved, and more efficient data writing and system performance improvements are achieved.

CN120492199APending Publication Date: 2025-08-15XINHUASAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510569981.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

When traditional data reconstruction mechanism faces concurrent execution of client I/O and reconstruction tasks, there are significant performance bottlenecks, especially in write operation scenarios, resulting in increased latency and risk of data consistency.

Method used

By dynamically switching the storage unit, the target storage unit that is not in the reconstructed state is selected as the new write object, the parallel execution of data reconstruction and client write operations is realized, the storage space is processed using a garbage collection strategy, and the processing flow of overwriting and approximating writes is optimized according to the write type.

Benefits of technology

It effectively avoids the increase in delay caused by waiting for reconstruction, avoids the risk of data consistency, and improves the overall performance and load balancing of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492199A_ABST
    Figure CN120492199A_ABST
Patent Text Reader

Abstract

The invention provides a data reconstruction method, device and equipment and a readable storage medium, and the method comprises the steps: responding to a data writing request, and obtaining the running state of a specific storage unit, the specific storage unit being a writing object associated with the data writing request; in response to the condition that the running state of the specific storage unit is a data reconstruction state, selecting a target storage unit of which the running state is not in a data reconstruction state as a new write-in object from the placement group to which the storage unit belongs; and writing data associated with the data writing request into the target storage unit. According to the technical scheme, parallel execution of data reconstruction and client write-in operation is achieved by dynamically switching the storage units, and when the specific storage unit is in the reconstruction state, the target storage unit which is not reconstructed in the same placement group is actively selected as a new write-in object, so that time delay increase caused by waiting for reconstruction is avoided, and the storage efficiency is improved. And the data consistency risk is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of communication technology, and in particular to a data reconstruction method, apparatus, device, and readable storage medium. Background Art

[0002] Typically, distributed storage systems implement multi-copy backup or erasure coding (EC) protection of data by dividing the data into multiple logical units (such as PlacementGroup, placement group PG) and distributing them to different storage nodes (Object Storage Device, OSD). When a hard disk fails or the system is expanded or reduced in capacity, the missing data needs to be redistributed to the target OSD through a data reconstruction process to maintain the preset redundancy and data balance. This process not only involves cross-node data transmission and calculation (such as check block reconstruction in EC scenarios), but also needs to ensure data consistency and compatibility with client business requests during the reconstruction process. However, traditional data reconstruction mechanisms often have significant performance bottlenecks when facing concurrent execution of client I / O requests and reconstruction tasks, especially in write operation scenarios, where their limitations are particularly prominent.

[0003] To address conflicts between client I / O and reconstruction tasks, a strategy of "waiting for reconstruction to complete before processing I / O" is often adopted. For example, in a three-replica architecture, when an OSD fails, the remaining two healthy replicas are elected as the authoritative data source. Client read operations can directly obtain data from the authoritative replicas, unaffected by the reconstruction progress. However, for write operations, the system must ensure that newly written data is consistent across all replicas. Therefore, the reconstruction task must complete and the data from the missing replicas must be complete before the write operation can be performed. Similarly, in an EC scenario, if a data shard or parity shard is lost, the system must first use the EC algorithm to restore the complete data before allowing the write operation. While this mechanism can strictly guarantee data consistency, the reconstruction task itself often involves large-scale data transmission and computation (such as processing data blocks larger than 1MB) and requires coordination across network nodes, which can result in reconstruction latency of seconds. For latency-sensitive client services (such as high-frequency trading systems or real-time databases), such delays can trigger a large number of retry requests, significantly reducing overall system performance (such as IOPS and throughput). Summary of the Invention

[0004] In view of this, the present disclosure provides a data reconstruction method, device, electronic device, and readable storage medium to improve the problem of write delay caused by the above-mentioned data reconstruction.

[0005] The specific technical solutions are as follows:

[0006] This specification provides a data reconstruction method, applied to a distributed storage device, the method comprising: in response to a data write request, obtaining the operating status of a specific storage unit, the specific storage unit being a write target associated with the data write request; in response to the operating status of the specific storage unit being in a data reconstruction state, selecting a target storage unit in a placement group to which the storage unit belongs whose operating status is not in a data reconstruction state as a new write target; and writing data associated with the data write request to the target storage unit.

[0007] As a technical solution, the step of writing data associated with the data write request to the target storage unit includes: using a garbage collection strategy to process storage space of a specific storage unit.

[0008] As a technical solution, after the step of responding to the specific storage unit's operating state being in a data reconstruction state, the method further includes: obtaining a write type associated with a data write request; locating a data cache area of a memory corresponding to the specific storage unit in the data reconstruction process based on the result of obtaining that the write type is an overwrite write, overwriting the data associated with the data write request into the data cache area, and then continuing to perform data reconstruction associated with the specific storage unit; and stopping the step of selecting a target storage unit in the placement group to which the storage unit belongs that is not in a data reconstruction state as a new write target, and writing the data associated with the data write request to the target storage unit.

[0009] As a technical solution, in response to the specific storage unit's operating state being in a data reconstruction state, selecting a target storage unit in the placement group to which the storage unit belongs whose operating state is not in a data reconstruction state as a new write target includes: in response to the specific storage unit's operating state being in a data reconstruction state, obtaining a write type associated with a data write request; and based on the obtained result that the write type is an append write, selecting a target storage unit in the placement group to which the storage unit belongs whose operating state is not in a data reconstruction state as the new write target.

[0010] This specification also provides a data reconstruction device for use in a distributed storage device. The device includes: a first module for obtaining the operating status of a specific storage unit in response to a data write request, where the specific storage unit is a write target associated with the data write request; a second module for selecting a target storage unit in the placement group to which the storage unit belongs whose operating status is not in the data reconstruction state as a new write target in response to the operating status of the specific storage unit being in the data reconstruction state; and a third module for writing data associated with the data write request to the target storage unit.

[0011] As a technical solution, the step of writing data associated with the data write request to the target storage unit includes: using a garbage collection strategy to process storage space of a specific storage unit.

[0012] As a technical solution, after the second module responds to the specific storage unit's operating state being in a data reconstruction state, it further includes: obtaining a write type associated with a data write request; based on the acquisition result that the write type is an overwrite write, locating a data cache area corresponding to the memory of the specific storage unit in the data reconstruction, overwriting the data associated with the data write request into the data cache area, and then continuing to execute the data reconstruction associated with the specific storage unit; stopping calling the second module and the third module to execute the above, selecting a target storage unit in the placement group to which the storage unit belongs that is not in a data reconstruction state as a new write object, and writing the data associated with the data write request to the target storage unit.

[0013] As a technical solution, in response to the specific storage unit's operating state being in a data reconstruction state, selecting a target storage unit in the placement group to which the storage unit belongs whose operating state is not in a data reconstruction state as a new write target includes: in response to the specific storage unit's operating state being in a data reconstruction state, obtaining a write type associated with a data write request; and based on the obtained result that the write type is an append write, selecting a target storage unit in the placement group to which the storage unit belongs whose operating state is not in a data reconstruction state as the new write target.

[0014] This specification also provides an electronic device, including a processor and a readable storage medium, wherein the readable storage medium stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the aforementioned data reconstruction method.

[0015] This specification also provides a readable storage medium, which stores machine-executable instructions. When the machine-executable instructions are called and executed by a processor, the machine-executable instructions prompt the processor to implement the aforementioned data reconstruction method.

[0016] The above technical solutions provided in this specification bring at least the following beneficial effects:

[0017] By dynamically switching storage units, data reconstruction and client write operations can be executed in parallel. When a specific storage unit is in the reconstruction state, the unreconstructed target storage unit in the same placement group is actively selected as the new write object, avoiding the increase in latency caused by waiting for reconstruction and avoiding data consistency risks. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the implementation methods of this specification or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the implementation methods of this specification or the description of the prior art. Obviously, the drawings described below are only some implementation methods recorded in this specification. For ordinary technicians in this field, other drawings can also be obtained based on these drawings of the implementation methods of this specification.

[0019] Figure 1 is a flow chart of a data reconstruction method in one embodiment of this specification;

[0020] Figure 2 This is a schematic diagram of an architecture in one embodiment of this specification;

[0021] Figure 3 This is a schematic diagram of an architecture in one embodiment of this specification;

[0022] Figure 4 This is a structural diagram of a data reconstruction device in one embodiment of this specification;

[0023] Figure 5 This is a hardware structure diagram of an electronic device in one embodiment of this specification.

[0024] Reference numerals: first module 21 , second module 22 , third module 23 . DETAILED DESCRIPTION

[0025] The terms used in the embodiments of this specification are only for the purpose of describing specific embodiments and are not intended to limit this specification. The singular forms "a," "the," and "the" used in this specification and claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to any or all possible combinations of one or more of the associated listed items.

[0026] It should be understood that although the terms first, second, third, etc. may be used in the embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of this specification, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" may also be interpreted as "when...", "when...", or "in response to determining."

[0027] This specification provides a data reconstruction method, device, electronic device, and readable storage medium to at least improve one of the above technical problems.

[0028] The specific technical solution is described below.

[0029] In one embodiment, this specification provides a data reconstruction method applied to a distributed storage device, the method comprising: in response to a data write request, obtaining the operating status of a specific storage unit, the specific storage unit being a write target associated with the data write request; in response to the operating status of the specific storage unit being a data reconstruction state, selecting a target storage unit in a placement group to which the storage unit belongs whose operating status is not in data reconstruction as a new write target; and writing data associated with the data write request to the target storage unit.

[0030] Specifically, if Figure 1 , including the following steps:

[0031] Step S11 : in response to a data write request, obtaining an operating status of a specific storage unit, where the specific storage unit is a write target associated with the data write request.

[0032] When a data write request arrives at a distributed storage system, the system first needs to determine the specific location where the data should be written, namely the specific storage unit. This specific storage unit is closely associated with the write request and is the intended destination for the data. In the traditional data write process, if the specific storage unit is in the data reconstruction state, the client's write operation is forced to wait until the reconstruction process is complete. This waiting mechanism often becomes a bottleneck for system performance when faced with large-scale data storage and frequent reconstruction operations.

[0033] Step S12 , in response to the operating state of the specific storage unit being in the data reconstruction state, a target storage unit in the placement group to which the storage unit belongs and whose operating state is not in the data reconstruction state is selected as a new write target.

[0034] When a specific storage unit is detected as being in a data reconstruction state, the system does not stall the write operation. Instead, it actively searches for other normally functioning storage units within the placement group (PG) to which the storage unit belongs as alternative write targets. A placement group, as the logical data sharding unit in a distributed storage system, manages the distribution of objects across multiple storage devices (OSDs). By dynamically reallocating objects within the same placement group, the system enables rapid data writes without compromising data consistency and integrity.

[0035] Step S13: writing data associated with the data write request into the target storage unit.

[0036] For example, in a three-replica distributed storage system, suppose a client initiates a write request for object OBJ1.1.1. OBJ1.1.1 is originally distributed across three different OSDs, forming three replicas. If one of these OSDs fails, triggering data reconstruction, the system initiates peering to select an authoritative data source. Typically, one of the two remaining healthy OSDs is selected as the authoritative replica.

[0037] When it detects that OSD1, where OBJ1.1.1 resides, is in the rebuilding state, it searches for other functioning storage units in the placement group to which OBJ1.1.1 belongs. Assume that, in addition to OSD1, the placement group also includes OSD2 and OSD3, and that they are not currently participating in the rebuild process. The system selects one of the non-rebuild OSDs as the new write target, for example, OSD4. At this point, the client's write request is redirected to OBJ4.1.1 on OSD4, enabling faster data writes. This process eliminates the need to wait for OSD1's rebuild to complete, significantly reducing write latency.

[0038] In the system architecture, each placement group typically contains multiple storage units, providing ample room for dynamic object selection. When a specific storage unit cannot immediately respond to a write request due to reconstruction, the system uses an internal scheduling algorithm to quickly locate a suitable replacement within the placement group. This algorithm comprehensively considers factors such as the current load of the storage unit, the balance of data distribution, and the efficiency of network communication to ensure that the selected target storage unit can complete the data write task with optimal performance.

[0039] In a large-scale distributed storage cluster, a placement group may contain dozens or even more storage units. When a storage unit initiates a reconstruction due to a hard drive failure, the system can select a unit with lower load and lower network latency from the remaining units as the new write target. This approach not only avoids client I / O waiting but also achieves load balancing for write operations, further improving overall system performance.

[0040] In one embodiment, Figure 2, for the optimization of the append write scenario, the system determines whether it is in the reconstruction stage by detecting the status of the target storage unit. When a specific storage unit (such as OBJ1.1.1 in PG1) triggers reconstruction due to a failure, if a write request to the unit is received from the client at this time, the system will actively search for target units that are not involved in the reconstruction (such as OBJ2.1.1) in other storage units belonging to PG1. This process relies on the preset rules of the CRUSH algorithm for the distribution of storage units within PG to ensure that the selected target unit meets both redundancy requirements (such as the number of replicas or EC slices) and availability (not marked as failed or under reconstruction).

[0041] In a three-replica architecture, suppose OSD3, where OBJ1.1.1 resides, fails. The system needs to recover data from healthy replicas (such as OSD1 and OSD2) to a new node, OSD4. If a client attempts to write data to OBJ1.1.1, the system will detect that its status is "reconstructing." It will then select the unreconstructed OBJ2.1.1 within PG1 as a replacement and redirect the client request to that target unit. Because the append write feature allows new data to be written to a new location on the storage medium without modifying the original data, this operation does not compromise data consistency.

[0042] The client originally planned to write 8KB of data to location B of OBJ1.1.1. The system then changed the plan to write the same content to location B of OBJ2.1.1. The GC mechanism then reclaimed the free space in location B of OBJ1.1.1. This dynamic switching mechanism effectively circumvents the synchronization limitation of traditional solutions that require "waiting for reconstruction to complete before writing," allowing reconstruction tasks to execute in parallel with client I / O.

[0043] To optimize overwrite scenarios, a prioritized reconstruction strategy is required to resolve conflicts caused by directly replacing old data locations. When the target storage unit is in the reconstruction state and the client requests an overwrite, the system must first read the authoritative data source (such as the healthy data shard and check shard in the EC scenario), merge the overwrite content into the data required for reconstruction, and then complete the reconstruction and write operations simultaneously. For example, in an EC4+2 configuration, if OSD3, where OBJ1.1.1 (data shard) is located, fails, the system must extract the remaining data shards and check shards from OSD1-OSD4 to calculate the lost data shards. At this time, if the client attempts to overwrite location A of OBJ1.1.1, the system will first read the complete data of OSD1-OSD4, update the overwrite content to the corresponding location, and then regenerate the complete six shards. The updated data shards and check shards are simultaneously written to the corresponding objects of OSD3 (the target OSD), OSD1, and OSD2 (such as OBJ1.1.1, OBJ2.1.1, and OBJ3.1.1). Although this process involves a full data reorganization (such as amplifying an 8KB overwrite write to a 1MB object write), the reconstruction and business request can be completed through a single network transmission and disk write, avoiding the two independent operations required for "secondary writing after reconstruction" in traditional solutions (such as first reconstructing 1MB of data and then appending 8KB of overwrite data), thereby significantly reducing latency fluctuations.

[0044] In one embodiment, writing data associated with the data write request to the target storage unit includes: processing storage space of a specific storage unit using a garbage collection strategy.

[0045] In one embodiment, after the step of responding to the operating state of the specific storage unit being a data reconstruction state, the method further includes: obtaining a write type associated with the data write request; locating a data cache area of the memory corresponding to the specific storage unit in the data reconstruction based on the result of obtaining that the write type is an overwrite write, overwriting the data associated with the data write request into the data cache area, and then continuing to execute the data reconstruction associated with the specific storage unit; and stopping the step of selecting a target storage unit in the placement group to which the storage unit belongs that is not in a data reconstruction state as a new write object, and writing the data associated with the data write request to the target storage unit.

[0046] When an object is detected as being reconstructed and an overwrite write request is received, the system prioritizes reconstruction. For example, in an EC4+2 erasure coding scenario, suppose the OSD hosting a shard of OBJ1.1.1 fails. The system initiates a reconstruction process to recover the lost data from the remaining healthy shards and parity shards. At this point, if a client initiates an overwrite write request for OBJ1.1.1, the system will not wait for the reconstruction to complete.

[0047] Specifically, the complete object data is first read from the authoritative data source into memory. The overwritten data is then integrated into the in-memory object data to form a new data version. This updated data version is then simultaneously written to all relevant storage units, including the target OSD being rebuilt and all other healthy OSDs. This approach allows data reconstruction and overwriting to be performed simultaneously in a single write operation, avoiding the cumulative latency associated with two writes in traditional methods.

[0048] In one embodiment, in response to the operating state of the specific storage unit being in a data reconstruction state, selecting a target storage unit in the placement group to which the storage unit belongs whose operating state is not in data reconstruction as a new write object includes: in response to the operating state of the specific storage unit being in a data reconstruction state, obtaining a write type associated with a data write request; and based on the obtained result that the write type is an append write, selecting a target storage unit in the placement group to which the storage unit belongs whose operating state is not in data reconstruction as the new write object.

[0049] In one embodiment, when reconstruction and client IO access the same object of a PG at the same time, the object can be replaced by writing, which not only avoids the performance impact on the client, but also greatly reduces the complexity caused by reconstruction and IO mutual exclusion.

[0050] Taking the replica method as an example, in practice, object OBJ1.1.1 typically has three states: waiting for reconstruction, being reconstructed, and reconstructed. When OBJ1.1.1 is in state 3 (reconstructed), its data is complete and requires no additional processing; write operations can proceed normally. For the waiting for reconstruction and reconstructing states, the traditional approach requires client I / O to wait until the reconstruction is complete before writing.

[0051] Before reconstruction, client I / O is normally writing to location A of object OBJ1.1.1. However, a failure occurs, requiring reconstruction of OBJ1.1.1. After peering, the system selects OSD3 as the authority. OBJ1.1.1 then enters the Pending Reconstruction or Reconstructing state, with the reconstruction data source being OBJ3.1.1 on OSD3. During reconstruction, the OBJ3.1.1 data is transferred to OSD1 and written to OBJ1.1.1. The client, as originally planned, is supposed to write to location B of OBJ1.1.1. At this point, another client I / O write request arrives. The system detects that OBJ1.1.1 is being reconstructed and returns the request to the upper layer, notifying it to write to a different object. The upper layer retries and selects an object with complete redundancy. For example, if OBJ2.1.1 is selected, the object is written normally and returns a normal response. Because OBJ2.1.1 is a normal object, the overall write latency is normal and unaffected by the reconstruction. Similarly, the reconstruction is unaffected by client I / O. Data is copied from OBJ3.1.1 to OBJ1.1.1 normally. When OBJ1.1.1 receives client I / O again, the upper layer has already recorded that OBJ1.1.1 is undergoing reconstruction, so subsequent I / O is directly sent to OBJ2.1.1 for writing without a second attempt. Since the B space of OBJ1.1.1 has not been written to, this space may be temporarily wasted. The GC scans the object in the background and discovers that there is a large amount of free space, so it reclaims and moves the space.

[0052] In one embodiment, Figure 3 When OBJ1.1.1 is in the state of having been reconstructed, it means that its data is complete and does not require additional processing. The normal overwrite operation can be performed.

[0053] For scenarios where data is pending or being reconstructed, let's take the example of data originally planned to be written to region A of OBJ1.1.1. While OBJ1.1.1 is still in the pending or reconstructing states, an overwrite occurs to region A of OBJ1.1.1. Since the system has peered and selected the authoritative data source, OBJ3.1.1, reconstruction is performed using OBJ3.1.1 as the data source. At this point, the entire object data of OBJ3.1.1 is read into memory. The contents of the overwritten data in A are written to the data read into memory in OBJ3.1.1, resulting in data E. This data is the data that currently needs to be reconstructed and has already been overwritten by data in A. Data E is then sent over the network to the PG objects OBJ3.1.1, OBJ1.1.1, and OBJ2.1.1 on each OSD where three replicas are required. Upon successful write, a notification is sent to the host confirming the successful overwrite of data A. Since the data in OBJ1.1.1 has been reconstructed and the overwrite of data in A has been completed, the object is marked as reconstructed, and subsequent data can be written normally.

[0054] In one embodiment, Figure 4 This specification also provides a data reconstruction device, which is applied to a distributed storage device. The device includes: a first module, which is used to obtain the operating status of a specific storage unit in response to a data write request, where the specific storage unit is a write target associated with the data write request; a second module, which is used to select a target storage unit in the placement group to which the storage unit belongs, whose operating status is not in the data reconstruction state, as a new write target in response to the operating status of the specific storage unit being in the data reconstruction state; and a third module, which is used to write data associated with the data write request to the target storage unit.

[0055] In one embodiment, writing data associated with the data write request to the target storage unit includes: processing storage space of a specific storage unit using a garbage collection strategy.

[0056] In one embodiment, after the operating state of the specific storage unit is in a data reconstruction state, the second module further includes: obtaining a write type associated with the data write request; based on the obtained result that the write type is an overwrite write, locating a data cache area of the memory corresponding to the specific storage unit in the data reconstruction, overwriting the data associated with the data write request into the data cache area, and then continuing to execute the data reconstruction associated with the specific storage unit; stopping calling the second module and the third module to execute the steps of selecting a target storage unit in the placement group to which the storage unit belongs that is not in a data reconstruction state as a new write target, and writing the data associated with the data write request to the target storage unit.

[0057] In one embodiment, in response to the operating state of the specific storage unit being in a data reconstruction state, selecting a target storage unit in the placement group to which the storage unit belongs whose operating state is not in data reconstruction as a new write object includes: in response to the operating state of the specific storage unit being in a data reconstruction state, obtaining a write type associated with a data write request; and based on the obtained result that the write type is an append write, selecting a target storage unit in the placement group to which the storage unit belongs whose operating state is not in data reconstruction as the new write object.

[0058] In one embodiment, it is first necessary to establish a global state perception mechanism to track the operating status of each storage unit (such as OSD) in real time through the metadata cluster, including health status, reconstruction progress and load indicators. The metadata service uses a distributed consistency protocol (such as Raft) to ensure the strong consistency of the status information of each node. When the client initiates a data write request, the request is routed to the primary OSD of the placement group PG to which the target object belongs. The primary OSD initiates a query to the metadata service to obtain the detailed status mark of the target storage unit. For example, when the client tries to write the object Object-A in PG-100, the primary OSD-2 reads from the metadata service that the physical storage unit OSD-1 currently mapped to Object-A is in the "reconstructing" state, and its reconstruction progress is 45%, with a remaining data transmission volume of approximately 600MB. At this time, the system will trigger the write path redirection logic instead of blocking the client request to avoid the surge in latency caused by waiting for the reconstruction to complete in the traditional solution.

[0059] The dynamic selection process for the target storage unit relies on a multi-dimensional weighted election algorithm that comprehensively considers factors such as real-time load, network topology, and historical stability. For example, consider PG-100 with a three-replication policy. Assume its members include OSD-2 (primary replica, load score 30), OSD-3 (load score 50), and OSD-4 (load score 75). The load score is calculated using a 3:4:3 weighting of CPU utilization, disk queue depth, and network bandwidth utilization. The system prioritizes OSD-3, which has the lowest load score, as the target storage unit. However, the actual selection process also requires topology affinity optimization: if OSD-3 and the client are located in the same rack with a network latency of 0.1ms, while OSD-4 has a cross-rack latency of 0.5ms, then even if OSD-3 has a slightly higher load (e.g., a score of 55), it may still be selected due to its more optimal network location, thus reducing end-to-end write latency. Furthermore, historical failure rates are factored into the election strategy. For example, if OSD-4 experiences two timeout errors within the past hour, the system will impose a penalty weight on it, increasing its overall score by 20%, thereby reducing its probability of selection. These dynamic adjustments are made through real-time monitoring of data streams, ensuring that the election results meet both immediate performance requirements and long-term system stability.

[0060] After selecting the target storage unit, the system creates a new object (such as Object-A') on the target OSD and writes data. This process requires synchronously updating metadata to maintain the mapping relationship between logical objects and physical objects. For example, a client write request points to the logical object " / bucket / file1", which was originally mapped to Object-A@OSD-1 and is updated to Object-A'@OSD-3 after redirection. Metadata updates are ensured atomically through a two-phase commit protocol: the primary OSD-2 initiates a pre-commit request to the metadata cluster, and only after confirmation by a majority of nodes will the new mapping relationship be persisted and a successful write response returned to the client. This mechanism prevents data inconsistency caused by failed updates on some nodes. At the same time, the system records the original object Object-A as "to be recycled" in the PG-level metadata and marks its invalid data interval (such as the 512KB to 1MB interval of Object-A, which is recyclable because it has not been written) for subsequent garbage collection processes. For erasure coding (EC) scenarios, such as PG-200 under the 4+2 strategy, if a data shard of object Object-B is under reconstruction, the system will redirect the write to the OSD where the healthy shard is located in the same PG and dynamically calculate the new checksum shard to ensure the integrity of the EC shard group. This process is accelerated by the parallel computing framework, for example, distributing data sharding and checksum update tasks to multiple computing nodes and using GPUs to accelerate encoding operations, reducing the checksum recalculation time from milliseconds to microseconds.

[0061] The garbage collection (GC) module periodically scans PG metadata to identify "pending reclaim" objects and their inactive space, then merges the space and frees up resources. For example, if the GC detects that the 512KB to 1MB range of Object-A is unused, and that the 0KB to 512KB range of the adjacent object Object-C is also unused, it merges the two into a new object, Object-D (1MB in size), updates the metadata to point to Object-D, and marks the space of the original Object-A and Object-C as reusable. In EC scenarios, the GC also verifies shard validity: If a data shard in PG-200 is undergoing reconstruction, the GC delays reclamation until the reconstruction is complete and ensures data recoverability by querying the checksums of the remaining shards. For cross-node reclamation operations, such as migrating fragmented space from OSD-1 to OSD-5, the system uses incremental copy and checksum technology, allowing concurrent reads and writes during the migration process. A version number mechanism ensures data consistency before and after the migration. Client read requests are automatically routed to the new and old objects based on the version number until the migration is complete.

[0062] The exception handling mechanism covers edge cases such as full PG reconstruction, network partitions, and version conflicts. When all OSDs in a PG-100 are undergoing reconstruction (e.g., when all three replicas fail simultaneously), the system activates emergency write mode: client data is temporarily stored in a cross-PG cache pool (e.g., a temporary storage area based on memory or high-speed SSDs), and operation logs are recorded. After at least one OSD is rebuilt, the data is migrated back to the original PG via log replay. This process uses logical timestamps to ensure operation order and avoid data corruption. In the event of a network partition, if the primary OSD-2 loses connection to some replicas, the system re-elects a new primary OSD (e.g., OSD-3) through a majority agreement and re-establishes metadata consistency based on version numbers, ensuring that the redirection mapping table is correctly merged after partition recovery. For version conflicts caused by concurrent writes (e.g., two clients modifying the same logical object simultaneously), the system assigns an incremental logical clock value to each write. When reading, the data with the highest version number is automatically returned. Writes with lower versions are marked as historical and cleaned up asynchronously by the garbage collector, thus maintaining eventual data consistency without blocking clients.

[0063] Actual scenario tests show that this implementation significantly optimizes performance under various load conditions. In an additional write-intensive scenario, a financial trading system needs to continuously write log data. When the OSD-6 of PG-300 triggers reconstruction due to capacity expansion, after adopting this method, the write is redirected to OSD-8, the latency is stabilized at 0.35ms, the retry rate is reduced to 0.5%, and the storage fragmentation rate is controlled below 3% by GC. In an all-flash NVMe cluster, an AI training platform needs to write a large number of small files (4KB to 16KB) with high concurrency. When multiple PGs are concurrently reconstructed due to node failures, this method distributes the writes to 10 low-load OSDs through dynamic load balancing. The system throughput is maintained at 92% of the baseline level, and the latency fluctuation does not exceed 8%, which is significantly better than the traditional solution where the throughput plummets to 40% and the latency fluctuates by more than 200%. In addition, this method also performs well in cross-regional distributed storage: when the OSD cluster in a certain geographical area triggers a large-scale reconstruction due to network interruption, the system redirects writes to the remote replica with the lowest latency through topology-aware election, and combines it with the incremental synchronization protocol to quickly replenish data after the network is restored, achieving data consistency recovery in minutes.

[0064] The technical solution's implementation details are further reflected in the optimized design of the underlying data structures. For example, the metadata cluster utilizes an LSM-Tree (Log-Structured Merge Tree) storage architecture, writing frequently updated redirection mapping tables to a memory table and asynchronously flushing them to disk. Bloom filters accelerate queries, keeping the latency of status check operations stable at under 5 microseconds. The target election algorithm incorporates a machine learning model to predict future load trends for each OSD based on historical load patterns and dynamically adjust the weight calculation strategy. For example, if the model detects a periodic increase in load on OSD-7 between 3:00 PM and 5:00 PM daily, the system will proactively reduce its election weight during this time, thus preemptively avoiding potential performance bottlenecks. In the data write path, the system utilizes zero-copy technology, transferring client data directly to the target OSD's persistent memory area (such as Intel Optane PMem), bypassing kernel buffers and reducing memory copy overhead, resulting in write latency down to nanoseconds. These optimizations are deeply integrated with the core redirection logic, forming an end-to-end performance improvement loop.

[0065] In one embodiment, this specification provides an electronic device, including a processor and a readable storage medium, wherein the readable storage medium stores machine executable instructions that can be executed by the processor, and the processor executes the machine executable instructions to implement the aforementioned data reconstruction method. From a hardware perspective, the hardware architecture diagram can be found in Figure 5 shown.

[0066] In one embodiment, this specification provides a readable storage medium, wherein the readable storage medium stores machine-executable instructions. When the machine-executable instructions are called and executed by a processor, the machine-executable instructions prompt the processor to implement the aforementioned data reconstruction method.

[0067] Here, the readable storage medium can be any electronic, magnetic, optical or other physical storage device that can contain or store information, such as executable instructions, data, etc. For example, the readable storage medium can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drive (such as hard disk drive), solid state drive, any type of storage disk (such as optical disk, DVD, etc.), or similar storage media, or a combination thereof.

[0068] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer, which may be in the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email transceiver, game console, tablet computer, wearable device, or any combination of these devices.

[0069] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0070] Those skilled in the art will appreciate that embodiments of this specification may be provided as methods, systems, or computer program products. Thus, this specification may take the form of a fully hardware implementation, a fully software implementation, or an implementation combining software and hardware. Furthermore, embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0071] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0072] Furthermore, these computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0073] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0074] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Thus, this specification may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0075] The foregoing is merely an embodiment of the present invention and is not intended to limit the present invention. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be included within the scope of the claims of the present invention.

Claims

1. A data reconstruction method, characterized in that: Applied to a distributed storage device, the method includes: In response to a data write request, obtaining an operating status of a specific storage unit, the specific storage unit being a write target associated with the data write request; In response to the specific storage unit being in a data reconstruction state, selecting a target storage unit in a placement group to which the storage unit belongs that is not in a data reconstruction state as a new write target; Write the data associated with the data write request to the target storage unit.

2. The method according to claim 1, characterized in that The step of writing data associated with the data write request to the target storage unit includes: Use a garbage collection policy to process the storage space of a specific storage unit.

3. The method according to claim 1, characterized in that After the step of responding that the operating state of the specific storage unit is a data reconstruction state, the method further includes: Get the write type associated with the data write request; Locating, based on the result obtained that the write type is overwrite, a data cache area of the memory corresponding to the specific storage unit in the data reconstruction, overwriting the data associated with the data write request into the data cache area, and then continuing to perform data reconstruction associated with the specific storage unit; Stop executing the step of selecting a target storage unit in the placement group to which the storage unit belongs, whose operating state is not in data reconstruction, as a new write object, and writing the data associated with the data write request to the target storage unit.

4. The method according to claim 1, wherein In response to the specific storage unit being in a data reconstruction state, selecting a target storage unit in a placement group to which the storage unit belongs that is not in a data reconstruction state as a new write target includes: In response to the operating state of the specific storage unit being a data reconstruction state, obtaining a write type associated with the data write request; According to the acquisition result of the write type being append write, a target storage unit that is not in data reconstruction operation status is selected as a new write target in the placement group to which the storage unit belongs.

5. A data reconstruction device, characterized in that: Applied to a distributed storage device, the apparatus comprises: A first module is configured to obtain, in response to a data write request, an operating status of a specific storage unit, where the specific storage unit is a write target associated with the data write request; The second module is configured to, in response to the operating state of the specific storage unit being in a data reconstruction state, select a target storage unit in the placement group to which the storage unit belongs that is not in a data reconstruction state as a new write target; The third module is configured to write data associated with the data write request into the target storage unit.

6. The device according to claim 5, characterized in that The step of writing data associated with the data write request to the target storage unit includes: Use a garbage collection policy to process the storage space of a specific storage unit.

7. The device according to claim 5, characterized in that After the operating state of the specific storage unit is in the data reconstruction state, the second module further includes: Get the write type associated with the data write request; Locating, based on the result obtained that the write type is overwrite, a data cache area of the memory corresponding to the specific storage unit in the data reconstruction, overwriting the data associated with the data write request into the data cache area, and then continuing to perform data reconstruction associated with the specific storage unit; Stop calling the second module and the third module to execute the steps of selecting a target storage unit that is not in data reconstruction as a new write target in the placement group to which the storage unit belongs, and writing data associated with the data write request to the target storage unit.

8. The device according to claim 5, characterized in that In response to the specific storage unit being in a data reconstruction state, selecting a target storage unit in a placement group to which the storage unit belongs that is not in a data reconstruction state as a new write target includes: In response to the operating state of the specific storage unit being a data reconstruction state, obtaining a write type associated with the data write request; According to the acquisition result of the write type being append write, a target storage unit that is not in data reconstruction operation status is selected as a new write target in the placement group to which the storage unit belongs.

9. An electronic device, characterized in that: include: A processor and a readable storage medium, wherein the readable storage medium stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the method according to any one of claims 1 to 4.

10. A readable storage medium, characterized in that: The readable storage medium stores machine-executable instructions. When the machine-executable instructions are called and executed by a processor, the machine-executable instructions prompt the processor to implement the method according to any one of claims 1 to 4.