Data reconstruction method, electronic device, medium and product

CN122653899APending Publication Date: 2026-08-28INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611139965.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-29
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

然而,传统RAID重构机制的重构速率仅基于底层输入/输出(Input/Output,简称为I/O)负载指标动态调整,这导致重构I/O与AI训练I/O争抢资源,导致AI训练效率低下

Benefits of technology

[0017] The data reconstruction method, electronic device, medium, and product provided in this application acquire the input/output feature data of a preset training task after detecting a faulty disk, determine the current target training stage of the training task, determine the target reconstruction rate based on the target training stage, and reconstruct the data in the faulty disk based on the target reconstruction rate. This enables the reconstruction process to actively avoid the critical stages of the training process, ensuring sufficient training resources. Based on this, the resource conflict between reconstruction I/O and AI training I/O is reduced, thereby improving AI training efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122653899A_ABST
    Figure CN122653899A_ABST
Patent Text Reader

Abstract

The application discloses a data reconstruction method, an electronic device, a medium and a product. The method comprises the following steps: in response to detecting a faulty disk, obtaining input / output characteristic data of a preset training task; determining a target training phase in which the training task currently locates according to the input / output characteristic data of the training task; determining a target reconstruction rate according to the target training phase; and performing reconstruction processing on data in the faulty disk based on the target reconstruction rate. According to the data reconstruction method provided by the application, the resource conflict between reconstruction I / O and AI training I / O can be reduced, and the AI training efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of storage systems and data processing technology, and in particular to a data reconstruction method, electronic device, medium and product. Background Technology

[0002] In deep learning training scenarios, the backend storage system of artificial intelligence (AI) training clusters typically relies on Redundant Array of Independent Disks (RAID) technology to ensure data reliability and performance. RAID technology combines multiple disks into logical storage units, achieving data striping and redundancy verification, and reconstructing lost data through XOR operations in the event of disk failure. However, the reconstruction rate of traditional RAID reconstruction mechanisms is dynamically adjusted based solely on the underlying input / output (I / O) load metrics. This leads to competition for resources between reconstruction I / O and AI training I / O, resulting in low AI training efficiency. Summary of the Invention

[0003] This application provides a data reconstruction method, electronic device, medium, and product that can reduce resource conflicts between reconstruction I / O and AI training I / O, thereby improving AI training efficiency.

[0004] This application provides a data reconstruction method, including:

[0005] In response to the detection of a faulty disk, the input / output feature data of a preset training task are acquired;

[0006] Based on the input / output feature data of the training task, determine the current target training stage of the training task;

[0007] Determine the target reconstruction rate based on the target training phase;

[0008] The data in the faulty disk is reconstructed based on the target reconstruction rate.

[0009] This application also provides a data reconstruction apparatus, including:

[0010] The acquisition module is used to acquire the input / output feature data of a preset training task in response to the detection of a faulty disk;

[0011] The determination module is used to determine the current target training stage of the training task based on the input / output feature data of the training task.

[0012] The determination module is also used to determine the target reconstruction rate based on the target training phase;

[0013] The processing module is used to reconstruct the data in the faulty disk based on the target reconstruction rate.

[0014] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for implementing the above-described data reconstruction method when executing the computer program.

[0015] This application also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of the above-described data reconstruction method.

[0016] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described data reconstruction method.

[0017] The data reconstruction method, electronic device, medium, and product provided in this application acquire the input / output feature data of a preset training task after detecting a faulty disk, determine the current target training stage of the training task, determine the target reconstruction rate based on the target training stage, and reconstruct the data in the faulty disk based on the target reconstruction rate. This enables the reconstruction process to actively avoid the critical stages of the training process, ensuring sufficient training resources. Based on this, the resource conflict between reconstruction I / O and AI training I / O is reduced, thereby improving AI training efficiency. Attached Figure Description

[0018] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a schematic diagram of a scenario provided for an embodiment of this application;

[0020] Figure 2 Flowchart of the data reconstruction method provided in the embodiments of this application Figure 1 ;

[0021] Figure 3 Flowchart of the data reconstruction method provided in the embodiments of this application Figure 2 ;

[0022] Figure 4 Flowchart of the data reconstruction method provided in the embodiments of this application Figure 3 ;

[0023] Figure 5 Flowchart of the data reconstruction method provided in the embodiments of this application Figure 4 ;

[0024] Figure 6 Flowchart of the data reconstruction method provided in the embodiments of this application Figure 5 ;

[0025] Figure 7 Flowchart of the data reconstruction method provided in the embodiments of this application Figure 6 ;

[0026] Figure 8 Flowchart of the data reconstruction method provided in the embodiments of this application Figure 7 ;

[0027] Figure 9 Flowchart of the data reconstruction method provided in the embodiments of this application Figure 8 ;

[0028] Figure 10 Flowchart of the data reconstruction method provided in the embodiments of this application Figure 9 ;

[0029] Figure 11 Flowchart of the data reconstruction method provided in the embodiments of this application Figure 10 ;

[0030] Figure 12 Flowchart of the data reconstruction method provided in the embodiments of this application Figure 10 one;

[0031] Figure 13 Flowchart of the data reconstruction method provided in the embodiments of this application Figure 10 two;

[0032] Figure 14 This is a schematic diagram of the overall process of the data reconstruction method provided in the embodiments of this application;

[0033] Figure 15 This is a schematic diagram of the structure of the data reconstruction apparatus provided in the embodiments of this application;

[0034] Figure 16 A schematic diagram of the structure of the electronic device provided in this application. Detailed Implementation

[0035] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, other embodiments obtained by those of ordinary skill in the art without creative effort are all within the protection scope of this application.

[0036] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0037] It should also be noted that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," "counterclockwise," "axial," "radial," and "circumferential," etc., indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application. The terms "installed," "connected," and "linked" should be interpreted broadly, for example, they can be fixed connections, detachable connections, or integral connections; they can be mechanical connections or electrical connections; they can be direct connections or indirect connections through an intermediate medium; they can be internal connections between two elements. The terms "parallel," "perpendicular," and "equal" include the described situation and situations similar to the described situation, the range of which is within an acceptable deviation range, wherein the acceptable deviation range is determined by those skilled in the art taking into account the measurement under discussion and the error associated with the measurement of a particular quantity (i.e., the limitations of the measurement system). For example, "parallel" includes absolute parallelism and approximate parallelism, where an acceptable deviation range for approximate parallelism can be, for example, within 5°; "perpendicular" includes absolute perpendicularity and approximate perpendicularity, where an acceptable deviation range for approximate perpendicularity can also be, for example, within 5°. "Equal" includes absolute equality and approximate equality, where an acceptable deviation range for approximate equality can be, for example, a difference between the two equal items being less than or equal to 5% of either one. Those skilled in the art will understand the specific meaning of the above terms in this application based on the specific circumstances.

[0038] Figure 1 This is a schematic diagram illustrating a scenario provided in an embodiment of this application. In deep learning training scenarios, the backend storage system of an AI training cluster typically relies on RAID technology to ensure data reliability and performance. RAID technology combines multiple disks into a single unit, improving data read / write speed and reliability while saving space. For example, Figure 1The example illustrates multiple disks, including disk 1, disk 2, disk 3, disk 4 up to disk n. RAID divides files into multiple data blocks and stores them on different disks in a "data striping" format. In addition, RAID also stores additional parity information; if one of the disks fails, for example, ... Figure 1 As shown, if disk 3 fails, all data on disk 3 can be recalculated based on data and verification information from other disks and written to a new spare disk. Then, the spare disk replaces disk 3, completing the reconstruction of the data on disk 3.

[0039] Currently, traditional RAID reconstruction mechanisms dynamically adjust their reconstruction rate based solely on underlying input / output (I / O) load metrics and recover data from failed disks according to a predetermined reconstruction process. The fundamental premise of this approach is to utilize a background reconstruction mechanism to achieve redundant recovery, ensuring the availability and reliability of stored data. However, relying solely on underlying load metrics for reconstruction control makes it difficult to identify the differentiated I / O characteristics at different stages of the training task, such as data loading, parameter saving, and training computation. Therefore, this leads to resource contention between reconstruction I / O and AI training I / O, resulting in low AI training efficiency.

[0040] The data reconstruction method provided in this application obtains the input / output feature data of a preset training task after detecting a faulty disk, determines the current target training stage of the training task based on this data, determines the target reconstruction rate based on the target training stage, and reconstructs the data in the faulty disk based on the target reconstruction rate. This enables the reconstruction process to actively avoid the critical stages of the training process, ensuring sufficient training resources. Based on this, the resource conflict between reconstruction I / O and AI training I / O is reduced, thereby improving AI training efficiency.

[0041] The embodiments of this application provide a data reconstruction method, and the system is described in detail below: Figure 2 Flowchart of the data reconstruction method provided in the embodiments of this application Figure 1 ,like Figure 2 As shown, it includes:

[0042] S201. In response to the detection of a faulty disk, acquire the input / output feature data of a preset training task.

[0043] For example, a failed disk is a disk unit in a RAID storage system that is detected as unable to continue providing normal data services and triggers a redundancy recovery process. This can correspond to member disks with physical damage, link failure, unreadable media, persistent verification failures, or logical errors reaching a preset threshold. The input / output feature data of the training task is used to characterize the underlying I / O information initiated to the backend storage during the training task's execution. After detecting a failed disk, I / O information related to the preset training task can be collected and processed to obtain input / output feature data characterizing the current storage access status of the training task.

[0044] S202. Based on the input / output feature data of the training task, determine the current target training stage of the training task.

[0045] For example, the target training stage is a stage label identified from multiple preset training stages based on the current input / output feature data of the training task. In this example, the preset training stages may include the data loading stage, the keypoint writing stage, and the training computation stage. The data loading stage typically corresponds to the stage where the training process centrally reads sample data; the keypoint writing stage typically corresponds to the stage where model parameters, optimizer states, or intermediate checkpoints are centrally written to disk; and the training computation stage typically corresponds to the stage where computing resources perform forward and backward computations while storage access is relatively sparse. In this example, based on the input / output feature data obtained in S201, a comprehensive analysis of the input / output feature data can be performed to obtain the current target training stage of the training task.

[0046] S203. Determine the target reconstruction rate based on the target training phase.

[0047] For example, a target reconstruction rate can be determined for the current reconstruction task based on the target training phase, ensuring that the I / O requirements of the background reconstruction and training tasks are coordinated at the current stage. If the target training phase changes, the target reconstruction rate can be adjusted accordingly. For instance, the data loading phase and keypoint writing phase are typically critical phases of the training task. When the target training phase is the data loading phase or the keypoint writing phase, the reconstruction rate can be appropriately reduced to achieve the target reconstruction rate; conversely, when the target training phase is the training computation phase, the reconstruction rate can be appropriately increased to achieve the target reconstruction rate.

[0048] S204. Reconstruct the data in the faulty disk based on the target reconstruction rate.

[0049] For example, using a target reconstruction rate, the data on the faulty disk is reconstructed based on the data and verification information stored on other normal disks, using an XOR algorithm. For instance, combining... Figure 1Since the faulty disk is disk 3, the data stored on the other disks besides disk 3 can be XORed based on the parity information stored in the RAID array to calculate the data in disk 3. In this example, reconstructing the data on the faulty disk according to the target reconstruction rate keeps the impact of background reconstruction on the foreground training task within a relatively controllable range.

[0050] Based on the above analysis, the data reconstruction method provided in this example obtains the input / output feature data of the preset training task after detecting a faulty disk, determines the current target training stage of the training task, determines the target reconstruction rate based on the target training stage, and reconstructs the data in the faulty disk based on the target reconstruction rate. This enables the reconstruction process to actively avoid the critical stages of the training process, ensuring sufficient training resources. Based on this, the resource conflict between reconstruction I / O and AI training I / O is reduced, thereby improving AI training efficiency.

[0051] Optional, Figure 3 Flowchart of the data reconstruction method provided in the embodiments of this application Figure 2 ,like Figure 3 As shown, S202 includes:

[0052] S301. Using a sliding window of a preset size, feature extraction is performed on the input / output feature data to obtain the distribution characteristics of read / write requests, the spatial characteristics of read / write access, and the activity of input / output corresponding to each sliding window.

[0053] For example, in this example, the input / output feature data includes at least the distribution characteristics of read / write requests, the spatial characteristics of read / write access, and the activity level of input / output. The distribution characteristics of read / write requests represent the proportion of read and write requests within different request size ranges; the spatial characteristics of read / write access represent the continuity, hopping frequency, or hotspot concentration of access addresses in the logical block address space; and the activity level of input / output represents at least one of the following per unit time: number of requests, throughput, queue occupancy rate, or request arrival rate. The size of the sliding window can be determined according to the actual situation; for example, a 5-minute sliding window can be used in this example. Therefore, for feature extraction of the input / output feature data, the input / output feature data is sliced ​​with a 5-minute sliding window. Within each time window, the number of read / write requests, request block size, access address difference, and number of bytes transferred per unit time are collected. The collected number of read / write requests, request block size, access address difference, and number of bytes transferred per unit time are then mapped to the distribution characteristics of read / write requests, the spatial characteristics of read / write access, and the activity level of input / output.

[0054] S302. Based on the distribution characteristics of read and write requests, spatial characteristics of read and write access, and input / output activity corresponding to each sliding window, determine the training stage corresponding to each sliding window.

[0055] For example, the training phase includes the data loading phase, the keypoint writing phase, or the training computation phase. The distribution characteristics of read and write requests, the spatial characteristics of read and write access, and the activity of input / output are different for each training phase. Therefore, the training phase corresponding to each window is determined based on the distribution characteristics of read and write requests, the spatial characteristics of read and write access, and the activity of input / output for each sliding window.

[0056] Optional, Figure 4 Flowchart of the data reconstruction method provided in the embodiments of this application Figure 3 ,like Figure 4 As shown, S302 includes:

[0057] S401. For each sliding window, if the distribution characteristics of the read and write requests of the sliding window are such that the proportion of small file read requests is greater than or equal to the first preset threshold, the spatial characteristics of read and write access are such that random skipping reads, and the activity of input / output is greater than or equal to the second preset threshold, then the training stage corresponding to the sliding window is determined to be the data loading period, wherein a small file read request is defined as a read request involving a data amount less than or equal to the third preset threshold.

[0058] For example, the first, second, and third preset thresholds can all be determined based on actual conditions. For instance, the first preset threshold could be 70%, the second preset threshold could be 500 MB / s, and the third preset threshold could be 1 MB. In other words, in the distribution characteristics of read / write requests within the sliding window, if read requests less than or equal to 1 MB account for more than or equal to 70% of all I / O requests within the sliding window, and in the spatial characteristics of read / write access within the sliding window, read requests exhibit a highly random skipping pattern, and the input / output activity of the sliding window is greater than or equal to 500 MB / s, then the training phase of the sliding window can be considered the data loading phase.

[0059] S402. For each sliding window, if the distribution characteristics of the read and write requests of the sliding window are such that the proportion of large file write requests is greater than or equal to the first preset threshold, the spatial characteristics of read and write access are such that the access is sequential, and the activity of input / output is greater than or equal to the second preset threshold, then the training stage corresponding to the sliding window is determined to be the key point writing period. Among them, large file write requests are those in which the amount of data involved in the write request is greater than or equal to the fourth preset threshold.

[0060] Similarly, the fourth preset threshold can be determined according to the actual situation, such as 100MB. That is to say, in the distribution characteristics of read and write requests in the sliding window, if write requests greater than or equal to 100MB account for greater than or equal to 70% of all I / O requests in the sliding window, and in the spatial characteristics of read and write access in the sliding window, the write requests are in the form of sequential access, and the input / output activity of the sliding window is greater than or equal to 500MB / s, then the training phase of the sliding window can be determined as the key point writing period.

[0061] S403. For each sliding window, if the distribution characteristics of read and write requests, the spatial characteristics of read and write access, and the activity of input / output of the sliding window do not conform to the data characteristics of the data loading period or the data characteristics of the key point writing period, then the training phase corresponding to the sliding window is determined to be the training calculation period.

[0062] For example, if the distribution characteristics of read and write requests in the sliding window show no obvious read / write bias, and the overall input / output activity of the sliding window is at a low level, such as below 100 MB / s, then the training phase of the sliding window can be determined as the training computation period. Alternatively, if the training phase of the sliding window cannot be determined based on the distribution characteristics of read and write requests, the spatial characteristics of read and write access, and the input / output activity, then the training phase of the sliding window can also be determined as the training computation period.

[0063] Based on the method provided in this example, the determination of the sliding window training phase does not rely on a single load metric, but simultaneously considers the distribution characteristics of read and write requests, the spatial characteristics of read and write access, and the activity of input / output. This ensures that the determination result of the current training phase is consistent with the actual I / O pattern of the training task, thereby improving the accuracy of the determination of the training phase.

[0064] S303. Determine the target training stage of the training task based on the training stage corresponding to each sliding window.

[0065] For example, the identification results of multiple consecutive windows can be further combined to select the stage with the highest proportion or the strongest continuity as the target training stage for the current training task. This example uses the identification results of the training stage at the sliding window level as the basis for data reconstruction, which can keep the process of reconstructing the faulty disk consistent with the foreground training I / O rhythm. This can reduce the bandwidth resource consumption of reconstruction during the critical read and write phases of training.

[0066] Optionally, S303 includes:

[0067] Based on the training phase corresponding to each sliding window, if multiple consecutive sliding windows correspond to the data loading phase, the target training phase is determined to be the data loading phase; if multiple consecutive sliding windows correspond to the keypoint writing phase, the target training phase is determined to be the keypoint writing phase; if multiple consecutive sliding windows correspond to the training computation phase, the target training phase is determined to be the training computation phase.

[0068] For example, when several adjacent sliding windows exhibit consistent training phases, this continuous training phase can be considered a stable phase rather than a single instantaneous fluctuation, thus defining the target training phase as this continuous training phase. The number of consecutive consistent training phases can be determined based on the actual situation. For instance, it could be three times. That is, if three consecutive sliding windows are determined to be data loading phases, the output data loading phase is taken as the target training phase; if three consecutive sliding windows are determined to be keypoint writing phases, the output keypoint writing phase is taken as the target training phase; if three consecutive sliding windows are determined to be training computation phases, the output training computation phase is taken as the target training phase. Alternatively, if no training phase occurs three times consecutively, the target training phase can be defined as the training computation phase.

[0069] The method of determining the target training phase uses the consistency of continuous windows as the criterion, which can limit the impact of short-term jitter and instantaneous burst access on the training phase to a controllable range and improve the accuracy of determining the target training phase.

[0070] Optionally, S303 includes:

[0071] Based on the training stage corresponding to each sliding window, if the proportion of the training stage that is the data loading stage is greater than or equal to the fifth preset threshold, then the target training stage is determined to be the data loading stage; if the proportion of the training stage that is the keypoint writing stage is greater than or equal to the fifth preset threshold, then the target training stage is determined to be the keypoint writing stage; if the proportion of the training stage that is the training calculation stage is greater than or equal to the fifth preset threshold, then the target training stage is determined to be the training calculation stage.

[0072] For example, the fifth preset threshold can be determined according to the actual situation, such as 60%. After obtaining the training stages corresponding to multiple sliding windows, the proportion of different training stages in the training stages corresponding to multiple sliding windows can be determined. For example, based on the training stages of the most recent five sliding windows, it can be determined whether there is a training stage with a proportion exceeding 60% in these five sliding window training stages. For example, if there are at least three data loading periods in the training stages of the most recent five sliding windows, then the target training stage is determined to be the data loading period. Similarly, if there are at least three keypoint writing periods in the training stages of the most recent five sliding windows, then the target training stage is determined to be the keypoint writing period. If there are at least three training computation periods in the training stages of the most recent five sliding windows, then the target training stage is determined to be the training computation period. In addition, if there is no training stage with a proportion exceeding 60% in the training stages of the most recent five sliding windows, then the target training stage can be determined to be the training computation period.

[0073] The method for determining the target training stage is jointly determined by the training stage distribution results of multiple sliding windows. This method can reflect the dominant I / O characteristics of the training business over a period of time and reduce the impact of fluctuations in a single sliding window on the determination of the target training stage, thereby improving the accuracy of determining the target training stage.

[0074] Optional, also includes:

[0075] In the initial stage of data reconstruction, the target training stage is determined to be the training computation period.

[0076] For example, the first five minutes after data reconstruction begins can be defined as the initial phase of data reconstruction. The duration of the initial phase can be adjusted according to the actual situation. Within the initial phase, the target training phase can be defined as the training computation period. Based on the method provided in this example, in the initial phase of data reconstruction, it is necessary to moderately accelerate the process to quickly restore redundancy capabilities. Therefore, the target training phase at this time can be directly defined as the training computation period to quickly respond to the data reconstruction processing and thus improve the efficiency of data reconstruction.

[0077] Optional, Figure 5 Flowchart of the data reconstruction method provided in the embodiments of this application Figure 4 ,like Figure 5 As shown, S203 includes:

[0078] S501. Determine the target baseline reconstruction rate corresponding to the target training phase.

[0079] For example, first select the corresponding target baseline reconstruction rate according to the target training stage, such as setting different baseline values ​​during the data loading period, key point writing period and training calculation period.

[0080] Optional, Figure 6 Flowchart of the data reconstruction method provided in the embodiments of this application Figure 5 ,like Figure 6 As shown, S501 includes:

[0081] S601. If the target training phase is the data loading phase, then the first preset benchmark rate is determined as the target benchmark reconstruction rate.

[0082] Optionally, since the data loading period is a critical training phase, the first preset baseline rate should be selected as a relatively small value. For example, the first preset baseline rate can be set to 10MB / s, or it can be adjusted according to the actual situation.

[0083] S602. If the target training phase is the key point writing phase, then the second preset benchmark rate is determined as the target benchmark reconstruction rate, wherein the second preset benchmark rate is less than the first preset benchmark rate.

[0084] Optionally, since the keypoint writing period is also a critical training phase, reconstruction processing usually needs to be paused during the keypoint writing period. Therefore, the second preset baseline rate can be set to 0MB / s, or it can be adjusted according to the actual situation.

[0085] S603. If the target training phase is the training computation period, then the third preset benchmark rate is determined as the target benchmark reconstruction rate, wherein the third preset benchmark rate is greater than the first preset benchmark rate.

[0086] Optionally, since the training computation period is a non-critical training phase, the third preset baseline rate should be selected as a relatively large value. For example, the first preset baseline rate can be set to 80MB / s, or it can be adjusted according to the actual situation.

[0087] The data reconstruction method improved in this application determines different target baseline reconstruction rates for different training stages, which can reserve sufficient bandwidth for batch reading during the data loading period, reduce reconstruction occupancy during the key point writing period, and increase recovery strength during the training computation period, thereby improving the adaptability of the reconstruction process of faulty disk data to each training stage.

[0088] S502, collect load status data and reconstruction progress.

[0089] For example, load status data can characterize I / O pressure, and reconstruction progress is the progress of reconstructing data in the failed disk.

[0090] Optionally, S502 includes:

[0091] According to the preset collection interval, real-time load status data and real-time reconstruction progress are collected.

[0092] For example, the data collection interval can be determined based on actual conditions, such as 5 seconds. Collecting current load status data and reconstruction progress every 5 seconds provides real-time load status data and reconstruction progress. Based on the method provided in this example, and using a preset collection interval, the collected load status data and reconstruction progress are guaranteed to be timely and accurately reflect the current load status and reconstruction progress.

[0093] S503. Based on the load status data and reconstruction progress, adjust the target baseline reconstruction rate to obtain the target reconstruction rate.

[0094] For example, when I / O pressure is high, the reconstruction rate can be appropriately reduced from the target baseline reconstruction rate; when I / O pressure is low, the reconstruction rate can be appropriately increased from the target baseline reconstruction rate. Furthermore, the reconstruction rate can be appropriately increased in the early stages of reconstruction to restore redundancy as quickly as possible, while the reconstruction rate should be appropriately reduced as the reconstruction nears completion to avoid sudden pressure on the storage system. By combining load status data and reconstruction progress, the target baseline reconstruction rate can be appropriately adjusted to obtain the final target reconstruction rate.

[0095] Based on the method provided in this example, the target baseline reconstruction rate corresponding to the training phase can be linked with real-time load status data and reconstruction progress. This can maintain the continuity of data recovery for failed disks and ensure that the training task does not compete for resources, thus ensuring the smooth progress of the training task.

[0096] Optional, Figure 7 Flowchart of the data reconstruction method provided in the embodiments of this application Figure 6 ,like Figure 7 As shown, S503 includes:

[0097] S701. Determine the first adjustment ratio based on real-time load status data.

[0098] For example, after collecting real-time load status data, the load status data is mapped to a first adjustment ratio. This mapping can be done using a linear function, a piecewise function, or a lookup table. The first adjustment ratio can be positive or negative. A positive first adjustment ratio indicates that the reconstruction rate can be appropriately increased based on the current real-time load status data; the larger the value of the first adjustment ratio, the greater the increase in reconstruction rate. For example, if the first adjustment ratio is 10%, the reconstruction rate will be increased by 10% based on the target baseline reconstruction rate. A negative first adjustment ratio indicates that the reconstruction rate can be appropriately decreased based on the current real-time load status data. If the first adjustment ratio is -10%, the reconstruction rate will be decreased by 10% based on the target baseline reconstruction rate.

[0099] S702. Determine the second adjustment ratio based on the real-time reconstruction progress.

[0100] Similarly, after collecting real-time reconstruction progress data, the load status data is mapped to a second adjustment ratio. Since the reconstruction rate is adaptively increased in the early stages and adaptively decreased in the later stages, the second adjustment ratio is positive when the reconstruction progress is below 50%, and the smaller the reconstruction progress, the larger the second adjustment ratio, indicating a greater increase in the target baseline reconstruction rate. When the reconstruction progress is above 50%, the second adjustment ratio is negative, and the larger the reconstruction progress, the larger the absolute value of the second adjustment ratio, indicating a greater decrease in the target baseline reconstruction rate. When the reconstruction progress is exactly 50%, the second adjustment ratio can be zero.

[0101] S703. Determine the first adjustment weight corresponding to the load status data and the second adjustment weight corresponding to the reconstruction progress.

[0102] For example, the impact of load status data and reconstruction progress on reconstruction rate is usually based on load status data, so the first adjustment weight is greater than the second adjustment weight. For example, the first adjustment weight can be 0.7 and the second adjustment weight can be 0.3.

[0103] S704. Based on the first adjustment weight and the second adjustment weight, perform a weighted summation of the first adjustment ratio and the second adjustment ratio to obtain the real-time target adjustment ratio.

[0104] For example, if the first adjustment ratio is 10% and the second adjustment ratio is 20%, then after weighting and summing the first and second adjustment ratios according to the first and second adjustment weights, the target adjustment ratio is 13%.

[0105] S705. Adjust the target baseline reconstruction rate in real time according to the real-time target adjustment ratio to obtain the target reconstruction rate.

[0106] For example, if the target adjustment ratio is 13%, then the target reconstruction rate is obtained by increasing the target baseline reconstruction rate by 1%.

[0107] Based on the method provided in this example, the weights can be adjusted according to the load status data and the reconstruction progress, respectively. The impact of load status data and reconstruction progress on the reconstruction rate can be integrated to improve the accuracy of the final target reconstruction rate.

[0108] Optional, Figure 8 Flowchart of the data reconstruction method provided in the embodiments of this application Figure 7 ,like Figure 8 As shown, S705 includes:

[0109] S801. Adjust the target baseline reconstruction rate in real time according to the real-time target adjustment ratio to obtain the adjusted real-time reconstruction rate.

[0110] For example, based on the aforementioned target adjustment ratio of 13%, the result of increasing the target baseline reconstruction rate by 1% can be determined as the real-time reconstruction rate. For instance, taking the target training phase as the training computation period, if the target baseline reconstruction rate during the training computation period is 80MB / s, then increasing 80MB / s by 1% will result in a real-time reconstruction rate of 90.4MB / s.

[0111] S802. Limit the adjusted real-time reconstruction rate according to the preset reconstruction rate range to obtain the target reconstruction rate.

[0112] For example, the reconstruction rate range can be 5MB / s-100MB / s. By limiting the real-time reconstruction rate to the range of 5MB / s-100MB / s, the final target reconstruction rate can be obtained.

[0113] Based on the method provided in this example, the stability of the target reconstruction rate can be guaranteed by limiting the target reconstruction rate within a certain range.

[0114] Optional, Figure 9 Flowchart of the data reconstruction method provided in the embodiments of this application Figure 8 ,like Figure 9 As shown, S802 includes:

[0115] S901. Determine the upper limit and lower limit of the reconstruction rate range.

[0116] For example, based on the reconstruction rate range of the previous example, the upper limit of the reconstruction rate is 100MB / s, and the lower limit of the reconstruction rate is 5MB / s.

[0117] S902. If the adjusted real-time reconstruction rate is less than the lower limit of the reconstruction rate, then the lower limit of the reconstruction rate shall be determined as the target reconstruction rate.

[0118] For example, taking the data loading phase as the target training phase, the baseline reconstruction rate during the data loading phase is 10MB / s. If the target adjustment ratio is -60%, the adjusted real-time reconstruction rate will be 4MB / s. Since 4MB / s is less than 5MB / s, the target reconstruction rate can be set at 5MB / s.

[0119] S903. If the adjusted real-time reconstruction rate is greater than the upper limit of the reconstruction rate, then the upper limit of the reconstruction rate shall be determined as the target reconstruction rate.

[0120] For example, taking the training computation phase as the target training phase, the baseline reconstruction rate during the training computation phase is 80 MB / s. If the target adjustment ratio is 30%, the adjusted real-time reconstruction rate will be 104 MB / s. Since 104 MB / s is greater than 100 MB / s, the target reconstruction rate can be set at 100 MB / s.

[0121] S904. If the adjusted real-time reconstruction rate is greater than or equal to the lower limit of the reconstruction rate and less than or equal to the upper limit of the reconstruction rate, then the adjusted real-time reconstruction rate shall be determined as the target reconstruction rate.

[0122] For example, taking the training computation phase as the target training phase, the baseline reconstruction rate during the training computation phase is 80 MB / s. If the target adjustment ratio is 13%, the adjusted real-time reconstruction rate will be 90.4 MB / s. At this point, 90.4 MB / s falls within the range of 5 MB / s to 100 MB / s, so the target reconstruction rate can be directly determined as 90.4 MB / s.

[0123] It's worth noting that, taking the keypoint writing phase as an example during the target training stage, since the baseline writing rate during this phase is set at 0 MB / s, the result after adjusting the rate according to the target adjustment ratio will still be 0 MB / s. In this case, the target reconstruction rate can be set to 5 MB / s based on the limitations of the reconstruction rate range.

[0124] Based on the method provided in this example, the stability of the target reconstruction rate can be guaranteed by limiting the target reconstruction rate within a certain range.

[0125] Optional, Figure 10 Flowchart of the data reconstruction method provided in the embodiments of this application Figure 9 ,like Figure 10 As shown, S204 includes:

[0126] S1001. Identify faulty data in the faulty disk.

[0127] For example, the sector, block, or stripe location of the faulty data can be located first based on the logical address mapping relationship of the faulty disk. The faulty disk is divided into 64KB stripes, which are stored in different storage units.

[0128] S1002. Determine at least one faulty storage unit where the faulty data is located, and at least one normal storage unit other than the faulty storage unit.

[0129] For example, the storage unit containing the faulty data can be marked as a faulty storage unit, while storage units in the same faulty disk that do not store faulty data can be marked as normal storage units. There can be one or more faulty storage units, and there can be one or more normal storage units.

[0130] S1003. Reconstruct the data in each failed storage unit to the preset backup disk.

[0131] For example, data in a failed storage unit is reconstructed to a standby disk using an XOR algorithm at a target reconstruction rate. The standby disk can be a replacement disk of the same specifications as the failed disk, or it can be a disk of another model with the same capacity and interface protocol.

[0132] Optional, Figure 11 Flowchart of the data reconstruction method provided in the embodiments of this application Figure 10 ,like Figure 11 As shown, S1003 includes:

[0133] S1101. Generate each data reconstruction task for each faulty storage unit and put each reconstruction task into the main queue in sequence.

[0134] For example, after identifying multiple faulty storage units, a corresponding reconstruction task is generated for each faulty storage unit. Each reconstruction task includes at least the target faulty storage unit identifier, the address range of data to be recovered, and the target address information to be written to the backup disk.

[0135] S1102. Execute the reconstruction tasks in the main queue in sequence to reconstruct the data in each failed storage unit to the standby disk.

[0136] For example, the reconstruction tasks can then be sequentially sent to the main queue in a preset order. This preset order can be determined based on the faulty storage unit number, the amount of remaining data, or the order in which the tasks were created, to ensure the determinism of the subsequent execution process. The main queue can be maintained by a queue controller, which sequentially retrieves the reconstruction tasks and assigns them to the reconstruction execution units. The reconstruction execution units then execute each reconstruction task sequentially to complete the data reconstruction of each faulty storage unit.

[0137] Based on the method provided in this example, the recovery process of multiple failed storage units can be uniformly incorporated into the main queue, the task execution order is clear, the reconstruction process is continuous and controllable, thereby improving the organization and execution stability of data recovery, and keeping the reconstruction write process on the standby disk consistent.

[0138] S1004. Copy the data from each normal storage unit to the spare disk.

[0139] For example, for a normal storage unit, its raw data is read directly and copied to the same or preset mapped location on a spare disk.

[0140] Based on the method provided in this example, the damaged data and intact data in the failed disk are processed separately, the spare disk can take over the complete recovery results, and the reconstruction process and the copying process work together to maintain relatively stable reconstruction efficiency while meeting the data recovery requirements.

[0141] Optional, Figure 12 Flowchart of the data reconstruction method provided in the embodiments of this application Figure 10 Second, such as Figure 12 As shown, it also includes:

[0142] S1201. Detect whether there is hot data in each faulty storage unit.

[0143] For example, the hotness or coldness of data can be determined based on the historical access records of the data. Data that has been frequently accessed recently can be identified as hot data. For instance, if the data in a faulty storage unit is accessed more than 100 times within 10 seconds, it is determined that there is hot data in that faulty storage unit.

[0144] S1202. If there is hot data in the faulty storage unit, the data reconstruction task of the faulty storage unit is placed in the delayed processing queue.

[0145] For example, when hot data exists in a faulty storage unit, reconstruction will not be initiated immediately. Instead, a corresponding data reconstruction task will be generated and written to a delayed processing queue, while retaining the logical address range, the location of the reconstruction source data, and the identifier of the target spare disk.

[0146] S1203. During the training computation period of the target training phase, and when the load status data representation resources are idle, the reconstruction task in the delayed processing queue is executed to reconstruct the data in the faulty storage unit to the spare disk.

[0147] For example, when the target training phase is the training computation period and the load status data indicates that the current storage resources are idle, the corresponding task is retrieved from the delayed processing queue, and the data is reconstructed to the spare disk according to the data mapping relationship of the faulty storage unit.

[0148] This example enables the reconstruction tasks corresponding to hot data to be executed during the training computation period when resources are idle. This matches the access intensity of the reconstruction behavior with that of the training task and keeps the impact on front-end data access to a low level. At the same time, the delayed processing queue centrally stores and uniformly releases hot tasks, keeping the fault recovery process controllable, thereby improving the coordination and consistency between the reconstruction process and training operations.

[0149] Optional, also includes:

[0150] After each failed storage unit's data is successfully reconstructed to a preset backup disk, the reconstruction progress is updated.

[0151] For example, once the reconstruction of a failed storage unit is completed and it is confirmed that the data in that unit has been successfully written to the standby disk, the reconstruction progress is updated. Optionally, the progress percentage can be recalculated based on the number of completed failed storage units and the total number of failed storage units, and the latest progress percentage can be updated to reflect the latest reconstruction progress. During the update, the identifier, write timestamp, and verification result of the corresponding failed storage unit can also be recorded simultaneously, allowing subsequent tasks to continue execution within the completed range without having to repeatedly process successfully reconstructed units.

[0152] This example updates the reconstruction progress immediately after a single faulty storage unit is completed. This allows for flexible adjustments to the reconstruction rate based on the latest progress and enables accurate identification of completed and incomplete reconstruction ranges during interrupt recovery. This approach ensures consistency between the progress record and the actual write results, reducing the impact of repeated reconstructions and state deviations on the reconstruction task, thereby improving the controllability and continuity of the fault recovery process.

[0153] Optional, Figure 13 Flowchart of the data reconstruction method provided in the embodiments of this application Figure 10 Second, such as Figure 13 As shown, it also includes:

[0154] S1301. If a data write anomaly is detected during the process of reconstructing the data in the faulty storage unit to the backup disk, the reconstruction process of the data in the faulty storage unit is restarted.

[0155] In the example, during the process of writing data from a failed storage unit to a backup disk, the write and verification results are continuously monitored. If a write request fails, the backup disk times out, or the verification value after writing is inconsistent with the source data, it is determined that the data write is abnormal, and the current round of reconstruction task is terminated. Subsequently, the reconstruction task corresponding to the failed storage unit is recreated, the source address, target address, and completed offset context information are restored, and the reconstruction write is performed again.

[0156] S1302. If the number of restarts reaches the sixth preset threshold, the faulty storage unit is marked as requiring worker intervention.

[0157] For example, each restart increments the restart count and compares it to a sixth preset threshold. If the number of restarts does not exceed the sixth preset threshold, automatic reconstruction continues. If the number of restarts equals the sixth preset threshold, automatic reconstruction of the faulty storage unit stops, and a "requires manual intervention" flag is added to the faulty storage unit. This flag can trigger an alarm message sent to the cluster management platform, prompting relevant personnel to manually inspect or replace the faulty storage unit. This example, by performing a restart reconstruction when data write anomalies occur and switching to manual intervention after the number of restarts exceeds the threshold, enables automatic tiered handling of persistently failing reconstruction processes. This allows transient data write anomalies to continue to have a chance of recovery, while repeatedly failing storage units promptly enter the manual intervention path, thereby improving the certainty of fault handling and the stability of reconstruction control.

[0158] Figure 14 This is a schematic diagram of the overall process of the data reconstruction method provided in the embodiments of this application, as shown below. Figure 14 As shown, upon detecting a faulty disk, a reconstruction operation is initiated to move the data from the faulty disk to a backup disk. First, the input / output feature data of a preset training task is acquired, and the training phase is identified based on this data. If the input / output feature data meets the characteristics of a data loading phase, the target training phase is determined to be the data loading phase; if it meets the characteristics of a keypoint writing phase, the target training phase is determined to be the keypoint writing phase; otherwise, the target training phase is determined to be the training computation phase. Then, after determining the target training phase, the target baseline reconstruction rate is adjusted based on the target baseline reconstruction rate corresponding to the target training phase, combined with load status data and reconstruction progress, to obtain the target reconstruction rate. During the reconstruction of data on the faulty disk, if hot data exists on the faulty disk, the reconstruction task for that faulty disk is added to a delayed processing queue. The reconstruction task for the faulty disk is then executed only when the target training phase is in the training computation phase and the load status data indicates that resources are idle. If there is no hot data on the failed disk, the reconstruction task for that failed disk will be added to the main queue. The reconstruction tasks in the main queue can be processed immediately in sequence until the reconstruction tasks of all failed disks are completed, and then the standby disk will be started.

[0159] Based on the method provided in this embodiment, by obtaining the input / output feature data of the preset training task after detecting the faulty disk, and determining the target training stage of the training task, and then determining the target reconstruction rate based on the target training stage and reconstructing the data in the faulty disk based on the target reconstruction rate, the reconstruction process can actively avoid the key stages of the training process, ensuring sufficient training resources. Based on this, the resource conflict between reconstruction I / O and AI training I / O is reduced, thereby improving AI training efficiency.

[0160] Figure 15 This is a schematic diagram of the structure of the data reconstruction apparatus provided in the embodiments of this application, as shown below. Figure 15 As shown, it includes:

[0161] The acquisition module 151 is used to acquire the input / output feature data of a preset training task in response to the detection of a faulty disk;

[0162] The determination module 152 is used to determine the current target training stage of the training task based on the input / output feature data of the training task.

[0163] The determination module 152 is also used to determine the target reconstruction rate based on the target training phase;

[0164] Processing module 153 is used to reconstruct data in the faulty disk based on the target reconstruction rate.

[0165] The determination module 152 is specifically used to extract features from input / output feature data using a sliding window of a preset size, so as to obtain the distribution features of read / write requests, the spatial features of read / write access, and the activity of input / output corresponding to each sliding window;

[0166] The determination module 152 is further used to determine the training stage corresponding to each sliding window based on the distribution characteristics of read and write requests, the spatial characteristics of read and write access, and the activity of input / output corresponding to each sliding window.

[0167] The determination module 152 is further used to determine the target training stage of the training task based on the training stage corresponding to each sliding window.

[0168] The determination module 152 is further used to determine that for each sliding window, if the distribution characteristics of the read and write requests of the sliding window are such that the proportion of small file read requests is greater than or equal to the first preset threshold, the spatial characteristics of read and write access are such that random skipping reads, and the activity of input / output is greater than or equal to the second preset threshold, then the training stage corresponding to the sliding window is determined to be the data loading period, wherein a small file read request is defined as a read request involving a data amount less than or equal to the third preset threshold.

[0169] The determination module 152 is further used for each sliding window. If the distribution characteristics of the read and write requests of the sliding window are such that the proportion of large file write requests is greater than or equal to the first preset threshold, the spatial characteristics of read and write access are such that sequential access is present, and the activity of input / output is greater than or equal to the second preset threshold, then the training stage corresponding to the sliding window is determined to be the key point writing period. Among them, large file write requests are defined as write requests involving a data amount greater than or equal to the fourth preset threshold.

[0170] The determination module 152 is further used to determine that for each sliding window, if the distribution characteristics of the read and write requests, the spatial characteristics of the read and write access, and the activity of the input / output do not conform to the data characteristics of the data loading period or the data characteristics of the key point writing period, then the training phase corresponding to the sliding window is determined to be the training calculation period.

[0171] The determination module 152 is further used to determine the target training stage as the data loading stage if multiple consecutive sliding windows correspond to the data loading stage; if multiple consecutive sliding windows correspond to the keypoint writing stage; and if multiple consecutive sliding windows correspond to the training calculation stage.

[0172] The determination module 152 is further used to determine the target training stage as the data loading stage if the proportion of the training stage corresponding to each sliding window is greater than or equal to the fifth preset threshold; if the proportion of the training stage as the key point writing stage is greater than or equal to the fifth preset threshold, the target training stage is determined as the key point writing stage; if the proportion of the training stage as the training calculation stage is greater than or equal to the fifth preset threshold, the target training stage is determined as the training calculation stage.

[0173] The determination module 152 is also used to determine the target training phase as the training computation period in the initial stage of data reconstruction.

[0174] Module 152 is specifically used to determine the target baseline reconstruction rate corresponding to the target training phase.

[0175] Module 152 is specifically used to collect load status data and reconstruction progress.

[0176] The determination module 152 is further used to adjust the target baseline reconstruction rate based on load status data and reconstruction progress to obtain the target reconstruction rate.

[0177] The determination module 152 is further used to determine the first preset benchmark rate as the target benchmark reconstruction rate if the target training phase is the data loading phase.

[0178] The determination module 152 is further used to determine the second preset benchmark rate as the target benchmark reconstruction rate if the target training phase is the key point writing period, wherein the second preset benchmark rate is less than the first preset benchmark rate.

[0179] The determination module 152 is further used to determine the third preset benchmark rate as the target benchmark reconstruction rate if the target training phase is the training calculation period, wherein the third preset benchmark rate is greater than the first preset benchmark rate.

[0180] The determination module 152 is specifically used to collect real-time load status data and real-time reconstruction progress according to a preset collection interval.

[0181] The determination module 152 is further used to determine the first adjustment ratio based on real-time load status data;

[0182] Module 152 is specifically used to determine the second adjustment ratio based on the real-time reconstruction progress;

[0183] The determination module 152 is further used to determine the first adjustment weight corresponding to the load status data and the second adjustment weight corresponding to the reconstruction progress;

[0184] The determination module 152 is further used to perform a weighted summation of the first adjustment ratio and the second adjustment ratio according to the first adjustment weight and the second adjustment weight to obtain the real-time target adjustment ratio.

[0185] The determination module 152 is specifically used to adjust the target baseline reconstruction rate in real time according to the real-time target adjustment ratio to obtain the target reconstruction rate.

[0186] The determination module 152 is further used to adjust the target baseline reconstruction rate in real time according to the real-time target adjustment ratio, so as to obtain the adjusted real-time reconstruction rate.

[0187] The determination module 152 is further used to limit the adjusted real-time reconstruction rate according to a preset reconstruction rate range to obtain the target reconstruction rate.

[0188] The determination module 152 is specifically used to determine the upper limit and lower limit of the reconstruction rate range;

[0189] The determination module 152 is further used to determine the lower limit of the reconstruction rate as the target reconstruction rate if the adjusted real-time reconstruction rate is less than the lower limit of the reconstruction rate.

[0190] The determination module 152 is further used to determine the upper limit of the reconstruction rate as the target reconstruction rate if the adjusted real-time reconstruction rate is greater than the upper limit of the reconstruction rate.

[0191] The determination module 152 is further used to determine the adjusted real-time reconstruction rate as the target reconstruction rate if the adjusted real-time reconstruction rate is greater than or equal to the lower limit of the reconstruction rate and less than or equal to the upper limit of the reconstruction rate.

[0192] Processing module 153 is specifically used to identify faulty data in a faulty disk;

[0193] The processing module 153 is further configured to determine at least one faulty storage unit where the faulty data is located, and at least one normal storage unit other than the faulty storage unit.

[0194] The processing module 153 is also specifically used to reconstruct the data in each failed storage unit to a preset backup disk;

[0195] The processing module 153 is also used to copy data from each normal storage unit to a backup disk.

[0196] The processing module 153 is also used to generate data reconstruction tasks for each faulty storage unit and put each reconstruction task into the main queue in sequence.

[0197] The processing module 153 is also used to sequentially execute the reconstruction tasks in the main queue to reconstruct the data in each failed storage unit to the backup disk.

[0198] The processing module 153 is also used to detect whether there is hot data in each faulty storage unit;

[0199] The processing module 153 is also used to put the data reconstruction task of the faulty storage unit into the delayed processing queue if there is hot data in the faulty storage unit.

[0200] The processing module 153 is also used to execute the reconstruction task in the delayed processing queue when the target training phase is in the training computation period and the load status data representation resource is idle, so as to reconstruct the data in the faulty storage unit to the backup disk.

[0201] The processing module 153 is also used to update the reconstruction progress after each data in any failed storage unit is successfully reconstructed to a preset backup disk.

[0202] The processing module 153 is also used to restart the reconstruction process of the data in the faulty storage unit if a data write abnormality is detected during the process of reconstructing the data in the faulty storage unit to the backup disk.

[0203] The processing module 153 is also used to mark the faulty storage unit as requiring worker intervention if the number of restarts reaches a sixth preset threshold.

[0204] The description of the features of the data reconstruction device provided in this embodiment can be found in the relevant description of the data reconstruction method embodiment, and will not be repeated here.

[0205] Figure 16 A schematic diagram of the structure of the electronic device provided in this application. Figure 16 As shown, the electronic device 50 provided in this embodiment includes at least one processor 501 and a memory 502. Optionally, the electronic device 50 further includes a communication component 503. The processor 501, memory 502, and communication component 503 are connected via a bus.

[0206] In a specific implementation, at least one processor 501 executes computer execution instructions stored in memory 502, causing at least one processor 501 to execute the above-described data reconstruction method embodiment.

[0207] The specific implementation process of processor 501 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.

[0208] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.

[0209] The memory may include random access memory (RAM) and non-volatile memory (NVM), such as at least one disk storage device.

[0210] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.

[0211] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described data reconstruction method embodiments at runtime.

[0212] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), removable disk, magnetic disk, or optical disk.

[0213] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described data reconstruction method embodiments.

[0214] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described data reconstruction method embodiments.

[0215] It should be noted that the division of units is merely a logical functional division. In actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.

[0216] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0217] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0218] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to related technologies, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0219] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0220] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.

[0221] It should be understood that the above-described device embodiments are merely illustrative, and the device of this application can also be implemented in other ways. For example, the division of units / modules in the above embodiments is only a logical functional division, and there may be other division methods in actual implementation.

[0222] The foregoing has provided a detailed description of the data reconstruction method, electronic device, medium, and product provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and its core ideas. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A data reconstruction method, characterized in that, include: In response to the detection of a faulty disk, the input / output feature data of a preset training task are acquired; Based on the input / output feature data of the training task, determine the current target training stage of the training task; The target reconstruction rate is determined based on the target training phase. The data in the faulty disk is reconstructed based on the target reconstruction rate.

2. The method according to claim 1, characterized in that, Determining the current target training stage of the training task based on the input / output feature data of the training task includes: Using a sliding window of a preset size, feature extraction is performed on the input / output feature data to obtain the distribution characteristics of read / write requests, the spatial characteristics of read / write access, and the activity of input / output corresponding to each sliding window; Based on the distribution characteristics of read and write requests, spatial characteristics of read and write access, and input / output activity corresponding to each sliding window, the training stage corresponding to each sliding window is determined. Based on the training stage corresponding to each sliding window, determine the current target training stage of the training task.

3. The method according to claim 2, characterized in that, The step of determining the training phase corresponding to each sliding window based on the distribution characteristics of the read / write requests, the spatial characteristics of read / write access, and the activity of input / output for each sliding window includes: For each sliding window, if the distribution characteristics of the read and write requests of the sliding window are such that the proportion of small file read requests is greater than or equal to the first preset threshold, the spatial characteristics of read and write access are such that random skipping reads, and the activity of input / output is greater than or equal to the second preset threshold, then the training stage corresponding to the sliding window is determined to be the data loading period, wherein the small file read request is defined as the amount of data involved in the read request being less than or equal to the third preset threshold. For each sliding window, if the distribution characteristics of the read and write requests of the sliding window are such that the proportion of large file write requests is greater than or equal to the first preset threshold, the spatial characteristics of read and write access are such that sequential access is present, and the activity of input / output is greater than or equal to the second preset threshold, then the training stage corresponding to the sliding window is determined to be the key point writing period, wherein the large file write request is defined as the amount of data involved in the write request being greater than or equal to the fourth preset threshold. If the distribution characteristics of read / write requests, the spatial characteristics of read / write access, and the activity of input / output for each sliding window do not conform to the data characteristics of the data loading period or the data characteristics of the key point writing period, then the training phase corresponding to the sliding window is determined to be the training computation period.

4. The method according to claim 2, characterized in that, The step of determining the target training stage of the training task based on the training stage corresponding to each sliding window includes: Based on the training phase corresponding to each sliding window, if multiple consecutive sliding windows correspond to the data loading phase, then the target training phase is determined to be the data loading phase; if multiple consecutive sliding windows correspond to the keypoint writing phase, then the target training phase is determined to be the keypoint writing phase; if multiple consecutive sliding windows correspond to the training computation phase, then the target training phase is determined to be the training computation phase.

5. The method according to claim 2, characterized in that, The step of determining the target training stage of the training task based on the training stage corresponding to each sliding window includes: Based on the training stage corresponding to each sliding window, if the proportion of the training stage being the data loading stage is greater than or equal to the fifth preset threshold, then the target training stage is determined to be the data loading stage; if the proportion of the training stage being the keypoint writing stage is greater than or equal to the fifth preset threshold, then the target training stage is determined to be the keypoint writing stage; if the proportion of the training stage being the training calculation stage is greater than or equal to the fifth preset threshold, then the target training stage is determined to be the training calculation stage.

6. The method according to claim 4 or 5, characterized in that, Also includes: In the initial stage of data reconstruction, the target training stage is determined to be the training computation period.

7. The method according to claim 6, characterized in that, Determining the target reconstruction rate based on the target training phase includes: Determine the target baseline reconstruction rate corresponding to the target training phase; Collect load status data and reconstruction progress; Based on the load status data and reconstruction progress, the target baseline reconstruction rate is adjusted to obtain the target reconstruction rate.

8. The method according to claim 7, characterized in that, Determining the target baseline reconstruction rate corresponding to the target training phase includes: If the target training phase is the data loading phase, then the first preset benchmark rate is determined as the target benchmark reconstruction rate; If the target training phase is the key point writing phase, then the second preset benchmark rate is determined as the target benchmark reconstruction rate, wherein the second preset benchmark rate is less than the first preset benchmark rate. If the target training phase is the training computation period, then the third preset benchmark rate is determined as the target benchmark reconstruction rate, wherein the third preset benchmark rate is greater than the first preset benchmark rate.

9. The method according to claim 7, characterized in that, The collected load status data and reconstruction progress include: According to the preset collection interval, real-time load status data and real-time reconstruction progress are collected.

10. The method according to claim 9, characterized in that, The step of adjusting the target baseline reconstruction rate based on the load status data and reconstruction progress to obtain the target reconstruction rate includes: The first adjustment ratio is determined based on the real-time load status data; The second adjustment ratio is determined based on the real-time reconstruction progress; Determine the first adjustment weight corresponding to the load status data and the second adjustment weight corresponding to the reconstruction progress; Based on the first adjustment weight and the second adjustment weight, the first adjustment ratio and the second adjustment ratio are weighted and summed to obtain the real-time target adjustment ratio. The target baseline reconstruction rate is adjusted in real time according to the real-time target adjustment ratio to obtain the target reconstruction rate.

11. The method according to claim 10, characterized in that, The step of adjusting the target baseline reconstruction rate in real time according to the real-time target adjustment ratio to obtain the target reconstruction rate includes: The target baseline reconstruction rate is adjusted in real time according to the real-time target adjustment ratio to obtain the adjusted real-time reconstruction rate. The adjusted real-time reconstruction rate is limited according to a preset reconstruction rate range to obtain the target reconstruction rate.

12. The method according to claim 11, characterized in that, The step of limiting the adjusted real-time reconstruction rate according to a preset reconstruction rate range to obtain the target reconstruction rate includes: Determine the upper limit and lower limit of the reconstruction rate within the reconstruction rate range; If the adjusted real-time reconstruction rate is less than the lower limit of the reconstruction rate, then the lower limit of the reconstruction rate is determined as the target reconstruction rate; If the adjusted real-time reconstruction rate is greater than the upper limit of the reconstruction rate, then the upper limit of the reconstruction rate is determined as the target reconstruction rate; If the adjusted real-time reconstruction rate is greater than or equal to the lower limit of the reconstruction rate and less than or equal to the upper limit of the reconstruction rate, then the adjusted real-time reconstruction rate is determined as the target reconstruction rate.

13. The method according to claim 1, characterized in that, The process of reconstructing the data in the faulty disk based on the target reconstruction rate includes: Identify faulty data in the faulty disk; Identify at least one faulty storage unit where the faulty data is located, and at least one normal storage unit other than the faulty storage unit; The data in each of the faulty storage units is reconstructed to a preset backup disk; The data in each of the normal storage units is copied to the spare disk.

14. The method according to claim 13, characterized in that, The step of reconstructing the data in each of the faulty storage units to a preset backup disk includes: Generate data reconstruction tasks for each of the faulty storage units, and put each of the reconstruction tasks into the main queue in sequence; The reconstruction tasks in the main queue are executed sequentially to reconstruct the data in each of the faulty storage units to the backup disk.

15. The method according to claim 14, characterized in that, Also includes: Detect whether there is hot data in each of the faulty storage units; If there is hot data in the faulty storage unit, the data reconstruction task of the faulty storage unit is placed in the delayed processing queue. During the target training phase, when the training computation period is underway and the load status data representation resources are idle, the reconstruction task in the delayed processing queue is executed to reconstruct the data in the faulty storage unit to the backup disk.

16. The method according to claim 13, characterized in that, Also includes: After each failed storage unit's data is successfully reconstructed to a preset backup disk, the reconstruction progress is updated.

17. The method according to claim 13, characterized in that, Also includes: If a data write anomaly is detected during the process of reconstructing the data in the faulty storage unit to the backup disk, the reconstruction process of the data in the faulty storage unit is restarted. If the number of restarts reaches the sixth preset threshold, the faulty storage unit will be marked as requiring worker intervention.

18. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the data reconstruction method as described in any one of claims 1 to 17.

19. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the steps of the data reconstruction method as described in any one of claims 1 to 17.

20. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the steps of the data reconstruction method as described in any one of claims 1 to 17.