Batch multi-period interactive scheduling method, device and medium for data protection task
Patent Information
- Application Number
- CN202610953261.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2046-06-30
AI Technical Summary
1、自动排程虽降低负载,但本质为"黑盒"运行,操作员无法理解决策依据;当业务需求临时变更(如"今晚必须提前完成某批备份")时,缺乏实时交互调整手段;
(1)本发明通过构建“批量多时段交互调整—异构资源增量重算—时段级动态预测反馈”的联动架构,将实时交互操作作为预测模型的动态输入,使预测结果随交互调整同步更新;用户在可视化界面上执行批量跨时段交互调整时,系统即时反馈调整对整体排期的影响,突破了传统自动排程的黑盒不可解释性与手动调整无法预知全局影响的局限,实现了交互操作与预测模型的实时耦合;
Smart Images

Figure CN122470327B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of disaster recovery and backup, and relates to a batch multi-time interactive scheduling method, apparatus and computer storage medium for data protection tasks. Background Technology
[0002] Batch multi-timeframe processing refers to the process where operators perform batch migrations or adjustments on multiple tasks across two or more time windows within a data protection batch processing window. Interactive processing means that operators participate in scheduling decisions in real time through a visual interface, with the system providing immediate feedback on the impact of adjustments. Scheduling refers to the planning and orchestration of the start and end times of data protection tasks (including backup, recovery, verification, archiving, etc.) based on resource constraints, task priorities, and time dependencies.
[0003] As enterprise-level digital transformation deepens, the task scale of data protection platforms is rapidly evolving from the traditional hundreds to thousands and even tens of thousands, with a large number of heterogeneous tasks often being deployed intensively during off-peak business periods. Current industry demands are evolving from unattended black-box scheduling to collaborative scheduling models that allow for human intervention and have real-time feedback capabilities. This requires the system to be able to instantly reflect changes in resource contention and provide reliable duration predictions when operators intervene to make adjustments.
[0004] Currently, to alleviate the complexity of scheduling large-scale tasks, mainstream technologies are developing along two paths: One is the automatic scheduling path. The system automatically generates task execution plans based on a rule engine or heuristic algorithm to optimize global resource utilization. This path performs well in batch processing scenarios with stable loads, enabling basic global optimization and unattended operation.
[0005] Secondly, there is manual route adjustment. Some systems allow operators to adjust individual tasks to new time slots via a visual interface. This route provides basic flexibility and is suitable for rapid response to sudden demands.
[0006] In terms of duration prediction, existing technologies generally use a static formula of "data volume ÷ average speed". This formula is simple and intuitive, has low computational overhead, and is suitable for scenarios where the network environment and data types are relatively homogeneous.
[0007] However, when the above mechanism is applied to batch multi-time interactive scheduling scenarios, its inherent workflow exposes the following defects: 1. Although automatic scheduling reduces the load, it is essentially a "black box" operation, and operators cannot understand the decision-making basis; when business needs change temporarily (such as "a certain batch of backups must be completed ahead of schedule tonight"), there is a lack of real-time interactive adjustment means. 2. Manual adjustment generally only supports isolated operations for a single task and a single time period. Although the system can provide simple predictions based on the attributes of a single task, this prediction mechanism cannot be adapted to batch multi-time period scenarios through linear expansion, and operators have difficulty knowing the true consequences of batch adjustments. 3. The static prediction formula does not take into account differences in data types, time-varying fluctuations in network bandwidth, and storage I / O bottlenecks. It has significant prediction deviations in weak network or I / O-intensive scenarios and cannot effectively guide peak-shifting scheduling.
[0008] More importantly, due to the technical inertia of the prediction layer and the scheduling layer being isolated from each other in the architecture, existing improvements are usually limited to offline optimization of the prediction algorithm or independent improvement of the interactive interface, without establishing a real-time feedback channel from the interactive operation to the prediction model at the architecture level. In batch multi-time scenarios, this isolation makes it impossible for the prediction model to obtain information on the dynamic changes of heterogeneous resources generated by the interactive operation, and the time-level prediction results are difficult to synchronize with the actual scheduling status.
[0009] In summary, in batch multi-time period interactive scheduling scenarios, how to enable operators to know the real impact of cross-time period batch interactive adjustments on the overall schedule in real time, and to obtain accurate duration predictions based on the dynamic changes in time period-level resource contention, is an urgent technical problem to be solved. Summary of the Invention
[0010] In order to solve the technical problems in the background art, the present invention provides a batch multi-time interactive scheduling method, apparatus and computer storage medium for data protection tasks.
[0011] The technical solution of the present invention to solve the above-mentioned technical problems is as follows: Firstly, a batch multi-time interactive scheduling method for data protection tasks is provided, the method comprising the steps of: The visualization interface presents multiple data protection task blocks and their corresponding time periods. In response to batch task configuration adjustments across multiple time periods, determine the set of affected time periods; Establish a nonlinear mapping relationship between heterogeneous task attributes, data volume and multidimensional resource consumption, characterize the resource consumption characteristics of heterogeneous tasks by preset coefficients, and determine the predicted value of resource consumption for each task in each resource dimension. Based on the set of affected time periods, each recalculation period is determined. Then, based on the predicted resource occupancy of each task in each resource dimension, the resource contention level of each recalculation period is determined. Based on the resource contention level, the incremental recalculation factor corresponding to each recalculation period is determined, forming an incremental recalculation factor vector. Based on the incremental recalculation of the factor vector, the estimated completion time of each task is updated, and the feedback is rendered in real time on the visualization interface.
[0012] Secondly, a batch multi-time interactive scheduling device for data protection tasks is provided, the device comprising: The task presentation module is used to present multiple data protection task blocks and their corresponding time period distribution on a visual interface; The batch multi-time period interaction module is used to respond to batch task multi-time period configuration adjustments and determine the set of affected time periods; The heterogeneous resource mapping module is used to establish a non-linear mapping relationship between heterogeneous task attributes, data volume and multi-dimensional resource consumption. It uses preset coefficients to characterize the resource consumption characteristics of heterogeneous tasks and determine the predicted value of each task's resource consumption in each dimension. The incremental recalculation module is used to determine each recalculation period based on the set of affected periods, then determine the resource contention level of each recalculation period based on the predicted resource occupancy of each task in each resource dimension, and determine the incremental recalculation factor corresponding to each recalculation period based on the resource contention level, forming an incremental recalculation factor vector. The prediction feedback module is used to update the estimated completion time of each task based on the incremental recalculation of the factor vector, and to render the feedback in real time on the visualization interface.
[0013] Thirdly, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements a batch multi-time interactive scheduling method for data protection tasks as described above.
[0014] The beneficial effects of this invention are: (1) This invention constructs a linkage architecture of “batch multi-period interactive adjustment - heterogeneous resource incremental recalculation - period-level dynamic prediction feedback”, and uses real-time interactive operation as the dynamic input of the prediction model, so that the prediction results are updated synchronously with the interactive adjustment; when the user performs batch cross-period interactive adjustment on the visual interface, the system provides real-time feedback on the impact of the adjustment on the overall schedule, which breaks through the limitations of the black box unexplainability of traditional automatic scheduling and the inability to predict the global impact of manual adjustment, and realizes the real-time coupling of interactive operation and prediction model; (2) This invention combines the nonlinear mapping relationship between heterogeneous task attributes, data volume and multidimensional resource occupation, and characterizes the resource consumption characteristics of heterogeneous tasks by setting a mapping coefficient. It also introduces an incremental recalculation factor to couple the dynamic changes of resource contention into the calculation process of the expected completion time. This enables the resource occupation prediction values under different module types and task types to be calculated differently, overcoming the defect that the traditional static prediction formula cannot reflect the dynamic changes of heterogeneous resources, and avoiding the distortion of resource contention assessment caused by linear superposition of data volume. (3) The present invention generates resource contention information based on the resource contention level of each recalculation period, and actively avoids task accumulation during periods of high resource contention based on the incremental recalculation factor vector; when the operator performs batch drag-and-drop adjustment across time periods, the operator can observe the changing trend of the expected completion time of each time period in real time, predict and avoid the risk of global scheduling collapse caused by local adjustment, and form a virtuous cycle of "prediction → scheduling → actual convergence". Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a schematic diagram of the batch multi-time interactive scheduling method for data protection tasks provided in Embodiment 1 of the present invention.
[0017] Figure 2 This is a schematic diagram of the overall architecture of batch multi-time interactive scheduling for data protection tasks provided in Embodiment 1 of the present invention.
[0018] Figure 3 This is a schematic diagram of the batch multi-time interactive scheduling method for data protection tasks provided in Embodiment 2 of the present invention.
[0019] Figure 4 This is a schematic diagram of the batch multi-time interactive scheduling method for data protection tasks provided in Embodiment 3 of the present invention.
[0020] Figure 5 This is a schematic diagram of time-duration coupling correction provided in Embodiment 3 of the present invention.
[0021] Figure 6 This is a schematic diagram of the incremental recalculation factor vector and task mapping provided in Embodiment 3 of the present invention.
[0022] Figure 7 This is a schematic diagram of the structure of the batch multi-time interactive scheduling device for data protection tasks provided in Embodiment 4 of the present invention.
[0023] Figure 8 This is a schematic diagram of the incremental recalculation module structure provided in Embodiment 4 of the present invention.
[0024] Figure 9 This is a schematic diagram of the predictive feedback module structure provided in Embodiment 4 of the present invention.
[0025] The attached diagram lists the components represented by each number as follows: 4001. Task Presentation Module; 4002. Batch Multi-Time Interaction Module; 4003. Heterogeneous Resource Mapping Module; 4004. Incremental Recalculation Module; 4005. Prediction Feedback Module; 40041. Buffer Range Determination Unit; 40042. Resource Contention Level Calculation Unit; 40043. Incremental Recalculation Factor Calculation Unit; 40051. Task Factor Determination Unit; 40052. Completion Time Prediction Unit; 40053. Rendering Feedback Unit. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0027] Example 1
[0028] Currently, to alleviate the complexity of scheduling large-scale tasks, mainstream technologies are developing along two paths: One is the automatic scheduling path. The system automatically generates task execution plans based on a rule engine or heuristic algorithm to optimize global resource utilization. This path performs well in batch processing scenarios with stable loads, enabling basic global optimization and unattended operation.
[0029] Secondly, there is manual route adjustment. Some systems allow operators to adjust individual tasks to new time slots via a visual interface. This route provides basic flexibility and is suitable for rapid response to sudden demands.
[0030] In terms of duration prediction, existing technologies generally use a static formula of "data volume ÷ average speed". This formula is simple and intuitive, has low computational overhead, and is suitable for scenarios where the network environment and data types are relatively homogeneous.
[0031] However, when the above mechanism is applied to batch multi-time interactive scheduling scenarios, its inherent workflow exposes the following defects: 1. Although automatic scheduling reduces the load, it is essentially a "black box" operation, and operators cannot understand the decision-making basis; when business needs change temporarily (such as "a certain batch of backups must be completed ahead of schedule tonight"), there is a lack of real-time interactive adjustment means. 2. Manual adjustment generally only supports isolated operations for a single task and a single time period. Although the system can provide simple predictions based on the attributes of a single task, this prediction mechanism cannot be adapted to batch multi-time period scenarios through linear expansion, and operators have difficulty knowing the true consequences of batch adjustments. 3. The static prediction formula does not take into account differences in data types, time-varying fluctuations in network bandwidth, and storage I / O bottlenecks. It has significant prediction deviations in weak network or I / O-intensive scenarios and cannot effectively guide peak-shifting scheduling.
[0032] In summary, in batch multi-time period interactive scheduling scenarios, how to enable operators to know the real impact of cross-time period batch interactive adjustments on the overall schedule in real time, and to obtain accurate duration predictions based on the dynamic changes in time period-level resource contention, is an urgent technical problem to be solved.
[0033] To address the aforementioned issues, this embodiment provides a batch, multi-period interactive scheduling method for data protection tasks. Figure 1 This is a flowchart illustrating the batch multi-time interactive scheduling method for data protection tasks provided in this embodiment. Figure 2 The schematic diagram shows the overall architecture of the batch multi-time interactive scheduling of data protection tasks provided in this embodiment. Figure 2 As shown, the overall architecture comprises a four-layer coupled structure: an interaction layer, a computation layer, a model layer, and a feedback layer. The interaction layer receives batch, multi-period interactive adjustment operations from the user and presents them on the visual interface. The computation layer performs heterogeneous resource mapping processing. The model layer handles the determination of each recalculation period and the calculation of incremental recalculation factor vectors. The feedback layer updates the estimated completion time, performs real-time rendering, and outputs resource contention alarms. Figure 1 As shown, this method includes: Step S101: Present multiple data protection task blocks and their corresponding time period distributions on the visualization interface.
[0034] This step relies on a visual interface to complete the visualization rendering of batch tasks and time-series information. The interface is divided into sections to display the entities of each data protection task block and their respective configured start and end time periods. Users can intuitively select a single task or multiple tasks and initiate batch time period adjustment operations.
[0035] It is understandable that the task block is a single data protection task carrier after logical packaging, and the time period distribution refers to two types of time-series data: the source configuration time period and the target change time period. This embodiment does not limit the concrete implementation methods such as interface layout and rendering controls, but only limits the presentation content and interaction basis.
[0036] This step serves as the preliminary data entry point for subsequent batch adjustments and affected time period filtering. It is essential to ensure that the information displayed on the interface maps one-to-one with the original parameters of the backend task to avoid a disconnect between the time periods displayed on the front end and the actual configuration data at the underlying level, thus preventing baseline deviations in the calculation of the affected time period set in subsequent steps.
[0037] Step S102: Respond to the batch task multi-time period configuration adjustment and determine the set of affected time periods.
[0038] This step responds to user-initiated time-period interaction adjustments for task blocks, calculates and outputs the set of affected time periods. Interaction adjustments include two typical implementation scenarios: single-task single-time-period adjustments and batch task multi-time-period adjustments. These two scenarios differentiate the task adjustment logic based on the scale of the change, serving as the implementation of different application scenarios for this step. All subsequent resource mapping and incremental recalculation use the set of affected time periods output by this step as the calculation benchmark. The two typical implementation scenarios are detailed below: (1) Single task single time period adjustment scenario When a user completes a migration operation from the source time period to the target time period for a single task, only the source time period and the target time period corresponding to the change are included in the affected time period. For example, if a single file backup task is adjusted from 1 o'clock to 3 o'clock, then 1 o'clock and 3 o'clock are the affected time periods. (2) Batch task multi-time period adjustment scenario For multi-task cross-time period migration, in addition to the source time period and target time period of each task, the time period during which the migration passes is also included in the affected time period. For example, if multiple backup tasks are moved in batches from the 2-hour interval to the 6-hour interval, the time periods during the 3-hour, 4-hour, and 5-hour intervals are all included in the statistics.
[0039] In this step, the core difference between the two types of scenarios lies in the scope of the affected time periods: single-task changes only include the two time periods before and after the change, while batch cross-time period changes additionally include the migration path time periods.
[0040] It is important to note that in practice, two key concepts need to be clarified: the affected period and the recalculation period. The affected period is the original period directly related to the task change, while the recalculation period is the period that actually participates in resource accounting after the affected period is superimposed on the buffer zone. The two have different definitions and should not be used interchangeably.
[0041] Although there are subtle differences in the rules for collecting affected time periods between the two scenarios, the scenario branches are only internal differences within this step, and the step ultimately outputs a standardized set of affected time periods. Based on the unified output, the entire processing framework for subsequent task resource mapping, incremental recalculation of factor vectors, and completion time prediction remains universal and unchanged, thus enabling a single architecture to be compatible with both single-point fine-tuning and large-scale cross-time period changes.
[0042] Step S103: Establish a nonlinear mapping relationship between heterogeneous task attributes, data volume and multidimensional resource consumption. Use preset coefficients to characterize the resource consumption characteristics of heterogeneous tasks and determine the predicted resource consumption value of each task in each resource dimension.
[0043] This step is also applicable to both single-task and batch-task pre-adjustment scenarios. The only difference between the two scenarios is the number of task inputs; the mapping calculation logic completely reuses the same architecture. The two typical implementation scenarios are as follows: (1) Single task single time period adjustment scenario The heterogeneous attributes and data volume of a single task are extracted and fed into the mapping model. The corresponding preset coefficients are matched to obtain the predicted value of the multidimensional resource occupancy of the task.
[0044] (2) Batch task multi-time period adjustment scenario Extract attribute parameters from all changed tasks in batches, and perform resource calculations for each task using the same mapping rules. Regardless of whether a single task or a batch of multiple task data is input, the mapping formula and preset coefficient matching rules remain consistent. The coefficients are pre-calibrated, generating differentiated resource results for the same data volume under different task attributes.
[0045] Step S104: Determine each recalculation period based on the set of affected periods, then determine the resource contention level of each recalculation period based on the predicted resource occupancy values of each task in each resource dimension, and determine the incremental recalculation factor corresponding to each recalculation period based on the resource contention level, forming an incremental recalculation factor vector.
[0046] This step determines the recalculation periods based on the affected time period set. As mentioned above, the recalculation period is not simply equivalent to the source period, target period, or transit period. The recalculation period is the actual calculation unit obtained by dynamically expanding the buffer range forward and backward according to the time period span and the data volume of each task. For example, when an operator drags a task from 20:00–22:00 to 19:00–21:00, the affected time period set includes the source period and the target period. The recalculation period may be extended forward and backward by one hour on this basis, forming a calculation window of 18:00–23:00 to ensure that resource fluctuations in the edge period are fully captured.
[0047] Furthermore, this step is implemented in two typical scenarios as follows: (1) Single task single time period adjustment scenario The predicted resource usage of individual tasks within a recalculation period is accumulated and combined with the baseline resources already occupied by other tasks within that period to determine the resource contention level for that period, and thus determine the corresponding incremental recalculation factor. In this scenario, the computational chain is simple, and the incremental recalculation factor only reflects the impact of a single task's migration in or out on local resource usage.
[0048] (2) Batch task multi-time period adjustment scenario Multiple tasks of different types migrate simultaneously across time periods. The predicted resource usage values of all tasks within each recalculation time period are comprehensively processed, and the resource contention level of each time period is determined in conjunction with baseline resources. This yields the incremental recalculation factor for each time period, forming an incremental recalculation factor vector. At this time, each time period may simultaneously carry multiple types of backup tasks, and its resource contention level needs to reflect the overall resource usage situation. The incremental recalculation factors of each time period are arranged in time period order to form a vector so that subsequent tasks can be mapped to the component corresponding to their respective time period.
[0049] Understandably, resource contention level is a comprehensive measure at the time-period level, reflecting the overall resource occupancy status of all tasks within that time period; the incremental recalculation factor is a time-period correction coefficient determined based on the resource contention level; and the incremental recalculation factor vector is a set formed by arranging the incremental recalculation factors corresponding to each recalculation time period in time-period order. The three constitute a progressive relationship of "time-period measurement → time-period correction → vector set," ensuring that the differences in resource contention in each time period are independently characterized.
[0050] Step S105: Based on the incremental recalculation of the factor vector, update the estimated completion time of each task and render the feedback in real time on the visualization interface.
[0051] Understandably, in this step, the incremental recalculation factor vector is a time-level output. Each task needs to be mapped to the component corresponding to its time period to correct the basic prediction time, obtain the updated expected completion time, and render it in real time on the interface.
[0052] Furthermore, this step is implemented in two typical scenarios as follows: (1) Single task single time period adjustment scenario Since it only involves the migration of a single task between two time periods, the incremental recalculation factor vector may only contain a small number of components. Directly map the task to the incremental recalculation factor of the corresponding time period to complete the update of the expected completion time, and render the latest scheduling information of the task on the interface in real time.
[0053] (2) Batch task multi-time period adjustment scenario The incremental recalculation factor vector contains components corresponding to multiple recalculation time periods. Each task needs to be mapped to the component corresponding to its respective time period, and the estimated completion time of each task is updated in batches. For example, when tasks A and B are in the 20:00–21:00 time period, and task C is in the 21:00–22:00 time period, tasks A and B are mapped to the incremental recalculation factor component corresponding to 20:00–21:00, and task C is mapped to the component corresponding to 21:00–22:00, thus correcting the estimated completion time. Simultaneously, time-level resource status information is rendered in real-time on the visualization interface, allowing operators to intuitively perceive the impact of batch adjustments on the overall scheduling.
[0054] It is important to note that the above-mentioned update process for estimated completion time follows the mapping logic from task to time period component and from time period component to correction factor in both scenarios. The specific correction process and visualization rendering sequence can be adjusted according to the actual system configuration, such as specifically selecting to render a time period-level resource utilization heatmap and alarm indicators. In addition, real-time rendering emphasizes millisecond-level response after the user releases an interactive operation, which should be distinguished from the second-level or minute-level latency required for full recalculation.
[0055] This embodiment constructs a linked architecture of "batch multi-period interactive adjustment—heterogeneous resource incremental recalculation—period-level dynamic prediction feedback," using real-time interactive operations as dynamic input to the prediction model, ensuring that prediction results are updated synchronously with interactive adjustments. When users perform batch cross-period interactive adjustments on the visual interface, the system provides immediate feedback on the impact of the adjustments on the overall schedule, overcoming the limitations of the black-box uninterpretability of traditional automatic scheduling and the inability to predict the global impact of manual adjustments, thus achieving real-time coupling between interactive operations and the prediction model. Simultaneously, combining the nonlinear mapping relationship between task heterogeneous attributes, data volume, and multi-dimensional resource consumption, a preset mapping coefficient characterizes the resource consumption characteristics of heterogeneous tasks, and an incremental recalculation factor is introduced to couple the dynamic changes in resource contention into the calculation process of the estimated completion time. This allows for differentiated calculation of resource contention prediction values under different module types and task types, overcoming the shortcomings of traditional static prediction formulas that cannot reflect the dynamic changes of heterogeneous resources, and avoiding the distortion in resource contention assessment caused by linear superposition of data volume.
[0056] Example 2
[0057] The above embodiments provide the overall process of interactive scheduling, mentioning that the predicted resource usage of each task in each dimension is determined based on the heterogeneous attributes and data volume of each task. However, if only data volume is used as the sole metric for linear calculation, incremental database backups and file backups will be considered to have equal resource usage under the same data volume, failing to reflect the essential differences in resource usage between the two, leading to a distortion in the assessment of resource contention levels during a given period. This embodiment, based on the aforementioned embodiments, refines the heterogeneous resource usage prediction in incremental recalculation.
[0058] like Figure 3 As shown in this embodiment, a batch multi-time interactive scheduling method for data protection tasks is provided. This method includes: Step S201: Present multiple data protection task blocks and the time period distribution corresponding to the data protection task blocks on the visualization interface; Step S202: Respond to the batch task multi-time period configuration adjustment, determine the set of affected time periods, which includes the source time period, target time period, and transit time period of the batch multi-time period interaction adjustment; Step S203: Based on the module type, task type and backup data volume of each task, establish a nonlinear mapping relationship between each task and storage I / O, network bandwidth and computing resources to determine the predicted resource usage value of each task in each resource dimension; the nonlinear mapping relationship is realized by preset mapping coefficients, which are associated with module type and task type, and different predicted resource usage values correspond to different module types or different task types. Step S204: Determine each recalculation period based on the set of affected periods, then determine the resource contention level of each recalculation period based on the predicted resource occupancy values of each task in each resource dimension, and determine the incremental recalculation factor corresponding to each recalculation period based on the resource contention level to form an incremental recalculation factor vector. Step S205: Based on the incremental recalculation of the factor vector, update the estimated completion time of each task and render the feedback in real time on the visualization interface.
[0059] In step S202 above, when responding to task block drag-and-drop interaction adjustments, the user's visual drag-and-drop behavior can be structured and analyzed, and the intuitive drag-and-drop actions on the interface can be decomposed into four types of structured spatiotemporal data: source time period set, target time period set, transit time period, and change task list, so as to realize the conversion of interface operation into backend computable spatiotemporal parameters.
[0060] By defining batch drag-and-drop as a holistic spatiotemporal transformation, rather than multiple scattered single-task modifications, a global correspondence between the source and target time periods is established. The advantage of this processing logic is that batch changes no longer fragment the calculation time periods for each task, but uniformly identify the affected and transit time periods from the overall spatiotemporal scope. This not only prevents the omission of transit time periods when batches cross gap time periods, but also allows for a unified set of affected time periods in both single-task and batch task scenarios, ensuring that subsequent resource mapping and recalculation factor calculations share a common processing framework.
[0061] It is worth noting that the rules for distinguishing transit time periods can be understood as follows: if there is no gap between the source time period and the target time period, there is no transit time period; if there is an unoccupied gap between the source and the target, the gap is the transit time period and is included in the affected time period.
[0062] For easier understanding, the specific rules for determining the transit time period are shown below: Example 1 for determining transit time periods: The source time periods are 20:00–22:00 and 23:00–24:00, and the target time periods are 19:00–21:00 and 21:00–22:00. The source and target time periods overlap and the gaps between them are completely covered. The entire relocation of the task does not cross the empty time periods between the two periods, so there are no transit time periods. The dragged task A (database incremental), task B (virtual machine whole machine), and task C (file backup) are synchronously bound. Based on the summary of the entire set of structured data, the affected time period set is generated to lock the input boundary for subsequent resource recalculation.
[0063] Example 2 for determining transit time periods: First, the source time period is 20:00–22:00, the target time period is 24:00–26:00, and the blank time period in between, 22:00–24:00, is the transit time period; Second, the batch source time periods are 10:00–12:00 and 15:00–17:00, and the whole batch is dragged to 13:00–14:00 and 18:00–20:00. The three gaps between the source and the target, 12:00–13:00, 14:00–15:00, and 17:00–18:00, are all included in the transit time period.
[0064] In step S203 above, the heterogeneous resource occupancy mapping step receives the affected time period and batch change task information output from S202 as input, based on module type M. i Task type P i Backup data volume D i The three core features can be supplemented by auxiliary features such as task execution priority and data source storage architecture as needed to construct a quantitative mapping relationship between task attributes and three-dimensional resource usage. The core calculation expression is: f R (M i ,P i D i )=D i ·k R (M i ,P i The symbols in the formula have the following meanings: M i Pi represents the module type of the i-th task (e.g., database agent, virtual machine agent, file agent, operating system agent, etc.). P i : The task type of the i-th task (e.g., full backup, incremental backup, differential backup, copy backup, recovery, verification, etc.); D i : The size of the backup data for the i-th task (in GB or MB); R∈{I, B, C}: resource dimension, corresponding to I (storage I / O), B (network bandwidth), and C (CPU / memory computing resources), respectively. k R (M) i P i ): Module type M i With task type P i The preset mapping coefficients (dimensionless) for the resource dimension R are obtained through fitting historical task logs or expert calibration. For example, K for incremental database backups. I Far exceeding the K of file backups and virtual machine full-system backups B K C At the same time, it is too high; f R (M)i P i D i ): The predicted resource usage of the i-th task in the resource dimension R.
[0065] This mapping model is uniformly adaptable to two scenarios: single-task fine-tuning and batch multi-task cross-time drag-and-drop. A single set of coefficient matching and calculation rules is used across all scenarios. R includes three dimensions: I (storage I / O), B (network bandwidth), and C (computing resources); the coefficient k... R (M i ,P i ) by M i P i It was jointly decided that all mapping coefficients would be generated based on fitting historical operation logs or calibration using industry expert experience. Regarding parameter definition, M... i This refers to the module type to which tasks such as database proxy, virtual machine proxy, and file proxy belong, P i For incremental backup, full backup, data recovery and other business types, D i This represents the actual backup data capacity for a single task. Different task types paired with the same module, or different modules paired with the same task, will all have independent coefficients configured under the same resource dimension.
[0066] To facilitate understanding, a complete calculation example of three batch tasks, A, B, and C, is provided: For each dragged task within the window, calculate its predicted occupancy of 3D resources: f R (M i ,P i D i )=D i ·k R (M i ,P i ), Preset mapping coefficient k R The following (fitted from historical logs): Table 1. Heterogeneous Resource Mapping Parameter Table A Database Proxy Incremental backup 500 0.8 0.1 0.3 B Virtual machine agent Full backup 200 0.4 0.6 0.5 C File Proxy Full backup 800 0.1 0.7 0.2 Calculate the resource usage of each task: Task A: f I =500 × 0.8 = 400, f B =500 × 0.1 = 50、f C =500 × 0.3 = 150; Task B: f I =200×0.4=80、f B =200 × 0.6 = 120, f C =200 × 0.5 = 100; Task C: f I=800 × 0.1 = 80, f B =800×0.7=560、f C =800×0.2=160.
[0067] The underlying principle of this step stems from the different resource access logics of heterogeneous tasks: incremental database backups involve a large amount of random disk reads and writes, resulting in a higher mapping coefficient for the IO dimension; full backups of virtual machines involve continuous streaming data transmission, leading to higher bandwidth and computing power coefficients; and backups of ordinary files have low disk overhead but high network overhead. Simply adding resources based solely on data volume, ignoring differences in modules and task types, would significantly underestimate or overestimate the actual resource load during a given period. This step relies on differentiated calibration coefficients to eliminate measurement biases caused by heterogeneity, achieving accurate conversion from a single data volume to three-dimensional indicators of IO, bandwidth, and computing power. Of course, in subsequent deployment scenarios, dimensions such as snapshot usage and metadata retrieval overhead can be added to the calculation simultaneously; this embodiment does not impose specific limitations.
[0068] The resource usage results for each task in this step are used as the standard input for subsequent incremental recalculation factors. Regardless of whether the front end involves dragging and dropping a single task or batch shifting multiple tasks, the mapping coefficient query logic and calculation formula remain unchanged. Relying on a general model, it is compatible with all business scenarios, ensuring both the accuracy of heterogeneous task resource measurement and maintaining a unified and universal architecture for the entire solution.
[0069] Unlike the above embodiments, this implementation establishes a nonlinear mapping relationship between heterogeneous task attributes and multidimensional resource consumption, and uses preset mapping coefficients to characterize the differentiated consumption characteristics of different module types and task types on each resource dimension. This enables the acquisition of differentiated resource consumption prediction values under the same data volume, avoiding the evaluation distortion caused by linear superposition of data volume, thereby providing an accurate input basis for subsequent incremental recalculation factor calculation.
[0070] Example 3 The above embodiments establish heterogeneous resource mapping relationships, enabling the acquisition of differentiated resource occupancy predictions. However, in scenarios involving batch multi-period interactive adjustments, if a fixed buffer window is used for incremental recalculation, it becomes difficult to adapt to the actual impact radius of different batch tasks. If the window is too small, it is easy to miss cross-period chain effects, while if the window is too large, invalid calculations will occur. At the same time, the existing prediction model and interactive operations are independent of each other, and resource contention between time periods is calculated in a fragmented manner, which cannot reflect the true scheduling impact after batch adjustments.
[0071] like Figure 4 As shown, in one embodiment, a batch multi-time interactive scheduling method for data protection tasks is provided. This method includes: Step S301: Present multiple data protection task blocks and the time period distribution corresponding to the data protection task blocks on the visualization interface; Step S302: In response to the batch task multi-time period configuration adjustment, determine the set of affected time periods; Step S303: Establish a nonlinear mapping relationship between heterogeneous task attributes, data volume and multidimensional resource consumption, characterize the resource consumption characteristics of heterogeneous tasks through preset coefficients, and determine the predicted value of resource consumption for each task in each resource dimension. Step S304: Based on the time span of the affected time period set and the data volume of each task, dynamically expand the buffer range, determine the incremental recalculation range of each affected time period, and obtain each recalculation time period; Step S305: For each recalculation period, the predicted resource usage values of each task in each resource dimension are accumulated, and combined with the baseline resources, the resource contention level for each recalculation period is generated. Step S306: Based on the resource contention level, preset weight coefficient and resource contention threshold, perform weighted penalty correction on each recalculation period to obtain the incremental recalculation factor corresponding to each recalculation period, and form an incremental recalculation factor vector. Step S307: Map each task to the corresponding component of the incremental recalculation factor vector to obtain the incremental recalculation factor for each task. Step S308: Perform a product correction based on the base time, queuing penalty factor, and incremental recalculation factor of each task to determine the estimated completion time of each task. Step S309: Render the estimated completion time in real time on the visualization interface to provide feedback on the impact of batch multi-period interactive adjustments on the overall schedule.
[0072] like Figure 5 As shown, in step S305, based on the time span of the affected time period set and the data volume of each task, the buffer range is dynamically expanded forward and backward to determine the incremental recalculation range of each affected time period, thus obtaining each recalculation time period; wherein, the affected time period set includes the source time period, target time period, and transit time period of the batch multi-time period interactive adjustment, the source time period is the task migration out time period, the target time period is the task migration in time period, and the transit time period is the continuous time period in the batch multi-time period interactive adjustment path where the resource status changes.
[0073] The following combination Figure 5 A specific example further illustrates this: The operator drags a batch of tasks A, B, and C from their original location, i.e., the source time period (20:00–22:00), to the target time period (19:00–21:00). Tasks A, B, and C are moved out of the source time period and into the target time period. The system identifies the source time period, the target time period, and the time periods between them for this batch adjustment. Based on the time period span and the data volume of each task, the system dynamically expands the buffer range and determines the recalculation time period (18:00–23:00) to fully capture the cross-time period resource disturbances caused by the task moving out, moving in, and passing through, and to avoid the chain reaction caused by a fixed window, resulting in omissions or invalid calculations.
[0074] It is important to note that in batch, multi-time-period scenarios, each time period is not an isolated resource island. When resource contention exceeds capacity due to task migration in a certain time period, some tasks will overflow into adjacent time periods, creating a chain reaction. If resource contention is still calculated independently for each time period, the boundary time periods will exhibit significant evaluation distortion because the upstream overflow is not taken into account.
[0075] Furthermore, by introducing time-period coupling correction, step S305 further includes: Step S3051: For each recalculation period, determine whether its resource contention ratio exceeds a preset threshold. If it does not exceed the threshold, determine that the resource overflow amount transmitted from the recalculation period to the adjacent period is zero; if it exceeds the threshold, determine that the resource overflow amount transmitted from the recalculation period to the adjacent period due to overload. Step S3052: Determine the coupling correction amount for adjacent time periods based on the resource overflow amount and the preset coupling coefficient, and correct the resource contention situation for adjacent time periods based on the coupling correction amount.
[0076] Specifically, the coupling correction amount is determined by the resource overflow amount and the preset coupling coefficient, and the calculation satisfies the relationship: Couple(s) j )=β·[Overflow(s j-1 )+Overflow(s j+1 The symbols in the formula have the following meanings: Couple(s j ): Overflow pairs between adjacent time periods s j Coupling correction amount; β: Coupling coefficient (0-1), reflecting the overflow propagation efficiency of the scheduling engine. If the engine is strictly isolated by time period (without delay), then β=0; Overflow(s j-1 (Time period s) j-1 Overflowing backwards to s due to overload j The task resource consumption; if s j-1 If the resource contention ratio is ≤1, then Overflow = 0; Overflow(s j+1 (Time period s) j+1 Overflowing forward to s due to overload j The task resource consumption; if s j+1 If the resource contention ratio is ≤1, then Overflow = 0.
[0077] Overflow represents the amount of resource consumption that is passed from one time period to the next due to overload. If the resource contention ratio does not exceed 1, then Overflow is 0. β is the coupling coefficient (0-1), which reflects the overflow propagation efficiency of the scheduling engine. If the engine strictly isolates time periods, then β is 0.
[0078] For example, suppose the original IO load quota for the 20:00–21:00 period is 80. Due to the migration of task B, severe IO contention occurs, and 20% of the IO load overflows into the 21:00–22:00 period. Then, the Overflow... IO (s j =80 × 20% = 16; If no overflow occurs in the bandwidth and CPU dimensions during the 20:00–21:00 period, then the corresponding overflow in that dimension is 0. Let β be 0.5, then the coupling correction obtained in the IO dimension during the 21:00–22:00 period is Couple IO (s j =0.5×(16+0)=8. By including this coupling correction in the resource baseline for the 21:00–22:00 period, the cross-period chain effect can be accurately captured, the overflow of adjacent periods can be included in the resource baseline of this period, the global resource situation after batch adjustment can be truly reflected, and the distortion of resource contention assessment due to the failure to include upstream overflow in boundary periods can be avoided.
[0079] It is important to note that in batch, multi-time-period scenarios, the same time period may simultaneously handle heterogeneous tasks such as incremental database backups (IO-intensive), full virtual machine backups (high IO and bandwidth), and file backups (bandwidth-intensive). Their resource consumption patterns differ significantly and cannot be linearly superimposed. This embodiment uses a weighted comprehensive resource utilization rate (dynamically configured with weights based on resource dimensions) combined with a threshold penalty mechanism. Non-linear expansion is triggered only when resource contention exceeds a critical point, allowing the incremental recalculation factor to more realistically depict the time-level resource contention situation of heterogeneous tasks coexisting.
[0080] Meanwhile, to transform the resource contention level into a quantifiable correction coefficient that can guide prediction, and to determine the incremental recalculation factor for each time period, further, in step S306, the incremental recalculation factor corresponding to each recalculation time period is determined by weighted penalty correction based on the resource contention level, preset weight coefficient, and resource competition threshold, and satisfies the following relationship: δ res (s j )=1+α·max(0,R w (s j )-θ), Where, δ res (s j R is the incremental recalculation factor. w (s j ) is the weighted comprehensive resource utilization rate determined based on the resource competition level and preset weight coefficients during the recalculation period, where α is the penalty coefficient and θ is the resource competition threshold.
[0081] Specifically, the weighted comprehensive resource utilization rate is determined by combining the utilization rate of each resource dimension and its weight coefficient. The weight coefficient is dynamically configured according to the backup task type. For example, the weight coefficient of storage I / O is higher in database backup scenarios, and the weight coefficient of network bandwidth is higher in file backup scenarios. The penalty coefficient α controls the growth slope of the factor, and the resource contention threshold θ is the critical point for triggering the penalty. When the weighted comprehensive utilization rate does not exceed this threshold, the difference between the weighted comprehensive utilization rate and the resource contention threshold is zero, and the incremental recalculation factor remains at 1. When it exceeds this threshold, the larger the difference, the larger the incremental recalculation factor, and the longer the estimated time is, thus truly reflecting the nonlinear impact of resource contention on task execution time.
[0082] To facilitate understanding, the following explanation uses specific numerical examples. δ is calculated for the incremental recalculation factor vector of each recalculation period. res Assuming the weighting coefficient is w I =0.5、w B =0.3、w C =0.2, threshold θ=0.6, penalty coefficient α=1.5, and the upper limit of resource capacity for each time period is I. cap =800、B cap =1000、C cap =400.
[0083] (1) For the period of 19:00–20:00: Utilization rates: I = 500 / 800 = 0.625, B = 250 / 1000 = 0.25, C = 230 / 400 = 0.575. Weighted average: (0.5×0.625+0.3×0.25+0.2×0.575)=0.5025 Coupling correction: Couple = 0 δ res =1+1.5×max(0,0.5025+0-0.6)=1.0 (No penalty if the threshold is not reached); (2) For the period of 20:00–21:00: Utilization rates: I = 600 / 800 = 0.75, B = 350 / 1000 = 0.35, C = 340 / 400 = 0.85. Weighted average: (0.5×0.75+0.3×0.35+0.2×0.85)=0.65 Coupling correction: Couple = 0 δ res =1+1.5×max(0,0.65-0.6)=1.075 (Slight overload, estimated time expansion 7.5%). (3) For the 21:00–22:00 time period (including coupling correction): Original utilization rates: I = 240 / 800 = 0.3, B = 830 / 1000 = 0.83, C = 330 / 400 = 0.825. Weighted average: (0.5×0.3+0.3×0.83+0.2×0.825)=0.564 Including coupling correction: 0.564 + (8 / 800 × 0.5) = 0.569 δ res =1+1.5×max(0,0.569-0.6)=1.0 (Threshold not reached).
[0084] Based on the above calculations, the system outputs a time-period-level incremental recalculation factor vector [δ]. res (19:00)=1.0,δ res (20:00)=1.075,δ res (21:00)=1.0]. The estimated completion time of each task is recalculated based on the corresponding component of its time period. For example, the estimated completion time of a task located in the 20:00–21:00 time period will be multiplied by a correction factor of 1.075 on top of the base time. By introducing an incremental recalculation factor, the system realizes dynamic quantitative correction of backup duration due to resource contention, providing an accurate basis for duration prediction for task scheduling after batch multi-time period interactive adjustments.
[0085] like Figure 6 As shown, in step S307, each task is mapped to the corresponding component of the incremental recalculation factor vector to obtain the incremental recalculation factor for each task. For example, assume the recalculation factor vector is organized as δ according to the time series. res (s1), δ res (s2) and δ res (s3) consists of three components corresponding to three consecutive time periods: 19:00–20:00, 20:00–21:00, and 21:00–22:00. During the first and last periods (s1 and s3), resource status is stable, and the factor value remains at the baseline of 1.0. During the middle period (s2), resource disturbances caused by batch dragging cause the factor value to rise to 1.305. This vector, as a time-period level output, can be understood as a deviation coefficient table indexed by time period, used to quantify the amplification / reduction ratio of each time period relative to the baseline load.
[0086] Similarly, Figure 6As shown, at the task-level mapping output level, each task extracts the corresponding component from the vector and substitutes it into the completion time prediction formula based on its scheduled target time period. It should be noted that the following task allocation is only illustrative, intended to visually demonstrate the mapping mechanism of "task → time period → factor component," and does not restrict the specific rules of task allocation in actual scheduling. The specific mapping relationship is as follows: Task A (database incremental) and Task B (virtual machine full) are scheduled by the system to 20:00–21:00, sharing the same vector component δ. res (s2)=1.305; Task C (full file) is scheduled for the 21:00–22:00 time slot, corresponding to the vector component δ. res (s3)=1.0.
[0087] It is important to note that the prediction model and interactive operation of existing backup scheduling systems are independent: after manual adjustment by the operator, only the task position is updated, and the predicted time is still calculated based on a static formula, which cannot reflect the actual resource contention after the adjustment. This embodiment maps each task to the corresponding time period component of the incremental recalculation factor vector, so that batch multi-time period interactive adjustments directly drive the incremental correction of the prediction model. After the operator releases the drag action, the batch update and real-time rendering of the estimated completion time of each task are completed. At this time, the base time reflects the task's own attributes, the queuing penalty factor reflects the queue waiting delay, and the incremental recalculation factor reflects the dynamic changes of time period-level resource contention—the three constitute a three-layer progressive correction of "ideal time consumption → waiting delay → resource expansion," translating the resource layer language into a time layer language that the operator can directly understand.
[0088] Meanwhile, in order to update the estimated completion time of each task, further, in step S308, the estimated completion time is determined by correcting the product of the base time of each task, the queuing penalty factor, and the incremental recalculation factor, and satisfies the following relationship: T pred (task i )=T base (task i )×γ(Q task_i )×δ res (S task_i ), where T pred (task i ) represents the estimated completion time; T base (task i γ(Q) represents the base time determined based on the amount of task data and the theoretical peak speed; task_i δ is the queuing penalty factor reflecting the delay of the task waiting in the queue; res (S task_i ) represents the incremental recalculation factor corresponding to the task.
[0089] To facilitate understanding, the following explanation uses specific numerical examples. Assume that task A has a base time of 120 minutes, a queuing penalty factor of 1.2 (there are 2 tasks ahead in the queue), and an incremental recalculation factor of 1.075 for the main time slot of 20:00–21:00. Then its estimated completion time is: T pred (A) = 120 × 1.2 × 1.075 = 154.8 minutes; Task B has a base time of 60 minutes, a queuing penalty factor of 1.0 (no queuing), and an incremental recalculation factor of 1.075 for the main time slot of 20:00–21:00. Therefore, its estimated completion time is: T pred (B) = 60 × 1.0 × 1.075 = 64.5 minutes; Given that Task C has a base time of 90 minutes, a queuing penalty factor of 1.1, and an incremental recalculation factor of 1.0 for its time slot of 21:00–22:00, its estimated completion time is: T pred (C) = 90 × 1.1 × 1.0 = 99 minutes.
[0090] By incorporating an incremental recalculation factor, the system can identify slight resource contention during the 20:00–21:00 period, and accordingly increase the predicted time by 7.5%, making the estimated result closer to the actual execution time. The system completes the above calculations after the user releases the interactive operation and renders the updated estimated completion time in real time on the visualization interface, completing the closed loop from technical calculation to interactive feedback.
[0091] To further understand the overall application process of the batch multi-time interactive time-series scheduling method in this embodiment, a comprehensive application scenario is given below: An operations and maintenance personnel needs to optimize and adjust the nighttime batch processing window, involving the following three types of heterogeneous tasks: Task X: Full database backup, data volume 3.2TB, module type is database proxy; Task Y: Full backup of virtual machines, data volume 1.8TB, module type is virtual machine agent; Task Z: Incremental backup of the file system, data volume 200GB, module type is file agent.
[0092] The operator drags a batch of tasks X, Y, and Z from the original time period to the target time period. The system dynamically determines the recalculation range based on the time period span and data volume, determines preset mapping coefficients based on the module type and task type of each task, and then determines the predicted values of storage I / O, network bandwidth, and computing resources occupied by each task.
[0093] During the incremental recalculation process, the system detects that a full database backup and log backup task already exists in the target time period. After the virtual machine full backup of task Y is migrated, the storage IO contention ratio will exceed the resource contention threshold, and the incremental recalculation factor for the corresponding time period will exceed the preset threshold. The system generates a resource contention alarm message on the visualization interface, prompting "It is recommended to postpone until after 0:30".
[0094] After the operator adjusts the task group to the suggested time period, the system recalculates the incremental recalculation factor vector for each time period, maps each task to the corresponding time period component, and updates the estimated completion time. Specifically, task X's estimated completion time is significantly extended due to the high IO preset mapping coefficient of a full database backup; task Z's impact on the overall time period is relatively small due to the low resource consumption of incremental file backups; task Y's expected completion time is restored after the time period adjustment. The system renders a 3D resource map of storage IO, network bandwidth, and computing resources, along with the updated time period-level estimated completion time, in real time on the interface to assist the operator in scheduling optimization.
[0095] Below, we provide some comparative experiments to further illustrate the embodiments of the present invention, as follows: (1) Experimental setup and subjects To verify the performance advantages of this invention's embodiments in batch, multi-time-segment task drag-and-drop scheduling, three typical test scenarios were set up. The hardware platform uniformly adopted an enterprise-grade backup server (Intel Xeon Gold 6248 / 64GB DDR4 / 10 Gigabit Ethernet), and the software environment was based on the Kubernetes scheduling engine to simulate time-segment resource quotas. The dataset covered 12-60 heterogeneous backup tasks, including full / incremental database backups, full / incremental virtual machine backups, and full / incremental file backups, simulating the real load of an enterprise-grade data protection platform.
[0096] (2) Definition of performance indicators Prediction Error (MAPE): Mean Absolute Percentage Error, reflecting the degree of deviation between the predicted completion time and the actual execution time; Resource contention misjudgment rate: The percentage of tasks whose assessed contention level deviates from the actual level by ≥1 level, characterizing the accuracy of resource conflict early warning; Response latency: The end-to-end latency from drag-and-drop release to the completion of all task forecast time updates; Overall resource utilization: average utilization of storage I / O, network bandwidth, and computing resources over a given period; Extreme scenario crash rate: The percentage of test rounds in which the scheduling engine fails due to timeout or resource overflow.
[0097] Each metric was automatically collected through scheduling engine logs and performance profiling tools. Each metric was measured 100 times, and the median and 95% confidence interval were taken after removing outliers.
[0098] (3) Comparison method Comparative example: The static prediction model uses a fixed volume coefficient multiplied by the base time of each task to obtain the estimated completion time. It does not introduce heterogeneous resource mapping, incremental recalculation factors, or completion time prediction based on incremental recalculation factor vectors. This model treats each task as an independent entity, without considering the actual mapping relationship of tasks on heterogeneous resources, without being aware of the dynamic changes in resource load at the time period level, and without establishing overflow transmission correlation between adjacent time periods. It belongs to a baseline scheduling strategy without time period coupling. Example 1: Introducing Heterogeneous Resource Mapping and Incremental Recalculation Factor Vector. Heterogeneous resource mapping maps heterogeneous tasks such as databases, virtual machines, and files to corresponding storage I / O, compute, and network resource dimensions. The incremental recalculation factor vector calculates the deviation coefficient of each time period relative to the baseline load, indexed by time period. The estimated completion time is corrected by multiplying the base time of each task, the queuing penalty factor, and the incremental recalculation factor. However, this example still uses a fixed time window for recalculation, failing to dynamically expand the buffer range based on the span of batch dragging and the data volume, and also failing to introduce overflow coupling correction between adjacent time periods, thus failing to capture the cascading transmission of resource disturbances across time periods. Experiment Example 2: Building upon Experiment Example 1, this experiment further incorporates dynamic buffer expansion and overflow coupling correction. Dynamic buffer expansion adaptively expands the buffer range forward and backward based on the time span of the affected time period set and the data volume of each task, determining the recalculation time period. Overflow coupling correction is achieved through the formula Couple(s)... j )=β·[Overflow(s j-1 ))+Overflow(s j+1 The process quantifies the resource consumption transferred from adjacent time periods due to overload to the current time period, and corrects the resource contention situation of adjacent time periods based on this coupling correction amount. Subsequently, an incremental recalculation factor vector is calculated based on the corrected resource contention status of each time period, and the estimated completion time of each task is then determined. These three elements work together to form a closed-loop scheduling mechanism: "resource mapping—dynamically expanding recalculation time periods—coupling correction—generating recalculation factor vectors—predicting completion time." Specifically, the estimated completion time of each task is determined by multiplying its base time, queuing penalty factor, and incremental recalculation factor. A more specific formula for calculating the estimated completion time can be: T pred (task i )=T base (task i )×γ(Q task_i )×δ res (S task_i ), where T pred (task i ) represents the estimated completion time; T base (task iγ(Q) represents the base time determined based on the amount of task data and the theoretical peak speed; task_i δ is the queuing penalty factor reflecting the delay of the task waiting in the queue; res (S task_i ) represents the incremental recalculation factor corresponding to the task. (4) Experimental procedure Before each scenario test, the scheduling engine state was reset and historical load records were cleared. Automated test scripts were used to simulate batch task group drag-and-drop operations, with three rounds of batch adjustments performed in each scenario (corresponding to 3, 5, and 10 drags respectively), corresponding one-to-one with the test scenarios in Tables 1-3. In each round of testing, the scheduling engine executed the following steps: First, it parsed the drag-and-drop instructions, identifying the source time period, target time period, and the time period between them for the task group, forming a set of affected time periods; then, it mapped each heterogeneous task to its corresponding resource dimension (database tasks mapped to storage IO, virtual machine tasks mapped to CPU, and file tasks mapped to network bandwidth); subsequently, based on the time period span and the data volume of each task, it adaptively expanded the buffer range forward and backward to determine the recalculation time period; it combined the overflow coupling correction amount (Couple) to correct resource contention; finally, it calculated the incremental recalculation factor vector δ_res for each time period; and finally, it substituted into the expected completion time formula for product correction and output to the feedback layer for real-time rendering and resource contention alarms. The test process covered four stages: initial load baseline determination, batch update of predicted time after batch drag-and-drop, comparison of actual execution time, and 30-minute steady-state resource monitoring. The entire process is recorded, including scheduling decision logs and resource contention snapshots.
[0099] (5) Test data To facilitate reproduction, the following core parameters are used uniformly across all scenarios. Each experimental case only executes the steps included in its corresponding technical solution; steps not included are not executed: Incremental recalculation factor formula δ res (s j )=1+α·max(0,R w (s j In the equation (α - θ), the competition sensitivity coefficient α is set to 0.15, and the competition threshold θ is set to 0.85. Overflow Coupling Correction Formula Couple(s) j )=β·[Overflow(s j-1 ))+Overflow(s j+1 In the figure, the coupling coefficient β is taken as 0.5; Resource quotas for each dimension are preset according to task type (storage IO base quota 80, CPU base quota 4 cores, network bandwidth base quota 10 Gigabit).
[0100] Scenario 1: Medium-load, light contention scenario (24 heterogeneous tasks, 3 batch drag-and-drop operations) Hardware: Intel Xeon Gold 6248 + 64GB DDR4 + 10 Gigabit Ethernet; Scheduling engine: Kubernetes + custom resource quota controller; Dataset: 24 tasks, 10 databases, 8 virtual machines, 6 files, spanning 6 hours.
[0101] Table 2. Comparison of experimental data in medium-load, low-contention scenarios Prediction Error (MAPE) 5.5% 8.2% 15.2% ↓64% ↓46% Resource contention misjudgment rate 1.5% 4.0% 8.3% ↓82% ↓52% Response delay 72ms 58ms 35ms +37ms +23ms Comprehensive resource utilization rate 65.8% 62.5% 58.2% ↑13.1% ↑7.4% Crash rate in extreme scenarios 0% 0% 0% — — Scenario 2: Heavy-load moderate contention scenario (36 heterogeneous tasks, 5 batch drags, time-limited overflow occurs) Hardware: Intel Xeon Gold 6248 + 64GB DDR4 + 10 Gigabit Ethernet; Scheduling engine: Kubernetes + custom resource quota controller; Dataset: 36 tasks, 15 databases, 12 virtual machines, and 9 files, spanning 8 hours.
[0102] Table 3. Comparison of experimental data in heavy-load, moderate contention scenarios Prediction Error (MAPE) 10.5% 15.0% 27.5% ↓62% ↓45% Resource contention misjudgment rate 6.5% 14.0% 32.5% ↓80% ↓57% Response delay 105ms 82ms 50ms +55ms +32ms Comprehensive resource utilization rate 72.0% 68.0% 63.0% ↑14.3% ↑7.9% Crash rate in extreme scenarios 0% 8% 15% The crash rate dropped to 0 ↓7% Scenario 3: Extreme high-pressure scenario (60 heterogeneous tasks, 10 batch drags, severe cross-time period cascading overflow) Hardware: Intel Xeon Gold 6248 + 64GB DDR4 + 10 Gigabit Ethernet; Scheduling engine: Kubernetes + custom resource quota controller; Dataset: 60 tasks, 25 databases, 20 virtual machines, 15 files, spanning 10 hours.
[0103] Table 4. Comparison of experimental data under extreme high-pressure scenarios Prediction Error (MAPE) 18.0% 26.0% 45.5% ↓60% ↓43% Resource contention misjudgment rate 12.0% 28.0% 65.0% ↓82% ↓57% Response delay 185ms 145ms 82ms +103ms +63ms Comprehensive resource utilization rate 81.2% 76.0% 71.0% ↑14.4% ↑7.0% Crash rate in extreme scenarios 0% 18% 65% The crash rate dropped to 0 ↓47% (6) Experiment Summary The experimental results above show that, in the scenario of batch multi-time interactive drag-and-drop scheduling, Experiment 1 has achieved a significant performance improvement compared to the comparative example by introducing heterogeneous resource mapping and incremental recalculation factor vector; Experiment 2 further superimposes dynamic buffer expansion and overflow coupling correction on this basis, forming a more complete closed-loop mechanism and achieving a systemic performance leap. a. Prediction Error: Experiment 1, by incrementally recalculating the factor vector, reduced MAPE from 15.2%-45.5% in the comparative example to 8.2%-26.0%, a reduction of 43%-46%, verifying the fundamental value of time-period bias coefficient correction; Experiment 2, through dynamic buffer expansion and overflow coupling correction, further reduced MAPE to 5.5%-18.0%, a reduction of 60%-64% compared to the comparative example, and a further reduction of about 3-8% compared to Experiment 1, demonstrating the cumulative gain of cross-time-period overflow quantization on prediction accuracy; b. False alarm rate of resource contention: After introducing resource mapping and factor vectors in Experiment Example 1, the false alarm rate decreased from 8.3%-65.0% in the comparative example to 4.0%-28.0%, a reduction of 52%-57%, indicating that resource dimension subdivision and time period-level correction can significantly improve the accuracy of early warning; Experiment Example 2 further reduced the false alarm rate to 1.5%-12.0% through overflow coupling correction, and in Scenario 3, it decreased from 28% in Experiment Example 1 to 12%, proving that dynamic buffer expansion is indispensable for capturing cross-time period chain overflows; c. Resource utilization rate: Experimental Example 1 improved by 7.0%-7.9% compared to the control group, reaching 62.5%-76.0%, indicating that the incremental recalculation factor vector can more fully tap into the idle quota of the time period; Experimental Example 2 further improved to 65.8%-81.2% on this basis, an improvement of 13.1%-14.4% compared to the control group, verifying the optimization depth of the complete mechanism for resource scheduling; d. Stability in extreme scenarios: In scenario 3, the crash rate of the comparative example was as high as 65% due to the lack of resource contention awareness; after introducing incremental recalculation factor vector in Experiment 1, the crash rate dropped to 18%, but scheduling timeouts and resource deadlocks still occurred because the fixed window could not completely capture the overflow propagation path; Experiment 2 quantified cross-time period overflow into the resource baseline increment of adjacent time periods through dynamic buffer expansion and coupling correction formula, and no crash occurred in all test rounds, verifying the robustness of the four links of "resource mapping - factor correction - dynamic buffer - overflow coupling" in coordination; e. Response latency: Experimental Example 1 showed an increase of 23ms-63ms compared to the control group, while Experimental Example 2 saw a further increase of 14ms-40ms due to dynamic buffer expansion and overflow coupling correction. All scenarios were kept within 185ms, far below the acceptable threshold of 200ms for real-time interaction. This latency increase represents a reasonable trade-off between limited response time and significantly improved prediction error, resource contention misjudgment rate, and crash rate.
[0104] In summary, Experiment 1 verified the significant improvement of static scheduling by heterogeneous resource mapping and incremental recalculation factor vector as the basic scheme; Experiment 2 verified the complete coverage capability of cross-time period chain overflow scenarios after adding dynamic buffer expansion and overflow coupling correction on this basis. Together, they constitute a technological progression path.
[0105] Unlike the previous embodiments, this embodiment replaces the fixed window with a dynamic buffer, adaptively defining the recalculation range based on the time span and data volume to avoid cascading effects due to an excessively small window or invalid calculations due to an excessively large window. Resource contention is dynamically coupled into the prediction model through time-level incremental recalculation factors, ensuring that prediction results are updated in real-time with interactive operations. Resource contention information is generated based on the resource contention level of each recalculation time period, and task backlog during periods of high resource contention is proactively avoided based on the incremental recalculation factor vector. When operators perform batch drag-and-drop adjustments across time periods, they can instantly observe the changing trends of the expected completion time for each time period, predict and avoid the risk of global scheduling collapse caused by local adjustments, forming a virtuous cycle of "prediction → scheduling → actual convergence".
[0106] Example 4 like Figure 7 As shown, in one embodiment, a batch multi-time interactive scheduling device for data protection tasks is provided. This device includes: The task presentation module 4001 is used to present multiple data protection task blocks and their corresponding time period distribution on the visualization interface; The batch multi-period interaction module 4002 is used to respond to batch task multi-period configuration adjustments and determine the set of affected period periods. The heterogeneous resource mapping module 4003 is used to establish a nonlinear mapping relationship between heterogeneous task attributes, data volume and multidimensional resource occupation. It uses preset coefficients to characterize the resource consumption characteristics of heterogeneous tasks and determine the predicted occupation value of each task in each resource dimension. The incremental recalculation module 4004 is used to determine each recalculation period based on the set of affected periods, then determine the resource contention level of each recalculation period based on the predicted value of each task's resource usage in each resource dimension, and determine the incremental recalculation factor corresponding to each recalculation period based on the resource contention level, forming an incremental recalculation factor vector. The prediction feedback module 4005 is used to update the estimated completion time of each task based on the incremental recalculation of the factor vector and to render the feedback in real time on the visualization interface.
[0107] Furthermore, the heterogeneous resource mapping module 4003 also includes: a nonlinear mapping relationship between each task and storage I / O, network bandwidth and computing resources based on the module type, task type and backup data volume of each task, so as to determine the predicted resource usage value of each task in each resource dimension; the nonlinear mapping relationship is realized by a preset mapping coefficient, which is associated with the module type and task type, and different predicted resource usage values correspond to different module types or different task types.
[0108] Furthermore, such as Figure 8 As shown, the incremental recalculation module 4004 further includes: The buffer range determination unit 40041 is used to dynamically expand the buffer range based on the time span of the affected time period set and the data volume of each task, determine the incremental recalculation range of each affected time period, and obtain each recalculation time period. The resource contention level calculation unit 40042 is used to accumulate the predicted values of each task's resource occupation in each resource dimension for each recalculation period, and combine them with the baseline resources to generate the resource contention level for each recalculation period. The incremental recalculation factor calculation unit 40043 is used to perform weighted penalty correction on each recalculation period based on the resource contention level, preset weight coefficient and resource contention threshold, to obtain the incremental recalculation factor corresponding to each recalculation period and form an incremental recalculation factor vector.
[0109] Furthermore, in the incremental recalculation factor calculation unit 40043, the incremental recalculation factor corresponding to each recalculation period is determined by weighted penalty correction based on the resource contention level, preset weight coefficient, and resource contention threshold, and satisfies the following relationship: δ_res(s_j)=1+α·max(0,R_w(s_j)-θ), Where δ_res(s_j) is the incremental recalculation factor, R_w(s_j) is the weighted comprehensive resource utilization rate determined based on the resource contention level and preset weight coefficients during the recalculation period, α is the penalty coefficient, and θ is the resource competition threshold.
[0110] Furthermore, such as Figure 9 As shown, the prediction feedback module 4005 further includes: The task factor determination unit 40051 is used to map each task to the corresponding component of the incremental recalculation factor vector to obtain the incremental recalculation factor corresponding to each task. The completion time prediction unit 40052 is used to perform product correction based on the base time of each task, the queuing penalty factor and the incremental recalculation factor to determine the expected completion time of each task. The rendering feedback unit 40053 is used to render the estimated completion time in real time on the visualization interface to provide feedback on the impact of batch, multi-period interactive adjustments on the overall schedule. Furthermore, in the completion time prediction unit 40052, the estimated completion time is determined by the product of the base time of each task, the queuing penalty factor, and the incremental recalculation factor, and satisfies the following relationship: T_pred(task_i) = T_base(task_i) × γ(Q_task_i) × δ_res(s_task_i), where T_pred(task_i) is the estimated completion time; T_base(task_i) is the base time determined based on the task data volume and theoretical peak speed; γ(Q_task_i) is the queuing penalty factor reflecting the delay of the task in the queue; and δ_res(s_task_i) is the incremental recalculation factor corresponding to the task.
[0111] It is worth noting that the various units and modules included in the above-mentioned batch multi-time interactive scheduling device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.
[0112] Example 5 In this embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When executed by a processor, the program implements a batch multi-time interactive scheduling method for data protection tasks as described in any one of embodiments 1 to 3.
[0113] The computer storage medium of this invention can be any combination of one or more computer-readable media. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0114] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.
[0115] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0116] Computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, Ruby, and Go, as well as conventional procedural programming languages such as "C" or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0117] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.
Claims
1. A batch multi-time interactive scheduling method for data protection tasks, characterized in that, Including the following steps: The visualization interface presents multiple data protection task blocks and their corresponding time periods. In response to batch task configuration adjustments across multiple time periods, determine the set of affected time periods; Establish a nonlinear mapping relationship between heterogeneous task attributes, data volume and multidimensional resource consumption, characterize the resource consumption characteristics of heterogeneous tasks by preset coefficients, and determine the predicted value of resource consumption for each task in each resource dimension. Based on the set of affected time periods, each recalculation period is determined. Then, based on the predicted resource occupancy values of each task across each resource dimension, the resource contention level of each recalculation period is determined. Furthermore, based on the resource contention level, the incremental recalculation factor corresponding to each recalculation period is determined, forming an incremental recalculation factor vector. The incremental recalculation factor is determined by weighted penalty correction of each recalculation period based on the resource contention level, preset weight coefficients, and resource contention threshold, and satisfies the following relationship: δ res (s j )=1+α·max(0,R w (s j )-θ), δ res (s j R is the incremental recalculation factor. w (s j ) is the weighted comprehensive resource utilization rate determined based on the resource contention level and preset weight coefficients during the recalculation period, where α is the penalty coefficient and θ is the resource competition threshold; Based on the incremental recalculation factor vector, the estimated completion time of each task is updated, and the feedback is rendered in real time on the visualization interface. The estimated completion time is determined by the product of the base time of each task, the queuing penalty factor, and the incremental recalculation factor, and satisfies the following relationship: T pred (task i )=T base (task i )×γ(Q task_i )×δ res (S task_i ), T pred (task i ) represents the estimated completion time; T represents the estimated completion time. base (task i γ(Q) represents the base time determined based on the amount of task data and the theoretical peak speed; task_i δ is the queuing penalty factor reflecting the delay of the task waiting in the queue; res (S task_i ) represents the incremental recalculation factor corresponding to the task.
2. The batch multi-time interactive scheduling method for data protection tasks according to claim 1, characterized in that, The set of affected time periods includes the source time period, target time period, and transit time period for batch multi-time period interactive adjustment.
3. The batch multi-time interactive scheduling method for data protection tasks according to claim 1, characterized in that, The step of establishing a nonlinear mapping relationship between heterogeneous task attributes, data volume, and multidimensional resource consumption, and determining the predicted resource consumption value of each task in each resource dimension by characterizing the resource consumption characteristics of heterogeneous tasks through preset coefficients, further includes: Based on the module type, task type, and backup data volume of each task, a nonlinear mapping relationship is established between each task and storage I / O, network bandwidth, and computing resources to determine the predicted resource usage value of each task in each resource dimension. The nonlinear mapping relationship is achieved through preset mapping coefficients, which are associated with the module type and task type. Different module types or different task types correspond to different predicted resource usage values.
4. The batch multi-time interactive scheduling method for data protection tasks according to any one of claims 1-3, characterized in that, The steps of determining each recalculation period based on the affected period set, determining the resource contention level of each recalculation period based on the predicted resource occupancy values of each task in each resource dimension, and determining the incremental recalculation factor corresponding to each recalculation period based on the resource contention level to form an incremental recalculation factor vector, further include: Based on the time span of the affected time period set and the data volume of each task, the buffer range is dynamically expanded to determine the incremental recalculation range of each affected time period, thus obtaining each recalculation time period; For each recalculation period, the predicted resource consumption values of each task in each resource dimension are summed up and combined with the baseline resources to generate the resource contention level for each recalculation period. Based on the resource contention level, preset weight coefficients, and resource competition thresholds, weighted penalty corrections are applied to each recalculation period to obtain the incremental recalculation factor corresponding to each recalculation period, forming an incremental recalculation factor vector.
5. The batch multi-time interactive scheduling method for data protection tasks according to claim 4, characterized in that, The step of summing up the predicted resource consumption values of each task across each resource dimension for each recalculation period and combining them with baseline resources to generate the resource contention level for each recalculation period also includes: For each recalculation period, determine whether its resource contention ratio exceeds a preset threshold. If it does not exceed the threshold, determine that the resource overflow amount passed from the recalculation period to the adjacent period is zero. If it exceeds the threshold, determine that the resource overflow amount passed from the recalculation period to the adjacent period due to overload. The coupling correction amount for adjacent time periods is determined based on the resource overflow amount and the preset coupling coefficient, and the resource contention situation for adjacent time periods is corrected according to the coupling correction amount.
6. The batch multi-time interactive scheduling method for data protection tasks according to any one of claims 1-3, characterized in that, The step of updating the estimated completion time of each task based on the incremental recalculation of the factor vector and rendering the feedback in real time on the visualization interface also includes: Each task is mapped to the corresponding component of the incremental recalculation factor vector to obtain the incremental recalculation factor for each task. The estimated completion time of each task is determined by multiplying and correcting the base time, queuing penalty factor, and incremental recalculation factor. The estimated completion time is rendered in real time on the visualization interface to provide feedback on the impact of batch, multi-period interactive adjustments on the overall scheduling.
7. A batch multi-time interactive scheduling device for data protection tasks, characterized in that, The device includes: The task presentation module is used to present multiple data protection task blocks and their corresponding time period distribution on a visual interface; The batch multi-time period interaction module is used to respond to batch task multi-time period configuration adjustments and determine the set of affected time periods; The heterogeneous resource mapping module is used to establish a non-linear mapping relationship between heterogeneous task attributes, data volume and multi-dimensional resource consumption. It uses preset coefficients to characterize the resource consumption characteristics of heterogeneous tasks and determine the predicted value of each task's resource consumption in each dimension. The incremental recalculation module is used to determine each recalculation period based on the set of affected periods, then determine the resource contention level of each recalculation period based on the predicted resource occupancy values of each task in each resource dimension, and determine the incremental recalculation factor corresponding to each recalculation period based on the resource contention level, forming an incremental recalculation factor vector. The incremental recalculation factor is determined by weighted penalty correction of each recalculation period based on the resource contention level, preset weight coefficients, and resource contention threshold, and satisfies the following relationship: δ res (s j )=1+α·max(0,R w (s j )-θ), δ res (s j R is the incremental recalculation factor. w (s j ) is the weighted comprehensive resource utilization rate determined based on the resource contention level and preset weight coefficients during the recalculation period, where α is the penalty coefficient and θ is the resource competition threshold; The prediction feedback module is used to update the estimated completion time of each task based on the incremental recalculation factor vector and render the feedback in real time on the visualization interface. The estimated completion time is determined by the product of the base time of each task, the queuing penalty factor, and the incremental recalculation factor, and satisfies the following relationship: T pred (task i )=T base (task i )×γ(Q task_i )×δ res (S task_i ), T pred (task i ) represents the estimated completion time; T represents the estimated completion time. base (task i γ(Q) represents the base time determined based on the amount of task data and the theoretical peak speed; task_i δ is the queuing penalty factor reflecting the delay of the task waiting in the queue; res (S task_i ) represents the incremental recalculation factor corresponding to the task.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the batch multi-time interactive scheduling method for data protection tasks as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Resource scheduling method and device, equipment, medium and program product
CN121957898A
Backup scheduling method for distributing system load and computer program product of the same enables the users to log into the system and operate at different times
TW201721421A