Data training and optimizing system based on cloud AI model

By extracting training task logs and resource usage in the cloud artificial intelligence platform, dynamic matching and cross-node migration, the problem of incoherent resource allocation in cloud artificial intelligence data training is solved, and resource utilization and task execution efficiency are improved.

CN120371528AInactive Publication Date: 2025-07-25GUANGZHOU YUANHUAN INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510540499.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-07-25
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology has problems such as inaccurate task type identification, inadequate resource allocation, incoherent scheduling, and unstable task migration in the training of cloud artificial intelligence data, resulting in low resource utilization and low training efficiency, especially in environments where resource dynamic changes or high task concurrency is high.

Method used

The task feature recognition module extracts the scheduling trajectory and resource usage in the training task log, combines GPU and memory tags for dynamic matching, monitors node operation and interruption, identifies scheduling interrupt paragraphs, realizes cross-node migration, and optimizes task scheduling strategies to improve resource utilization and task continuity.

Benefits of technology

It realizes accurate identification of task types and sub-task characteristics, reduces resource idle rate and task queuing delay, enhances real-time monitoring and feedback mechanisms for scheduling, improves the robustness of task migration and resource utilization efficiency, and ensures the continuity of task execution and the accuracy of platform scheduling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371528A_ABST
    Figure CN120371528A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of cloud artificial intelligence, in particular to a data training and optimizing system based on a cloud AI model, and the system comprises a task feature recognition module, a resource scheduling execution module, a resource load monitoring module, a cross-cloud task migration module and a performance evaluation adaptation module. According to the method, by extracting a scheduling track, resource use and a model rhythm in a training task log, task type accurate identification and subtask characteristic subdivision are realized, task division adaptability is improved, dynamic resource matching is realized in combination with a GPU and a memory label, an idling rate and queuing delay are reduced, node operation and interruption conditions are monitored, and scheduling interruption paragraphs are identified; real-time scheduling feedback is enhanced, cross-node migration is achieved according to task progress and node states, task continuity is guaranteed, training log evaluation stability is compared, task and resource adaptation efficiency is improved, and resource utilization rate improvement, execution coherence enhancement and platform scheduling precision optimization are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cloud artificial intelligence, and particularly to a data training and optimization system based on a cloud AI model. Background Art

[0002] The technical field of cloud artificial intelligence includes multiple technical branches such as artificial intelligence algorithms, data storage and processing, cloud computing platforms and their services. The core content of this technical field is to provide computing power and storage resource support through cloud computing, and combine artificial intelligence algorithms to analyze, model and learn big data in order to achieve intelligent decision-making and task processing. The application scope of cloud artificial intelligence is very wide, covering multiple aspects such as automated data processing, intelligent recommendation systems, natural language processing, image recognition, machine learning and deep learning. With the improvement of computing power and the increase in data volume, cloud artificial intelligence has been widely applied in many fields such as transportation and security industries.

[0003] Among them, the data training and optimization system refers to a technical solution for data training and optimization of a cloud artificial intelligence platform. The technical matters involved in this patent theme include multiple aspects such as data acquisition, data preprocessing, model training, and model optimization. The data training and optimization system preprocesses the training data, removes redundant information, cleans the data, and adopts certain algorithm strategies for feature selection and model training. After the model training is completed, the model performance is improved by adjusting the optimization algorithm to better meet the specific application requirements. It also uses a dynamic monitoring and feedback mechanism to make real-time adjustments to the training process to ensure the training effect and optimization efficiency of the model. The technology involved in this patent solution is mainly completed through reasonable algorithm design and data processing flow, and does not involve specific algorithm modules or device configurations.

[0004] When dealing with cloud training tasks, the existing technology has a rough judgment problem in identifying the task type, relying on fixed resource indicators or task labels, lacking in-depth analysis of task operation characteristics, resulting in difficulty in accurately adapting resource allocation to the actual needs of the task. The resource scheduling method is mainly based on static configuration or simple priority sorting, which is difficult to cope with the dynamic changes in the node resource usage status, resulting in resource idling or extended task waiting time. In the process of task scheduling execution, there is a lack of monitoring mechanism for interruption points in the scheduling chain, resulting in hidden waiting and scheduling faults when multiple subtasks run on the same node, which in turn affects the overall training efficiency. In the face of incoherent cloud resource node scheduling, the existing technology lacks an effective cross-node migration solution and only relies on manual adjustment or preset strategies, which can easily lead to task state loss or training interruption after migration. In terms of model training stability evaluation, traditional methods use the overall training results as the measurement standard, ignoring the interruption frequency and local performance changes during the migration process. It is difficult to accurately judge the adaptation relationship between platform resources and tasks, which directly affects the execution efficiency of training tasks, platform resource utilization and model deployment effects, which is particularly evident in environments with drastic dynamic changes in resources or high task concurrency. Summary of the invention

[0005] The purpose of the present invention is to solve the shortcomings of the prior art and propose a data training and optimization system based on a cloud AI model.

[0006] In order to achieve the above objectives, the present invention adopts the following technical solutions: A data training and optimization system based on a cloud AI model includes:

[0007] The task feature recognition module extracts the task scheduling trajectory, resource usage, processing time, model loading frequency, and data reading rhythm based on the model training task log of the cloud AI platform. It classifies the task type according to the GPU usage frequency and data access density, analyzes the subtask operation characteristics, and obtains the subtask resource classification structure.

[0008] The resource scheduling execution module calls the subtask resource classification structure, compares the available tags of the current resource nodes of the cloud platform, matches the GPU identification nodes and the memory expansion nodes, sorts the same nodes by idle status, arranges tasks according to priority, and obtains the subtask node allocation path;

[0009] The resource load monitoring module extracts the task running time and interruption records within the node scheduling cycle based on the subtask node allocation path, determines the waiting interval between consecutive subtasks under the same node, identifies the idle segment and rescheduling segment in the resource queue, and obtains the resource scheduling coherence state;

[0010] The cross-cloud task migration module identifies the candidate mapping situations of cloud platform nodes according to the resource discontinuous nodes marked in the resource scheduling coherence state, determines whether the state switch is consistent with the task continuity, and generates a training task migration and switch list.

[0011] As a further solution of the present invention, the sub-task resource classification structure includes a computing intensity label, a data transmission intensity label, a model processing load distribution, a resource call hierarchy structure, and a data reading frequency identifier. The sub-task node allocation path includes a GPU node allocation sequence, a memory expansion node list, a task mapping priority identifier, and a node resource matching comparison table. The resource scheduling coherence state includes a resource idle segment identifier, a task rearrangement section identifier, a node scheduling discontinuity label, and a waiting time distribution structure. The training task migration and switch list includes a sub-task correspondence mapping, a node migration path structure, a task continuity maintenance label, and a migrability condition identifier.

[0012] As a further solution of the present invention, the task feature recognition module includes:

[0013] The scheduling trajectory analysis sub-module extracts the GPU call sequence, data loading timing, and task running duration based on the model training task log of the cloud AI platform, identifies the task density of the GPU call sequence in each time period, and obtains the task scheduling density value.

[0014] The resource classification judgment sub-module calls the task scheduling density value, combines the GPU usage frequency and data access frequency, analyzes the impact of the GPU frequency on the training cycle, determines whether it exceeds the GPU and data access thresholds, and uses the formula:

[0015]

[0016] Obtain the task resource attribute type value;

[0017] where T represents the task resource attribute type value, f i represents the GPU usage frequency in the i-th time period, represents the average value of the GPU usage frequency training cycle, d i represents the data access frequency in the i-th time period, s i represents the data access rhythm reference value in the i-th time period, and n represents the number of task running time periods;

[0018] The sub-task structure extraction sub-module extracts the sub-task log segments in the corresponding task according to the task resource attribute type value, determines the model loading frequency, data reading interval, and node distribution, analyzes the execution sequence and scheduling stability, and obtains the sub-task resource classification structure.

[0019] As a further solution of the present invention, the resource scheduling execution module includes:

[0020] The GPU node selection sub-module calls the sub-task resource classification structure, detects the idle status of GPU nodes according to the available labels of the current resource nodes in the cloud platform, filters the idle GPU nodes and sorts them according to the degree of idleness, and generates a list of idle GPU nodes;

[0021] The memory expansion node matching sub-module matches the memory expansion nodes in the cloud platform based on the list of idle GPU nodes, filters the qualified nodes according to the memory capacity and idle status, and obtains a list of memory expansion nodes;

[0022] The sub-task path optimization sub-module calls the list of idle GPU nodes and the list of memory expansion nodes, sorts the GPU and memory expansion nodes by task priority and idle status, and obtains the sub-task node allocation path.

[0023] As a further solution of the present invention, the resource load monitoring module includes:

[0024] The running extraction sub-module extracts the start, completion, and interruption times of tasks during the node scheduling period based on the sub-task node allocation path, merges the running sections, and generates a node running time series;

[0025] The interval judgment sub-module calls the node running time series, identifies the waiting time between adjacent tasks, filters the items outside the threshold range and classifies and judges them, and generates a list of abnormal interval segments;

[0026] The scheduling marking sub-module judges whether the corresponding node scheduling is continuous according to the list of abnormal interval segments, combines the path length, the number of abnormal segments, and the task density, and uses the formula:

[0027]

[0028] Calculate the resource scheduling coherence index, extract the low coherence node sequence, mark the nodes with discontinuous resource utilization, and obtain the resource scheduling coherence state;

[0029] Among them, U represents the resource scheduling coherence index, L represents the length of the node task allocation path, E represents the number of abnormal interval segments, M represents the node task density, C k represents the sub-task switching frequency within the kth idle segment, q is the total number of idle segments, and W represents the node scheduling waiting times.

[0030] As a further solution of the present invention, the cross-cloud task migration module includes:

[0031] The task record extraction sub-module extracts the sub-task records under each discontinuous node according to the resource scheduling coherence state, analyzes the execution progress information of the sub-tasks, including the current task state, execution duration, remaining duration, and task interruption records, identifies the execution records of related tasks, and generates a sub-task record set;

[0032] The mapping identification sub-module identifies the candidate mapping situations of the cloud platform nodes according to the sub-task record set, matches the resource load, task requirements, and bandwidth limitations between the nodes, judges the task suitability, filters the nodes that meet the requirements, and generates a candidate mapping result set;

[0033] The migration screening sub-module makes a state transition judgment on each mapping candidate according to the candidate mapping result set, filters the task items according to the task migration conditions and the platform resource status, and uses the formula:

[0034]

[0035] Calculate the task migration suitability index, identify the task items that meet the migration conditions, and generate a training task migration switch list;

[0036] Among them, P represents the task migration suitability index, R represents the resource requirement degree of the current task, N represents the number of task migrations, Q represents the bandwidth capacity of the candidate node, O represents the utilization rate of the node resources, and Y represents the resource load of the candidate node.

[0037] As a further solution of the present invention, the system further includes a performance evaluation and adaptation module:

[0038] The performance evaluation and adaptation module calls the task identifiers in the training task migration switch list, compares the training logs of the cloud AI model on the source node and the migration node, analyzes the training step duration and interruption frequency, filters the tasks with consistent stability, and marks them as platform optimization matching items to obtain an AI model scheduling matching mapping table;

[0039] The AI model scheduling matching mapping table includes training stability scoring results, task execution consistency identifiers, migration node performance matching degrees, and model scheduling optimization items.

[0040] As a further solution of the present invention, the performance evaluation and adaptation module includes:

[0041] The task migration comparison sub-module calls the task identifiers in the training task migration switch list, identifies the training logs of the cloud AI model on the source node and the migration node, and conducts a training process evaluation by comparing and analyzing the training step duration and interruption frequency, using the formula:

[0042]

[0043] Calculate the total training difference time and generate the task migration comparison result;

[0044] Among them, T train (j) represents the training time of the j-th step, T interrupt (j) represents the interruption time of the j-th step, m represents the total number of steps, and H represents the total training difference time;

[0045] The stability screening sub-module filters tasks with few interruption frequencies and stable training step durations during training based on the task migration comparison result, records the screening criteria, marks the eligible tasks, and generates a set of stable tasks after screening;

[0046] The scheduling matching generation sub-module filters tasks that match the platform optimization according to the set of stable tasks after screening, determines the matching priority of each task, and obtains the AI model scheduling matching mapping table.

[0047] Compared with the prior art, the advantages and positive effects of the present invention are as follows:

[0048] In the present invention, through the in-depth extraction of the scheduling trajectory, resource usage, and model operation rhythm of the training task log, the accurate identification of the training task type and the subdivision of the running characteristics of subtasks are realized, breaking the coarse-grained classification method of a single resource occupancy ratio, making the task division more targeted and adaptable. During the task arrangement process, the GPU and memory label information of the nodes are jointly matched, and dynamic sorting and mapping allocation are completed according to the idle state of the nodes, effectively reducing the resource idling rate and task queuing delay. The fine-grained extraction of the node running time and task interruption record during scheduling realizes the identification and marking of the continuous interruption paragraphs of tasks, enhancing the real-time monitoring ability and feedback mechanism of resource scheduling. For resource nodes with discontinuous scheduling, by recording the execution progress of subtasks and the node status, cross-node mapping and switching are performed, improving the robustness of task migration while maintaining the consistency of task status. Based on the training log difference between the source node and the migration node, stability judgment and performance scoring are carried out, improving the accurate adaptation efficiency of task scheduling and providing a more stable reference basis for platform-level model allocation. The overall solution realizes the collaborative optimization of resource utilization efficiency, task execution continuity enhancement, and platform scheduling accuracy through task identification refinement, resource matching intelligence, scheduling process monitoring precision, cross-node migration consistency guarantee, and stability evaluation feedback linkage. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 is the system flow chart of the present invention;

[0050] Figure 2 is the flow chart of the task feature recognition module in the present invention;

[0051] Figure 3It is the flowchart of the resource scheduling execution module in the present invention;

[0052] Figure 4 It is the flowchart of the resource load monitoring module in the present invention;

[0053] Figure 5 It is the flowchart of the cross-cloud task migration module in the present invention;

[0054] Figure 6 It is the flowchart of the performance evaluation and adaptation module in the present invention. Detailed implementation manners

[0055] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention.

[0056] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention. In addition, in the description of the present invention, the meaning of "a plurality of" is two or more, unless otherwise specifically defined.

[0057] Please refer to Figure 1 , a data training and optimization system based on a cloud AI model includes:

[0058] The task feature recognition module extracts the scheduling track, resource usage, processing duration, model loading frequency, and data reading rhythm of the task based on the model training task log of the cloud AI platform, classifies the task into a compute-intensive type and a transmission-intensive type according to the GPU usage frequency and data access density, and analyzes the running characteristics of each subtask to obtain the subtask resource classification structure;

[0059] The resource scheduling execution module calls the subtask resource classification structure, compares it with the available labels of the current resource nodes of the cloud platform, matches the GPU identification nodes and the memory expansion nodes, sorts the same type of nodes according to the idle state, and arranges the tasks in the order of priority to obtain the subtask node allocation path;

[0060] The resource load monitoring module assigns paths based on subtask nodes, extracts the task running time and task interruption records of nodes within the scheduling period, judges the waiting intervals between consecutive subtasks under the same node, identifies the idle segments and rearrangement segments in the resource queue, marks the nodes with discontinuous resource utilization, and obtains the resource scheduling coherence status;

[0061] The cross-cloud task migration module extracts the corresponding subtask records and execution progress information according to the resource discontinuous nodes marked in the resource scheduling coherence status, identifies the candidate mapping situations on the cloud platform nodes, judges whether the state switch and task continuity are consistent, filters the task items that meet the migration conditions, and generates a training task migration and switch list;

[0062] The performance evaluation and adaptation module calls the task identifiers in the training task migration and switch list, compares the training logs of the cloud AI model on the source node and the migration node, analyzes the time consumption of training steps and the interruption frequency, filters the tasks with consistent stability, and marks them as platform optimization matching items to obtain the AI model scheduling matching mapping table.

[0063] The subtask resource classification structure includes the computational intensity label, data transmission intensity label, model processing load distribution, resource call hierarchical structure, and data reading frequency identifier. The subtask node allocation path includes the GPU node allocation sequence, memory expansion node list, task mapping priority identifier, and node resource matching comparison table. The resource scheduling coherence status includes the resource idle segment identifier, task rearrangement section identifier, node scheduling discontinuity label, and waiting time distribution structure. The training task migration and switch list includes subtask correspondence mapping, node migration path structure, task continuity maintenance label, and transferable condition identifier. The AI model scheduling matching mapping table includes training stability scoring results, task execution consistency identifier, migration node performance matching degree, and model scheduling optimization items.

[0064] Please refer to Figure 2 , the task feature recognition module includes:

[0065] The scheduling trajectory analysis sub-module extracts the GPU call sequence, data loading time series, and task running duration based on the model training task logs of the cloud AI platform, identifies the task density of the GPU call sequence in each time period, and obtains the task scheduling density value;

[0066] Based on the model training task logs generated by the cloud AI platform, extract the GPU call sequence from the logs. For example, in a typical deep learning training process, the GPU call frequency and duration can be observed. By monitoring the GPU call sequence in each time period, the running intensity of the task in different time periods can be analyzed. For example, during the stage with large data loading, the GPU is called frequently, indicating that high-intensity computing is in progress. According to the records of the task running duration and data loading time sequence, further calculate the overall execution efficiency of the task. For example, by calculating the total duration from the start to the end of the task and analyzing the duration of each stage, it can be found that some stages are lengthened due to data waiting. By comprehensively analyzing the relationship between the statistical data of the GPU call sequence and the data loading time sequence, obtain the task scheduling intensity value, which can help the administrator understand the resource usage of each task and optimize the scheduling strategy accordingly.

[0067] The resource classification judgment sub-module calls the task scheduling intensity value, combines the GPU usage frequency and data access frequency, analyzes the impact of the GPU frequency on the training cycle, and determines whether it exceeds the GPU and data access thresholds, using the formula:

[0068]

[0069] Obtain the task resource attribute type value;

[0070] where, T represents the task resource attribute type value, f i represents the GPU usage frequency in the i-th time period, represents the average value of the GPU usage frequency in the training cycle, d i represents the data access frequency in the i-th time period, s i represents the data access rhythm reference value in the i-th time period, and n represents the number of task running time periods;

[0071] Obtain it through the overlap interval density between the GPU call sequence and the data loading time sequence, as a reference for the resource usage activity in the current task scheduling stage. Through the time sequence distribution of the GPU usage frequency and data access frequency in the task running cycle, collect the GPU usage frequency value and data access frequency value in each equal divided time period. Assume that the training cycle of a certain task is 120 seconds, divided into n = 6 time periods, each period is 20 seconds;

[0072] The monitored GPU usage frequency f i are 0.82, 0.79, 0.91, 0.76, 0.88, 0.81 respectively, and the unit is unified as the frequency value (times / second);

[0073] The data access frequency d iare 1.05, 0.92, 1.10, 0.89, 1.20, 0.95 (times / second), and this data access frequency is obtained by dividing the total number of data loading requests of the task in the corresponding time period by the duration of the time period;

[0074] For the data access rhythm reference value s of each time period i , it can be preset by the task scheduling platform according to the average data access frequency of transmission-intensive tasks in historical training tasks, and s1 - s6 are all taken as 1.00;

[0075] The periodic mean barf of the GPU usage frequency is the arithmetic mean of the above six segments of GPU frequencies, and the calculation is as follows:

[0076] Substitute each item into

[0077] The first item: (0.82 - 0.8283)·(1.05 - 1.00) = -0.000415;

[0078] The second item: (0.79 - 0.8283)·(0.92 - 1.00) = 0.003064;

[0079] The third item: (0.91 - 0.8283)·(1.10 - 1.00) = 0.00817;

[0080] The fourth item: (0.76 - 0.8283)·(0.89 - 1.00) = 0.007513;

[0081] The fifth item: (0.88 - 0.8283)·(1.20 - 1.00) = 0.01034;

[0082] The sixth item: (0.81 - 0.8283)·(0.95 - 1.00) = 0.000915;

[0083] After accumulation: T =

[0084] | -0.000415 + 0.003064 + 0.00817 + 0.007513 + 0.01034 + 0.000915 |

[0085] ≈0.0296;

[0086] This value is the deviation value of the resource utilization of the task within the current cycle. If the set GPU data resource judgment threshold is 0.025, then T > 0.025, indicating that the current task's resource utilization behavior deviates from the set standard range and should be judged as a transmission-intensive task. The setting basis of this resource judgment threshold lies in the statistical distribution results of the deviation values of historical training tasks on the platform, and the 95% confidence lower bound is selected. For example, if the average deviation value among 100 transmission-intensive tasks is 0.032 and the standard deviation is 0.006, then 0.025 is the empirical threshold;

[0087] Through the joint deviation judgment of GPU frequency and data frequency in the time dimension, it avoids the static classification method of the average value or the fluctuation range, and realizes the quantitative determination of the dynamic characteristics of the task resource attributes. The result shows that the current task belongs to the transmission-intensive type, and subsequent subtasks need to be classified into the high-data-scheduling priority queue, and the task resource classification structure will be further refined according to this classification result in the subsequent module.

[0088] The subtask structure extraction sub-module extracts the subtask log segments in the corresponding task according to the task resource attribute type value, judges the model loading frequency, data reading interval and node distribution, analyzes the execution sequence and scheduling stability, and obtains the subtask resource classification structure;

[0089] According to the task resource attribute type value, extract the subtask log segments of this type, and further deeply analyze the running characteristics of each subtask. For example, in a compute-intensive subtask, it is observed that the model loading frequency is relatively high while the data reading is relatively less. In a data transmission-intensive subtask, the frequency and rhythm of data reading may be more significant. By calling the data in the log segments, the execution sequence and scheduling stability of the subtask within the task running cycle can be analyzed. By analyzing the distribution trend during the model loading process and the continuity of data reading, it can help better understand the performance bottlenecks and optimization points of the subtask, and the analysis enables the generation of the subtask resource classification structure.

[0090] Please refer to Figure 3 , the resource scheduling execution module includes:

[0091] The GPU node selection sub-module calls the subtask resource classification structure, detects the idle status of GPU nodes according to the available labels of the current resource nodes on the cloud platform, filters out the idle GPU nodes and sorts them according to the degree of idleness, and generates a list of idle GPU nodes;

[0092] First, detect the idle state of each GPU node to determine whether the node can be used for the current resource scheduling task. The idle state is determined by monitoring the resource usage of the node, such as GPU memory occupancy rate, computing core utilization rate, etc. If the memory usage rate of a certain GPU node is lower than a certain threshold, for example, lower than 80%, then the node is regarded as an idle node. Based on the idle state, all GPU nodes are sorted according to the degree of idleness, and the nodes with higher idleness are ranked in the front. The nodes are more suitable for the current task. If there are two GPU nodes, the memory occupancy rate of node A is 70% and that of node B is 90%, then node A is considered to have a better idle state and is ranked before node B. Generate a list of idle GPU nodes for subsequent use, so that nodes with higher idleness can be preferentially selected during task scheduling.

[0093] The memory expansion node matching sub-module matches the memory expansion nodes in the cloud platform based on the list of idle GPU nodes, and filters out the qualified nodes according to the memory capacity and idle state to obtain a list of memory expansion nodes;

[0094] Match the memory expansion nodes with the GPU nodes according to their availability to ensure the matching of the computing power and memory resources of each GPU node. During the matching process, first detect the idle memory capacity of the memory expansion node. If the memory capacity is greater than a certain set benchmark value (such as 32GB), then the memory expansion node is regarded as a qualified node. The matching of GPU nodes and memory expansion nodes will be sorted according to the idle memory capacity, and the memory expansion nodes that meet the conditions will be selected. If the memory expansion node of node A has 48GB of memory and the memory expansion node of node B only has 16GB of memory, then the matching degree of node A is higher and it is preferentially matched. Through this process, a list of memory expansion nodes is obtained. Each memory expansion node in this list is the best match with the idle GPU node, ensuring that the resources of each node can meet the requirements of the subtasks.

[0095] The subtask path optimization sub-module calls the list of idle GPU nodes and the list of memory expansion nodes, sorts the GPU and memory expansion nodes according to the task priority and idle state to obtain the subtask node allocation path;

[0096] Sort the nodes according to the task priority and idle status to ensure that tasks can be executed on the most suitable nodes. Sort the tasks to be executed according to the priority. Tasks with higher priorities will be arranged to be executed on the nodes with the best idle status. Suppose the priority of task X is 1, the priority of task Y is 2, and the priority of task Z is 3. Then task Z will be arranged to be executed first. Match the tasks with the nodes according to the task priority, and preferentially select the nodes with a higher degree of idleness. If task X has a high demand for the GPU, but the free memory of this GPU node is small, then a GPU node with a larger free memory will be selected to execute task X to avoid resource bottlenecks, and obtain the subtask node allocation path. This path contains each subtask and its corresponding GPU and memory expansion nodes.

[0097] Please refer to Figure 4 , the resource load monitoring module includes:

[0098] Based on the subtask node allocation path, the running extraction sub-module extracts the start, completion, and interruption times of tasks within the node scheduling period, merges the running sections, and generates the node running time series.

[0099] Extract the task running time and interruption time of each node within the scheduling period. By recording the running time of the node tasks, the resource usage of each node can be accurately grasped. For example, if a certain node runs for 5 hours within the scheduling period, where task A runs for 3 hours and task B runs for 2 hours without interruption, then the task running time of this node is 5 hours. Extract the interruption time of each task. For example, if a certain task has a 1-hour pause due to scheduling interruption, record and mark it as an interruption period. By counting the start time, end time, and interruption time of the tasks, the system can calculate the overall running time and interruption frequency of the nodes. For example, in a node, task A starts running at 0 o'clock, task B starts running at 2 o'clock. After task A ends, task B does not start immediately but has a period of idle time, called the waiting time. Through this method, the resource utilization status of each node can be accurately marked. The information finally helps to generate the node running time series, providing detailed basic data for subsequent analysis and ensuring that the resource utilization of each node at a specific time is comprehensively recorded.

[0100] The interval judgment sub-module calls the node running time series, identifies the waiting time between adjacent tasks, screens out the items that exceed the threshold range and conducts classification judgment, and generates a list of abnormal interval segments.

[0101] By calculating the time difference between tasks, it is determined whether it meets the preset scheduling specifications. For example, the time difference between task A and task B is 15 minutes, while the set node waiting time threshold is 10 minutes. Task intervals exceeding this threshold will be marked as abnormal. Through this judgment, the task intervals under the node can be screened to identify which tasks have idle segments and rearrangement segments. The paragraph indicates that the node does not fully utilize time resources during scheduling. The idle segment exists because task B fails to start immediately after task A finishes running, resulting in idle time. The rearrangement segment is due to an unreasonable task scheduling order, which lengthens the interval time of task execution. According to the information, a list of abnormal interval segments is generated. This list indicates all the node time periods that do not conform to the normal scheduling time specifications, further optimizing the utilization efficiency of node resources.

[0102] The scheduling marking sub-module determines whether the corresponding node scheduling is continuous according to the list of abnormal interval segments. Combining the path length, the number of abnormal segments, and the task density, the formula is used:

[0103]

[0104] Calculate the resource scheduling coherence index, extract the low coherence node sequence, mark the nodes with discontinuous resource utilization, and obtain the resource scheduling coherence state;

[0105] Among them, U represents the resource scheduling coherence index, L represents the node task allocation path length, E represents the number of abnormal interval segments, M represents the node task density, C k represents the sub-task switching frequency within the kth idle segment, q is the total number of idle segments, and W represents the node scheduling waiting times;

[0106] Evaluate the scheduling coherence of each node, mainly considering the node task allocation path length, the number of abnormal interval segments, and the node task density. The task allocation path length refers to the complexity of the node task arrangement. The longer the path, the higher the complexity of task scheduling. The number of abnormal interval segments directly reflects the discontinuity of the node during scheduling. The more the number, the more frequent the interruption of node resource scheduling. The task density refers to the number of tasks processed by the node per unit time. A high task density means a high utilization rate of node resources. Calculate the resource scheduling coherence index of the node according to the parameters;

[0107] Task allocation path length L: This parameter is the total length of all task scheduling paths of the node, indicating the complexity of task scheduling. For example, if the scheduling path of node A includes tasks A, B, and C, and the execution duration of each task is 2 hours, 3 hours, and 1 hour respectively, the path length L is 2 + 3 + 1 = 6 hours;

[0108] Number E of abnormal interval segments: An abnormal interval segment represents an idle segment and a rearrangement segment existing in the task scheduling process. Assuming there are two idle time periods of 30 minutes and 45 minutes respectively between task A and task B in the task scheduling of node B, and there is an idle segment of 60 minutes between task B and task C, so E = 3;

[0109] Task density M: This parameter calculates the density of tasks executed by a node within a scheduling period. Assuming the scheduling period of node A is 10 hours and the total duration of tasks executed is 8 hours, then the task density is

[0110] Frequency C of idle segments k : The frequency of idle segments represents the number of times an idle segment occurs in the task scheduling process of a node. For example, assuming the node experiences 5 idle segments in its scheduling period, so C1 = 5, and there is no idle segment, that is, C2 = 0;

[0111] Number of scheduling waits W: This parameter represents the number of times a node waits during scheduling. If node D has 4 scheduling wait times, each waiting for 30 minutes, then W = 4;

[0112] Numeric assignment and formula calculation:

[0113] According to the actual situation, assume that the task assignment path length L of node A = 6 hours, the number of abnormal interval segments E = 3, the task density M = 0.8, the frequency of idle segments C1 = 5, the total number of idle segments q = 5, and the number of scheduling waits W = 4;

[0114] Substitute the values into the formula for calculation:

[0115]

[0116] The result shows that the scheduling coherence index of node A is 0.746, which is lower than the preset coherence threshold of 1.0. Therefore, this node is marked as a node with discontinuous resource utilization and requires further optimization of the scheduling strategy.

[0117] Please refer to Figure 5 , the cross-cloud task migration module includes:

[0118] The sub-module for extracting task records extracts the sub-task records under each discontinuous node according to the resource scheduling coherence status, analyzes the execution progress information of the sub-tasks, including the current task status, execution duration, remaining duration, and task interruption records, identifies the execution records of associated tasks, and generates a set of sub-task records;

[0119] First, extract the subtask records under the node, and obtain the detailed execution progress information of the subtasks of each discontinuous node from the historical scheduling data, including the current status, execution duration, remaining duration of the task, and the interruption record of the task. For example, task X in node A starts execution at 10 o'clock, is interrupted at 11 o'clock, resumes execution at 12 o'clock, and ends at 13 o'clock. Record the start time, end time, and interruption period of task X. When extracting the subtask records, associate the execution time of the task with the interruption information, record the progress status of each task, analyze each node, calculate the execution duration of each subtask, and determine whether there are abnormal situations such as idle periods and rearrangement periods. If task X is interrupted for more than 1 hour, mark this task as a discontinuous task. Obtain the execution data of all relevant tasks under the node, generate a subtask record set, and provide data support for the subsequent steps.

[0120] The mapping recognition sub-module identifies the candidate mapping situations of the cloud platform nodes according to the subtask record set, matches the resource load, task requirements, and bandwidth limitations between the nodes, judges the task adaptability, filters the nodes that meet the requirements, and generates a candidate mapping result set;

[0121] Identify the candidate mapping situations of the current cloud platform nodes, analyze the requirements of each task for computing resources, storage resources, and bandwidth, evaluate the adaptability by matching with the resource status of the candidate nodes, and calculate according to the resource load, task requirements, and bandwidth limitations of the nodes. For example, if the resource load of node B is 70% and the bandwidth is 100 Gbps, while task X requires 50% computing resources and 40 Gbps bandwidth, then this node will be considered a candidate node suitable for the migration of task X. According to the execution time, remaining duration, task density, etc. of task X, further screen the most suitable node. If the resource load of node A is low and the bandwidth resource is sufficient, preferentially select node A for task migration. By matching the node resources with the task requirements, generate a candidate mapping result set, and provide a reference basis for the subsequent migration decision.

[0122] The migration screening sub-module makes a status switching judgment on each mapping candidate according to the candidate mapping result set, screens the task items according to the task migration conditions and the platform resource status, and uses the formula:

[0123]

[0124] Calculate the task migration adaptability index, identify the task items that meet the migration conditions, and generate a training task migration switching list;

[0125] Among them, P represents the task migration adaptability index, R represents the resource requirement degree of the current task, N represents the number of migrations of the task, Q represents the bandwidth capacity of the candidate node, O represents the utilization rate of the node resources, and Y represents the resource load of the candidate node;

[0126] For each candidate node, perform a status transition judgment to ensure that the status transition during task migration is consistent with the continuity of the task. Analyze the status transition time of the current node and the task migration time, and check whether there are excessive pauses or interruptions during the task migration. Assume that when task X is running on node A, its task status has switched twice, and each switch takes 10 minutes. During the migration process, the status transition time shall not exceed 10 minutes, otherwise the task migration will fail. Verify the adaptability of the resources required for task migration and the node resources. During this process, calculate the resource requirements of the task, including computing resources and bandwidth, to ensure that the candidate node can meet the resource requirements of the migrated task. For example, if the computing requirement of task X is 50% of CPU resources and 40 Gbps of bandwidth, and node B provides 80% of CPU resources and 100 Gbps of bandwidth, then node B can meet the migration requirements. Perform an adaptability judgment based on the migration history of the task (such as the number of task migrations). Assume that the number of migrations of task X is 1. During the migration process, evaluate whether the node has sufficient bandwidth, computing resources, and its load condition to confirm whether the task can be migrated smoothly. For more accurate migration adaptability calculation;

[0127] R (Task Resource Requirement Degree): Represents the requirement of the task for computing resources. Assume that task X requires 50% of computing resources, and the value is 0.5;

[0128] N (Number of Task Migrations): The number of task migrations. Assume that the number of migrations of task X is 1, and the value is 1;

[0129] Q (Bandwidth Capacity of Candidate Node): Represents the bandwidth of the candidate node. Assume that the bandwidth capacity of node B is 100 Gbps, and the value is 100;

[0130] O (Node Resource Utilization Rate): Represents the resource utilization condition of the node. Assume that the resource utilization rate of node A is 60%, and the value is 0.6;

[0131] Y (Resource Load of Candidate Node): Represents the current load of the candidate node. Assume that the resource load of node B is 70%, and the value is 0.7;

[0132] Assume that task X requires 50% of computing resources and 40 Gbps of bandwidth, the bandwidth of node B is 100 Gbps, the number of task migrations N = 1, the resource utilization rate O of node A = 0.6, and the load Y of node B = 0.7. Substitute into the formula for calculation:

[0133]

[0134] This calculation shows that the migration adaptability of task X on node B is relatively low, with an adaptability index of 0.00499. Since this value is very low, the system determines that node B is not suitable for migrating task X. Therefore, it will continue to screen candidate nodes, generate a training task migration and switching list, and preferentially select a more suitable node for task migration based on the adaptability index.

[0135] Please refer to Figure 6 , the performance evaluation and adaptation module includes:

[0136] The task migration comparison sub-module calls the task identifiers in the training task migration and switching list to identify the training logs of the cloud AI model on the source node and the migration node. By comparing and analyzing the training step time consumption and interruption frequency, it conducts a training process evaluation using the formula:

[0137]

[0138] Calculate the total training difference time and generate a task migration comparison result;

[0139] Among them, T train (j) represents the training time of the jth step, T interrupt (j) represents the interruption time of the jth step, m represents the total number of steps, and H represents the total training difference time;

[0140] Call the task identifiers in the training task migration and switching list to obtain the training logs in the source node and the migration node of the cloud AI model, including information such as the training step time consumption and interruption frequency at each stage. The task identifiers will be used to accurately distinguish each task and obtain detailed training data related to the source node and the migration node. The data includes training time, steps, and interruption records, such as the start time, end time, interruption time, and execution duration of each step. Compare the time and interruption frequency of each step in the source node and the migration node, and conduct a comprehensive analysis of the training step time consumption and interruption frequency. First, it is necessary to ensure the unity of all time units. During the calculation process, all time is in seconds to ensure the consistency of dimensions;

[0141] In task A on the source node, the training time consumption of step 1 is 600 seconds (10 minutes). In task A on the migration node, the time consumption of step 1 is 900 seconds (15 minutes). There is no interruption during the training of step 1 in task A on the source node, but in task A on the migration node, there are 2 interruptions in step 1, and the duration of each interruption is 30 seconds. According to the data, the training time difference is that the training time of step 1 of task A on the migration node is 300 seconds (5 minutes) more than that on the source node, and the total interruption duration is 60 seconds (2 interruptions, 30 seconds each). Through the comparison of the data, it can be concluded that the efficiency and stability of the training steps of task A after migration are poor. Through data comparison and calculation;

[0142] Suppose the times of Task A at Step 1 and Step 2 in the migration node and the interruption time are as follows:

[0143] Step 1: The training time is 900 seconds and the interruption time is 60 seconds;

[0144] Step 2: The training time is 800 seconds and the interruption time is 0 seconds;

[0145] Then: H = (900 - 60) + (800 - 0) = 840 + 800 = 1640;

[0146] The total training time difference of Task A in the migration node in these two steps is 1640 seconds. Through this calculation method, the task migration comparison results are finally obtained, including the time differences in the task migration process, and then providing input for the subsequent stability screening sub-module. This step ensures the consistency of time units, the accuracy of data, and the comparability of the final calculation results.

[0147] Based on the task migration comparison results, the stability screening sub-module screens tasks with fewer interruption frequencies and stable training step durations during the training process, records the screening criteria, marks the tasks that meet the conditions, and generates a set of screened stable tasks;

[0148] Screen tasks that perform relatively stably during the training process. The stability of the tasks is judged based on the duration differences and interruption frequencies of the training steps. Pay special attention to whether there are significant time fluctuations, whether there are multiple interruptions, or whether the training time of some steps is much higher than that of other steps during the training process of each task. For the setting of screening conditions, a judgment criterion based on time fluctuations is adopted. If a certain task has a step duration exceeding 5 minutes and an interruption frequency higher than 3 times during the training, it is considered to have poor stability and needs to be excluded. In the migration node of Task B, the training time in Step 2 and Step 4 exceeds 10 minutes respectively and there are 3 interruptions, while the steps in the source node are relatively stable time periods. Through the screening rules, a set of screened stable tasks is generated.

[0149] According to the set of screened stable tasks, the scheduling matching generation sub-module screens tasks that match the platform optimization, determines the matching priority of each task, and obtains the AI model scheduling matching mapping table;

[0150] Verify the matching degree of each task with platform optimization. The process mainly evaluates according to the stability, efficiency of the task and its matching degree with platform requirements. By conducting matching analysis on each task, determine whether it meets the requirements of platform optimization, including the time efficiency, stability and resource requirements of the task, etc. During the scheduling matching generation process, according to the platform's resource management and task scheduling algorithms, adjust the priority of tasks, allocate nodes, etc., to ensure the efficiency and rationality of the overall scheduling process. For example, for task C, the task log after migration shows that its training time and interruption frequency are both within a reasonable range, and the resource requirements meet the resource availability of the platform. Finally, it is marked as an optimized matching item and corresponds to the corresponding resource node in the mapping table, obtaining the AI model scheduling matching mapping table.

[0151] The above are only the preferred embodiments of the present invention, and do not limit the present invention in other forms. Any person skilled in the art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A data training and optimization system based on a cloud AI model, characterized in that, The system includes: Based on the model training task logs of the cloud AI platform, the task feature recognition module extracts task scheduling trajectories, resource usage, processing duration, model loading frequency, and data reading rhythm, classifies task types according to GPU usage frequency and data access density, analyzes the running characteristics of subtasks, and obtains the subtask resource classification structure; The resource scheduling execution module calls the subtask resource classification structure, compares it with the available labels of the current resource nodes in the cloud platform, matches the GPU identification nodes and memory expansion nodes, sorts the same type of nodes according to the idle state, and arranges tasks in the order of priority to obtain the subtask node allocation path; Based on the subtask node allocation path, the resource load monitoring module extracts the task running time and interruption records during the node scheduling period, judges the waiting interval between consecutive subtasks under the same node, and identifies the idle section and rearrangement section in the resource queue to obtain the resource scheduling coherence state; According to the resource discontinuous nodes marked in the resource scheduling coherence state, the cross-cloud task migration module identifies the candidate mapping situation of cloud platform nodes, judges whether the state switch is consistent with task continuity, and generates a training task migration and switch list.

2. The data training and optimization system based on the cloud AI model according to claim 1, wherein The subtask resource classification structure includes a computational intensity label, a data transmission intensity label, a model processing load distribution, a resource call hierarchy structure, and a data reading frequency identifier. The subtask node allocation path includes a GPU node allocation sequence, a memory expansion node list, a task mapping priority identifier, and a node resource matching comparison table. The resource scheduling coherence state includes a resource idle section identifier, a task rearrangement section identifier, a node scheduling discontinuity label, and a waiting time distribution structure. The training task migration and switch list includes a subtask correspondence mapping, a node migration path structure, a task continuity preservation label, and a migratable condition identifier.

3. The data training and optimization system based on the cloud AI model according to claim 1, characterized in that The task feature recognition module includes: Based on the model training task logs of the cloud AI platform, the scheduling trajectory analysis sub-module extracts the GPU call sequence, data loading timing, and task running duration, and identifies the task density of the GPU call sequence in each time period to obtain the task scheduling density value; The resource classification judgment sub-module calls the task scheduling density value, combines the GPU usage frequency and data access frequency, analyzes the impact of the GPU frequency on the training cycle, judges whether it exceeds the GPU and data access thresholds, and uses the formula: Obtain the task resource attribute type value; where T represents the task resource attribute type value, f i represents the GPU usage frequency in the i-th time period, and f represents the average value of the GPU usage frequency training cycle, d i represents the data access frequency in the i-th time period, s i represents the data access rhythm reference value in the i-th time period, and n represents the number of task running time periods; According to the task resource attribute type value, the subtask structure extraction sub-module extracts the subtask log segments in the corresponding tasks, judges the model loading frequency, data reading interval, and node distribution, analyzes the execution sequence and scheduling stability, and obtains the subtask resource classification structure.

4. The data training and optimization system based on the cloud AI model according to claim 3, wherein The resource scheduling execution module includes: The GPU node selection sub-module calls the subtask resource classification structure, detects the idle state of the GPU nodes according to the available labels of the current resource nodes in the cloud platform, filters out the idle GPU nodes and sorts them according to the degree of idleness to generate a list of idle GPU nodes; Based on the list of idle GPU nodes, the memory expansion node matching sub-module matches the memory expansion nodes in the cloud platform, filters the eligible nodes according to the memory capacity and idle status, and obtains the list of memory expansion nodes; The sub-task path optimization sub-module calls the list of idle GPU nodes and the list of memory expansion nodes, sorts the GPU and memory expansion nodes by task priority and idle status, and obtains the sub-task node allocation path.

5. The data training and optimization system based on the cloud AI model according to claim 4, wherein The resource load monitoring module includes: Based on the sub-task node allocation path, the running extraction sub-module extracts the start, completion, and interruption times of tasks during the node scheduling period, merges the running segments, and generates the node running time series; The interval judgment sub-module calls the node running time series, identifies the waiting time between adjacent tasks, filters the items outside the threshold range and classifies and judges them, and generates the list of abnormal interval segments; Based on the list of abnormal interval segments, the scheduling marking sub-module judges whether the corresponding node scheduling is continuous, combines the path length, the number of abnormal segments, and the task density, and uses the formula: Calculate the resource scheduling coherence index, extract the low coherence node sequence, mark the nodes with discontinuous resource utilization, and obtain the resource scheduling coherence status; Among them, U represents the resource scheduling coherence index, L represents the length of the node task allocation path, E represents the number of abnormal interval segments, M represents the node task density, and C k represents the sub-task switching frequency within the k-th idle segment, q is the total number of idle segments, and W represents the node scheduling waiting times.

6. The data training and optimization system based on the cloud AI model according to claim 5, wherein The cross-cloud task migration module includes: Based on the resource scheduling coherence status, the task record extraction sub-module extracts the sub-task records under each discontinuous node, analyzes the execution progress information of the sub-tasks, including the current task status, execution duration, remaining duration, and task interruption records, identifies the execution records of related tasks, and generates the sub-task record set; Based on the sub-task record set, the mapping identification sub-module identifies the candidate mapping situations of the cloud platform nodes, matches the resource load, task requirements, and bandwidth limitations between the nodes, judges the task adaptability, filters the eligible nodes, and generates the candidate mapping result set; Based on the candidate mapping result set, the migration screening sub-module makes a status switching judgment on each mapping candidate, screens the task items according to the task migration conditions and the platform resource status, and uses the formula: Calculate the task migration adaptability index, identify the task items that meet the migration conditions, and generate the training task migration switching list; Among them, P represents the task migration adaptability index, R represents the resource demand degree of the current task, N represents the number of task migrations, Q represents the bandwidth capacity of the candidate node, O represents the utilization rate of the node resources, and Y represents the resource load of the candidate node.

7. The data training and optimization system based on the cloud AI model according to claim 1, characterized in that, The system also includes a performance evaluation adaptation module: The performance evaluation adaptation module calls the task identifiers in the training task migration switching list, compares the training logs of the cloud AI model on the source node and the migration node, analyzes the training step duration and interruption frequency, screens the tasks with consistent stability, and marks them as platform optimization matching items to obtain the AI model scheduling matching mapping table; The AI model scheduling matching mapping table includes the training stability scoring result, the task execution consistency identifier, the migration node performance matching degree, and the model scheduling optimization item.

8. The data training and optimization system based on the cloud AI model according to claim 7, wherein, The performance evaluation adaptation module includes: The task migration comparison sub-module calls the task identifiers in the training task migration and switching list, identifies the training logs of the cloud AI model at the source node and the migration node, evaluates the training process by comparing and analyzing the time consumption of training steps and the interruption frequency, and uses the formula: Calculate the total training difference time and generate the task migration comparison result; where, T train (j) represents the training time of the j-th step, T interrupt (j) represents the interruption time of the j-th step, m represents the total number of steps, and H represents the total training difference time; The stability screening sub-module based on the task migration comparison result, screens the tasks with few interruption frequencies and stable training step time consumption during the training process, records the screening criteria, marks the eligible tasks, and generates a set of stable tasks after screening; The scheduling matching generation sub-module screens the tasks that match the platform optimization according to the set of stable tasks after screening, determines the matching priority of each task, and obtains the AI model scheduling matching mapping table.

Citation Information

Cited By

  • Monitoring management method and system for artificial intelligence development platform

    CN121722645A

  • A monitoring management method and system of an artificial intelligence development platform

    CN121722645B