Computing resource optimization method, system and storage medium based on AI big model
By building a task processing time prediction model and progress evaluation method, and using AI large models to dynamically adjust computing power resources, the problem of inability to regulate computing power resources in the existing technology is solved, and efficient resource utilization and task completion are achieved.
Patent Information
- Application Number
- CN202510112134.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-24
AI Technical Summary
The existing computing power resource regulation and optimization technologies cannot predict the task processing time while using multi-dimensional data to accurately evaluate the task progress, so as to dynamically regulate and optimize computing power resources according to real-time changes in the task progress during the task processing process.
By collecting task sample data of computer processing tasks, pre-processing data, building a task processing time prediction model and task processing progress evaluation method, and using AI large models to dynamically adjust computing resources during task processing.
While predicting the task processing time, it is realized that multi-dimensional data is used to accurately evaluate the task progress, so that during the task processing process, it dynamically regulates and optimizes computing resources according to real-time changes, improves resource utilization efficiency, and avoids idle or excessive use of resources.
Smart Images

Figure CN119557108B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computing power resource regulation and optimization, and specifically to a computing power resource optimization method, system and storage medium based on an AI big model. Background Art
[0002] Computing power resource regulation and optimization technology refers to a collection of technical means for effectively managing, regulating and optimizing computing power resources. With the rapid development of information technology, the amount of data has exploded, and the rational use of computing power resources has become the key to improving system performance and efficiency. If computing power resources cannot be effectively regulated and optimized, problems such as resource waste, task response delays, and overall system performance degradation may occur.
[0003] Existing computing power resource control and optimization technologies often predict the computing power resource requirements of the task, and then perform computing power resource control and optimization based on the prediction results. This computing power resource control and optimization method is often completed before the task, and cannot be controlled in real time during the task process, cannot cope with emergencies, and has low robustness. Other computing power resource control and optimization methods are based on fixed algorithms or manual intervention, and cannot adjust computing power resources in time according to real-time changes in task progress; for example, in the patent application with publication number CN117785487A, a computing power resource scheduling method, device, equipment and medium are disclosed. This solution uses a fixed rule algorithm to control computing power resources, which cannot be adjusted during the task. The computing resources are adjusted according to the real-time changes of task progress during the progress; and the existing task progress evaluation technology may only focus on one aspect or the evaluation method is too complicated and requires a lot of computing resources to be wasted, such as evaluating the progress based only on the completion of subtasks or only on the processing ratio of data volume; this one-sided evaluation may lead to the inability to accurately judge the task progress in some complex tasks, thereby affecting the reasonable regulation of computing resources; therefore, the existing computing resource regulation and optimization technology cannot use multi-dimensional data to perform a simple and accurate evaluation of task progress while predicting the processing time of the task, so as to dynamically regulate and optimize computing resources according to the real-time changes of task progress during task processing. Summary of the invention
[0004] The present invention aims to solve one of the technical problems in the prior art to at least a certain extent, by collecting task sample data of computer processing tasks, performing data preprocessing, and then constructing a task processing time prediction model to predict the time required for the computer to process the task; constructing a task processing progress evaluation method to perform calculation and evaluation on the progress of the computer processing task; and dynamically adjusting and optimizing computing power resources during the computer processing task; so as to solve the problem that the existing computing power resource regulation and optimization technology cannot accurately evaluate the task progress using multi-dimensional data while predicting the processing time of the task, thereby dynamically regulating and optimizing the computing power resources according to the real-time changes of the task progress during the task processing process.
[0005] To achieve the above objectives, in a first aspect, the present application provides a computing resource optimization method based on an AI big model, comprising the following steps:
[0006] Collecting task sample data of a computer processing task, and performing data preprocessing on the task sample data to obtain first sample data;
[0007] Building a task processing time prediction model based on the first sample data, predicting the time required for the computer to process the task, and obtaining task duration information;
[0008] Construct a task processing progress evaluation method to calculate and evaluate the progress of computer processing tasks and obtain task progress information;
[0009] Based on task duration information and task progress information, computing resources are dynamically adjusted and optimized during the computer processing task.
[0010] Furthermore, collecting task sample data of the computer processing task and performing data preprocessing on the task sample data to obtain first sample data includes the following sub-steps:
[0011] Obtaining task sample data of each processing task from the computer task execution record, the task sample data including: CPU requirement, GPU requirement, memory requirement, task data size, and time required to complete the task;
[0012] The total resource demand is calculated according to the CPU demand, GPU demand and memory demand of each processing task. The total resource demand calculation formula is as follows: total resource demand = CPU demand + GPU demand + memory demand. Then, the proportion of CPU demand, GPU demand and memory demand in the total resource demand is calculated respectively, and the proportions are compared to obtain the maximum proportion. The processing tasks are divided into task types according to the demand corresponding to the maximum proportion, and marked. The task types include: CPU demand type, GPU demand type and memory demand type. The CPU demand type is coded as 001, the GPU demand type is coded as 010, and the memory demand type is coded as 100.
[0013] The task data size, time required to complete the task, and task type of each processing task are stored according to the corresponding processing task classification and marked as first sample data.
[0014] Furthermore, constructing a task processing time prediction model based on the first sample data includes the following sub-steps:
[0015] The CPU demand, GPU demand, memory demand and task data size in the first sample data are normalized according to the data type, and the data size is scaled to [0, 1]. After completion, the second sample data is obtained;
[0016] The CPU demand, GPU demand, memory demand, task data size and task type under the same processing task in the second sample data are combined into a feature vector, denoted as Ai, where i represents the i-th processing task; after completion, all feature vectors and the corresponding task completion time are stored together and marked as the third sample data;
[0017] Constructing a task processing time prediction model includes: setting the number of input layer neurons to b1, setting the number of hidden layer neurons to b2, setting the number of output layer neurons to b3 based on a multi-layer perceptron, adding an attention mechanism layer between the input layer and the hidden layer, and obtaining a task processing time prediction model after completion.
[0018] Furthermore, predicting the time required for the computer to process the task and obtaining the task duration information includes the following sub-steps:
[0019] Divide the third sample data into a training sample set and a test sample set in a ratio of 3:1;
[0020] Set the learning rate of the task processing time prediction model training to c1, the number of training iterations to c2, and the batch size to c3; set the model training loss function as follows: , where M is the model training loss function, n is the number of feature vectors input to the model, Xi is the actual time required to complete the task, and Yi is the time required to complete the task predicted by the model;
[0021] The task processing time prediction model is trained using the training sample set. After completing c2 training iterations, the first time prediction model is obtained. The first time prediction model is tested using the test sample set, and the prediction fit index of the first time prediction model is calculated. The prediction fit index calculation formula is as follows: , where R1 is the prediction fit index, Xj is the actual time required to complete the task, Z is the average time required to complete the task corresponding to all feature vectors in the test sample set, Yj is the time required to complete the task predicted by the first-time prediction model, and m is the number of feature vectors input into the first-time prediction model.
[0022] Furthermore, predicting the time required for the computer to process the task and obtaining the task duration information includes the following sub-steps:
[0023] Set the prediction fitting threshold to R0. If R1 is less than R0, use the training sample set to train the first-time prediction model again. After completion, use the test sample set to test again and calculate the prediction fitting index R1 until R1 is not less than R0. If R1 is not less than R0, mark the first-time prediction model as the second-time prediction model.
[0024] When the computer receives a processing task, it obtains the processing task and the task data size, and obtains the task type of the processing task based on the CPU demand, GPU demand and memory demand; then the task data size and the task type of the processing task are combined into a feature vector, which is input into the second time prediction model to obtain the predicted time required to complete the task, which is marked as the task duration information.
[0025] Furthermore, a task processing progress evaluation method is constructed to calculate and evaluate the progress of the computer processing task, and obtaining the task progress information includes the following sub-steps:
[0026] Split the processing task received by the computer into several processing subtasks, and record the subtask data size corresponding to the processing subtask, record the total number of processing subtasks as v, record any processing subtask as Wk, and record the task data size corresponding to the processing subtask Wk as Sk, where k represents the kth processing subtask;
[0027] When the computer starts processing the processing task, it counts the number of completed processing subtasks in real time, denoted as u; and obtains the total size of the subtask data corresponding to the completed processing subtasks, denoted as Z1.
[0028] Furthermore, constructing a task processing progress evaluation method to calculate and evaluate the progress of the computer processing task to obtain task progress information also includes the following sub-steps:
[0029] Set the subtask progress weight to T1 and the data progress weight to T2;
[0030] The subtask processing progress P1 is calculated based on the total number of subtasks v and the number of subtasks u that have been processed. The subtask processing progress formula is as follows: ;
[0031] The data processing progress P2 is calculated based on the task data size Z0 of the processing task and the total size Z1 of the subtask data corresponding to the completed processing subtask. The data processing progress calculation formula is as follows: ;
[0032] The overall task processing progress P0 is calculated based on the subtask processing progress P1 and the data processing progress P2. The overall task processing progress calculation formula is as follows: ; Mark the overall task processing progress P0 as task progress information.
[0033] Furthermore, based on the task execution time prediction model and the task processing progress evaluation method, the dynamic adjustment and optimization of computing resources during the computer processing task includes the following sub-steps:
[0034] After the computer starts processing the task, it obtains the corresponding task progress information at a first time interval, and records the number of times the corresponding task progress information is obtained, which is recorded as Q, and the first time interval is e;
[0035] Based on the task duration information corresponding to the processing task, the task processing time deviation BC is calculated. The calculation formula for the task processing time deviation is as follows: , where A represents the predicted time required to complete the task;
[0036] If BC is less than 0, the computing power resources of the corresponding processing task will be reduced based on the initial total resource demand of the corresponding processing task. If BC is greater than 0, the computing power resources of the corresponding processing task will be increased based on the initial total resource demand of the corresponding processing task. If BC is equal to 0, no adjustment will be made. The calculation formula for the reduction ratio and the increase ratio is as follows , where when BL is less than 0, it represents a decrease ratio, and when BL is greater than 0, it represents an increase ratio.
[0037] In the second aspect, the present application provides a computing resource optimization system based on an AI big model, including a data acquisition module, a time prediction module, a progress evaluation module, and a computing power adjustment module;
[0038] The data acquisition module includes a collection unit and a processing unit, the collection unit is used to collect task sample data of the computer processing task, and the processing unit is used to perform data preprocessing on the task sample data to obtain first sample data;
[0039] The time prediction module includes a model unit and a prediction unit. The model unit constructs a task processing time prediction model based on the first sample data; the prediction unit is used to predict the time required for the computer to process the task and obtain task time information;
[0040] The progress evaluation module is used to construct a task processing progress evaluation method, calculate and evaluate the progress of computer processing tasks, and obtain task progress information;
[0041] The computing power adjustment module dynamically adjusts and optimizes computing power resources during the computer processing task based on task duration information and task progress information.
[0042] In a third aspect, the present application provides a storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps in the above method are performed.
[0043] Beneficial effects of the present invention: The present invention obtains first sample data by collecting task sample data of computer processing tasks and performing data preprocessing on the task sample data; constructs a task processing time prediction model based on the first sample data, predicts the time required for the computer to process the task, and obtains task time information; constructs a task processing progress evaluation method, calculates and evaluates the progress of the computer processing task, and obtains task progress information; dynamically adjusts and optimizes computing power resources during the computer processing task based on the task time information and task progress information; can accurately evaluate the task progress by using multi-dimensional data while predicting the processing time of the task, so as to dynamically adjust and optimize computing power resources according to the real-time changes of the task progress during the task processing process;
[0044] The present invention collects multiple feature data, uses AI technology, and adds an attention mechanism to predict the time required to complete a task. It can focus on key features, adapt to various task changes, and have good generalization and robustness while ensuring prediction accuracy. The processing progress is evaluated by combining the subtasks and data volume of the processing task. The advantage is that there is no need to perform a large amount of calculations, and the two important dimensions of the task structure and data processing are taken into account. While ensuring the accuracy of the evaluation, the complexity of the evaluation is reduced, and it can be applied to various types of computing tasks. By adjusting the computing power resources in real time during the task processing process through predicting time and progress evaluation, dynamic optimization of computing power resources can be achieved. Because the computing power resources required during the task processing process are not constant, this real-time regulation can improve the utilization efficiency of computing power resources, avoid idle or excessive use of resources, and ensure that the task is completed within the specified time. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 is a functional block diagram of the system of the present invention;
[0046] Figure 2 is a flow chart of the steps of the method of the present invention;
[0047] Figure 3 It is a structural diagram of the task processing time prediction model of the present invention;
[0048] Figure 4 is a flow chart of the progress evaluation strategy of the present invention;
[0049] Figure 5 It is a schematic structural diagram of the electronic device of the present invention. DETAILED DESCRIPTION
[0050] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0051] Example 1, please refer to Figure 1 As shown, the present application provides a computing resource optimization system based on an AI big model, including a data acquisition module, a time prediction module, a progress evaluation module, and a computing power adjustment module;
[0052] The data acquisition module includes a collection unit and a processing unit, the collection unit is used to collect task sample data of the computer processing task, and the processing unit is used to perform data preprocessing on the task sample data to obtain first sample data;
[0053] The collection unit is configured with a collection strategy, which includes: obtaining task sample data of each processing task from the computer task execution record, and the task sample data includes: the CPU demand, GPU demand, memory demand, task data size and time required for task completion of the processing task; the unit of CPU demand, GPU demand and memory demand is percentage, which reflects the proportion of the degree of use of these hardware resources at a certain moment to their total capacity or total performance, and the use of percentage as a unit is convenient for subsequent processing and calculation;
[0054] The processing unit is configured with a processing strategy, which includes: calculating the total resource demand according to the CPU demand, GPU demand and memory demand of each processing task, and the total resource demand calculation formula is as follows: total resource demand = CPU demand + GPU demand + memory demand, and then calculating the proportion of CPU demand, GPU demand and memory demand in the total resource demand respectively, and comparing the proportions to obtain the maximum proportion, and dividing the processing tasks into task types according to the demand corresponding to the maximum proportion, and marking them. The task types include: CPU demand type, GPU demand type and memory demand type. For example, the CPU demand of a task is 60%, the GPU demand is 9% and the memory demand is 12%, then the total resource demand = 60% + 9% + 12% = 81%, among which the CPU demand accounts for the highest proportion, then the task type is CPU demand type, and the CPU demand type is coded as 001, the GPU demand type is coded as 010, and the memory demand type is coded as 100; the task type is coded so that the computer can clearly distinguish the task type, which is convenient for subsequent model learning and prediction;
[0055] The CPU requirement, GPU requirement, memory requirement, task data size, time required to complete the task, and task type of each processing task are classified and stored according to the corresponding processing tasks, and marked as first sample data;
[0056] In the specific implementation process, by dividing the task types, we can distinguish tasks of different natures, such as data processing tasks, machine learning tasks, graphics rendering tasks, etc.; different types of tasks have very different ways and efficiencies in the use of computing resources; graphics rendering tasks usually have high demands on GPU resources for large-scale rendering operations, while data processing tasks may focus more on the efficient use of CPU and memory for reading, writing and converting data; task completion time is affected by many factors, and the importance and interrelationships of these factors vary for different task types; after dividing the task types, the key influencing factors of the completion time of different types of tasks can be analyzed separately, thereby improving the prediction accuracy of task completion time.
[0057] The time prediction module includes a model unit and a prediction unit. The model unit constructs a task processing time prediction model based on the first sample data. The prediction unit is used to predict the time required for the computer to process the task and obtain task time information.
[0058] The model unit is configured with a model building strategy, which includes: normalizing the CPU demand, GPU demand, memory demand and task data size in the first sample data according to the data type, scaling the data size to [0, 1], because the CPU demand, GPU demand and memory demand units are percentages, they can be directly converted into decimals to complete the normalization, and the task data size can be normalized according to the proportion of the maximum storage space of the computer; after completion, the second sample data is obtained;
[0059] The CPU demand, GPU demand, memory demand, task data size and task type under the same processing task in the second sample data are combined into a feature vector, recorded as Ai, where i represents the i-th processing task; for example, the CPU demand under a second processing task is 0.52, the GPU demand is 0.13, the memory demand is 0.23, the task data size is 0.20, and the task type is CPU demand type, which is coded as 001. Then the feature vector A2 of the processing task = {0.52, 0.13, 0.23, 0.20, 0, 0, 1}; after completion, all feature vectors are stored together with the corresponding task completion time, marked as the third sample data;
[0060] See also Figure 3 As shown, constructing a task processing time prediction model includes: setting the number of input layer neurons to b1, setting the number of hidden layer neurons to b2, setting the number of output layer neurons to b3 based on a multi-layer perceptron, adding an attention mechanism layer between the input layer and the hidden layer, and obtaining a task processing time prediction model after completion; in this embodiment, b1 is 7, because there are 7 components in the feature vector Ai; b2 is 12, b2 can be dynamically adjusted according to the actual application scenario, and b3 is 1, because the output result has only one prediction time;
[0061] The prediction unit is configured with a time prediction strategy, which includes: dividing the third sample data into a training sample set and a test sample set in a ratio of 3:1;
[0062] Set the learning rate of the task processing time prediction model training to c1, the number of training iterations to c2, and the batch size to c3; set the model training loss function as follows: , where M is the model training loss function, n is the number of feature vectors input to the model, Xi is the actual time required to complete the task, and Yi is the time required to complete the task predicted by the model; in this embodiment, c1=0.001, c2=200, c2=16, that is, 16 feature vectors are input each time;
[0063] The task processing time prediction model is trained using the training sample set. After completing c2 training iterations, the first time prediction model is obtained. The first time prediction model is tested using the test sample set, and the prediction fit index of the first time prediction model is calculated. The prediction fit index calculation formula is as follows: , where R1 is the prediction fit index, Xj is the actual time required to complete the task, Z is the average time required to complete the task corresponding to all feature vectors in the test sample set, Yj is the time required to complete the task predicted by the first-time prediction model, and m is the number of feature vectors input into the first-time prediction model; the prediction fit index measures the fit effect between the model's prediction results and the actual results. The closer its value is to 1, the better the model fit effect is, that is, the higher the accuracy of the model is;
[0064] The prediction fitting threshold is set to R0. If R1 is less than R0, the first-time prediction model is trained again using the training sample set. After completion, the test sample set is used again to test, and the prediction fitting index R1 is calculated until R1 is not less than R0; if R1 is not less than R0, the first-time prediction model is marked as the second-time prediction model; in this embodiment, R0 is 0.85. In some preliminary data analysis or simple prediction scenarios, if it can reach 0.5-0.6, it can be considered that the model has a certain prediction ability, and if it reaches 0.7-0.8 or above, it means that the model can better capture the relationship between input features and task completion time, and can provide relatively reliable predictions;
[0065] When the computer receives a processing task, it obtains the processing task and the task data size, and obtains the task type of the processing task according to the CPU requirement, GPU requirement and memory requirement; then, the task data size and the task type of the processing task are combined into a feature vector, which is input into the second time prediction model to obtain the predicted time required for task completion, which is marked as task duration information;
[0066] In the specific implementation process, the attention mechanism can automatically learn the relative importance of each input feature, such as CPU demand, GPU demand, memory demand, task type and task data size, to the task completion time. For example, for a deep learning training task, the model may find that the GPU demand is the most critical factor through the attention mechanism, and thus assign a higher weight to it. This ability to automatically focus on key resource requirements enables the model to better adapt to the characteristics of different types of tasks, because different tasks have different degrees of dependence on various resources. Even when the nature of the task changes, for example, from mainly data processing tasks to mainly graphics rendering tasks, the model can dynamically adjust the degree of attention to each feature according to the characteristics of the new task. This helps to improve the flexibility and adaptability of the model, so that it can still accurately capture the core factors that affect the task completion time when faced with a variety of task combinations, and improve the robustness of the model. The hidden layers can be increased or decreased according to the actual application scenario. For example, two hidden layers are added, and the number of neurons in each hidden layer can be set separately according to the actual application scenario.
[0067] The progress evaluation module is used to construct a task processing progress evaluation method, calculate and evaluate the progress of computer processing tasks, and obtain task progress information;
[0068] See also Figure 4 As shown, the progress evaluation module is configured with a progress evaluation strategy, which includes: splitting the processing task received by the computer into several processing subtasks, and recording the subtask data size corresponding to the processing subtasks, recording the total number of processing subtasks as v, recording any processing subtask as Wk, and recording the task data size corresponding to the processing subtask Wk as Sk, where k represents the kth processing subtask; for example, the image processing task is divided into subtasks such as image segmentation, grayscale processing, feature extraction, and classification recognition; the division should be as detailed as possible;
[0069] When the computer starts processing the processing task, it counts the number of subtasks that have been processed in real time, which is recorded as u; and obtains the total size of the subtask data corresponding to the subtasks that have been processed, which is recorded as Z1;
[0070] The subtask progress weight is set to T1 and the data progress weight is set to T2. In this embodiment, the subtask progress weight T1=0.6 and the data progress weight T2=0.4. The subtask processing progress P1 is calculated based on the total number of processing subtasks v and the number of processing subtasks u that have been processed. The subtask processing progress formula is as follows: ; For example, the total number of processing subtasks v = 10, and the number of processing subtasks that have been processed u = 2, then P1 = 2 / 10*100% = 20.0%;
[0071] The data processing progress P2 is calculated based on the task data size Z0 of the processing task and the total size Z1 of the subtask data corresponding to the completed processing subtask. The data processing progress calculation formula is as follows: ;in, , that is, the sum of the data of all processing subtasks is equal to the task data size of the processing task. For example, Z0=1245M, Z1=299M, then P2=298 / 1245*100%=23.9%;
[0072] The overall task processing progress P0 is calculated based on the subtask processing progress P1 and the data processing progress P2. The overall task processing progress calculation formula is as follows: ; Mark the overall task processing progress P0 as the task progress information; for example, T1=0.6, T2=0.4, P1=20.0%, P2=24.9%, then P0=0.6*20.0+0.4*23.9=21.59%; The advantage of the task processing progress evaluation method is that it balances complexity and accuracy by taking into account both the task structure and the data processing dimension with fixed weights, and it has wide applicability and can effectively avoid the problems of single dimension failure and narrow applicability;
[0073] In the specific implementation process, for the setting of subtask progress weights and data progress weights, usually the subtask progress weight is slightly greater than the data progress weight; because the completion of a subtask usually represents the advancement of the computing task in the key logical stage, the completion of a subtask, such as the successful completion of the compilation stage, is an important milestone, which ensures the integrity of the task in the overall architecture. In contrast, the data processing progress may be affected by the characteristics of the data itself, such as data quality, data distribution, etc.; taking the data cleaning task as an example, the data in the first half may be relatively regular and the data processing progress is very fast, but the data in the second half may contain a large number of errors or outliers, resulting in a slower data processing progress; and the completion of the subtask can relatively better reflect the progress of the task in the core process, so it is given a higher weight; but it can also be flexibly set according to the actual application scenario.
[0074] The computing power adjustment module dynamically adjusts and optimizes computing power resources during computer processing tasks based on task duration information and task progress information;
[0075] The computing power adjustment module is configured with a computing power adjustment strategy, which includes: after the computer starts processing the processing task, the corresponding task progress information is obtained at a first time interval, and the number of times the corresponding task progress information is obtained is recorded as Q, and the first time interval is e; in this embodiment, the first time interval e is 3 minutes;
[0076] Based on the task duration information corresponding to the processing task, the task processing time deviation BC is calculated. The calculation formula for the task processing time deviation is as follows: , where A represents the predicted time required to complete the task; for example, the first time interval e is 3 minutes, and the predicted time required to complete a certain processing task A is 25 minutes. When the fourth acquisition, i.e., Q=4, is obtained, the task progress information is obtained, i.e., the overall task processing progress P0=40.4%, then BC=4*3-25*40.4%=1.9;
[0077] If BC is less than 0, the computing power resources of the corresponding processing task will be reduced based on the initial total resource demand of the corresponding processing task. If BC is greater than 0, the computing power resources of the corresponding processing task will be increased based on the initial total resource demand of the corresponding processing task. If BC is equal to 0, no adjustment will be made. The calculation formula for the reduction ratio and the increase ratio is as follows , where BL is less than 0, it indicates a decrease ratio, and BL is greater than 0, it indicates an increase ratio; for example, for a processing task, Q=4, e=3 minutes, A=25 minutes, P0=40.4%, then BL=[(4*3 / 25*40.4%)-1]*100%=18.8%, and BL is greater than 0, it indicates an increase in computing resources. If the initial total resource requirements of the processing task are 60% of the CPU requirements, 9% of the GPU requirements, and 12% of the memory requirements, it means that the computing resources initially provided to the processing task are 60% of the CPU computing power, 9% of the GPU computing power, and 12% of the memory. Then, they are increased by 18% on the original basis, that is, they become 1.18 times the original. After adjustment, the computing resources provided become 1.18*60% of the CPU computing power, 1.18*9% of the GPU computing power, and 1.18*12% of the memory.
[0078] In the specific implementation process, by calculating in real time the difference between the actual task progress time and the theoretical task progress time, it is possible to promptly discover whether the task progress deviates from expectations; for example, in a big data analysis task, if the actual time is longer than the theoretical time, it means that the computing resources are insufficient; if the actual time is less than the theoretical time, there may be a waste of computing resources; by increasing or decreasing the computing resources for processing tasks in real time based on this gap, dynamic optimization of computing resources can be achieved; this real-time regulation can improve the utilization efficiency of computing resources, avoid idle or excessive use of resources, and ensure that tasks are completed within the expected time.
[0079] Example 2, please refer to Figure 2 As shown, the present application provides a computing resource optimization method based on an AI large model, comprising the following steps:
[0080] Step S1, collecting task sample data of a computer processing task, and performing data preprocessing on the task sample data to obtain first sample data; Step S1 includes the following sub-steps:
[0081] Step S101, obtaining task sample data of each processing task from the computer task execution record, the task sample data including: CPU requirement, GPU requirement, memory requirement, task data size and time required to complete the task;
[0082] Step S102, calculating the total resource requirement according to the CPU requirement, GPU requirement and memory requirement of each processing task, the total resource requirement calculation formula is as follows: total resource requirement = CPU requirement + GPU requirement + memory requirement;
[0083] Step S103, respectively calculating the proportions of the CPU demand, GPU demand, and memory demand in the total resource demand, and comparing the proportions to obtain the maximum proportion;
[0084] Step S104, classifying the processing tasks into task types according to the demand corresponding to the maximum proportion, and marking them. The task types include: CPU demand type, GPU demand type, and memory demand type. The CPU demand type is coded as 001, the GPU demand type is coded as 010, and the memory demand type is coded as 100.
[0085] Step S105 , the task data size, time required to complete the task, and task type of each processing task are stored according to the corresponding processing task classification and marked as first sample data.
[0086] Step S2, building a task processing time prediction model based on the first sample data, predicting the time required for the computer to process the task, and obtaining task duration information; Step S2 includes the following sub-steps:
[0087] Step S201, normalizing the CPU demand, GPU demand, memory demand and task data size in the first sample data according to the data type, scaling the data size to [0, 1], and obtaining the second sample data after completion;
[0088] Step S202, combining the CPU demand, GPU demand, memory demand, task data size and task type under the same processing task in the second sample data into a feature vector, denoted as Ai, where i represents the i-th processing task; after completion, all feature vectors and the corresponding task completion time are stored together, marked as the third sample data;
[0089] Step S203, constructing a task processing time prediction model includes: setting the number of input layer neurons to b1, setting the number of hidden layer neurons to b2, setting the number of output layer neurons to b3 based on a multi-layer perceptron, adding an attention mechanism layer between the input layer and the hidden layer, and obtaining a task processing time prediction model after completion;
[0090] Step S204, dividing the third sample data into a training sample set and a test sample set in a ratio of 3:1;
[0091] Step S205, setting the learning rate of the task processing time prediction model training to c1, the number of training iterations to c2, and the batch size to c3; setting the model training loss function as follows: , where M is the model training loss function, n is the number of feature vectors input to the model, Xi is the actual time required to complete the task, and Yi is the time required to complete the task predicted by the model;
[0092] Step S206, training the task processing time prediction model using the training sample set, and obtaining a first time prediction model after completing c2 training iterations;
[0093] Step S207, using the test sample set to test the first time prediction model, and calculating the prediction fit index of the first time prediction model, the prediction fit index calculation formula is as follows: , where R1 is the prediction fit index, Xj is the actual time required to complete the task, Z is the average time required to complete the task corresponding to all feature vectors in the test sample set, Yj is the time required to complete the task predicted by the first-time prediction model, and m is the number of feature vectors input into the first-time prediction model;
[0094] Step S208, setting the prediction fitting threshold to R0, if R1 is less than R0, then using the training sample set to train the first time prediction model again, and then using the test sample set to test again after completion, and calculating the prediction fitting index R1 until R1 is not less than R0; if R1 is not less than R0, then marking the first time prediction model as the second time prediction model;
[0095] Step S209, when the computer receives a processing task, it obtains the processing task and the task data size, and obtains the task type of the processing task based on the CPU requirement, GPU requirement and memory requirement; then the task data size and task type of the processing task are combined into a feature vector, which is input into the second time prediction model to obtain the predicted time required to complete the task, which is marked as the task duration information.
[0096] Step S3, constructing a task processing progress evaluation method, calculating and evaluating the progress of the computer processing task, and obtaining task progress information; Step S3 includes the following sub-steps:
[0097] Step S301, split the processing task received by the computer into several processing subtasks, and record the subtask data size corresponding to the processing subtasks, record the total number of processing subtasks as v, record any processing subtask as Wk, and record the task data size corresponding to the processing subtask Wk as Sk, where k represents the kth processing subtask;
[0098] Step S302, when the computer starts processing the processing task, it counts the number of subtasks that have been processed in real time, which is recorded as u; and obtains the total size of the subtask data corresponding to the subtasks that have been processed, which is recorded as Z1;
[0099] Step S303, setting the subtask progress weight to T1 and the data progress weight to T2;
[0100] Step S304, the subtask processing progress P1 is calculated based on the total number v of processing subtasks and the number u of processed subtasks that have been processed. The formula for the subtask processing progress is as follows: ;
[0101] Step S305, based on the task data size Z0 of the processing task and the total size Z1 of the subtask data corresponding to the completed processing subtasks, the data processing progress P2 is calculated. The data processing progress calculation formula is as follows: ;
[0102] Step S306, calculating the overall task processing progress P0 based on the subtask processing progress P1 and the data processing progress P2, the overall task processing progress calculation formula is as follows: ; Mark the overall task processing progress P0 as task progress information.
[0103] Step S4, based on the task duration information and the task progress information, dynamically adjust and optimize the computing resources during the computer processing task; Step S4 includes the following sub-steps:
[0104] Step S401, when the computer starts processing the task, it obtains the corresponding task progress information at a first time interval, and records the number of times the corresponding task progress information is obtained, which is recorded as Q, and the first time interval is e;
[0105] Step S402, based on the task duration information corresponding to the processing task, calculate the task processing time deviation BC, and the task processing time deviation calculation formula is as follows: , where A represents the predicted time required to complete the task;
[0106] Step S403, if BC is less than 0, the computing power resources of the corresponding processing task are reduced based on the initial total resource demand of the corresponding processing task; if BC is greater than 0, the computing power resources of the corresponding processing task are increased based on the initial total resource demand of the corresponding processing task; if BC is equal to 0, no adjustment is made; the calculation formula for the reduction ratio and the increase ratio is as follows , where when BL is less than 0, it represents a decrease ratio, and when BL is greater than 0, it represents an increase ratio.
[0107] Example 3, please refer to Figure 5 As shown, Figure 5 The structural diagram of an electronic device is illustrated, and the electronic device may include: a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus. The memory stores computer-readable instructions, and the processor can call the instructions in the memory. When the computer-readable instructions are executed by the processor, the steps in the computing power resource optimization method based on the AI large model are executed to achieve the following functions: collect task sample data of computer processing tasks, and perform data preprocessing on the task sample data to obtain first sample data; construct a task processing time prediction model based on the first sample data, predict the time required for the computer to process the task, and obtain task duration information; construct a task processing progress evaluation method, calculate and evaluate the progress of the computer processing task, and obtain task progress information; based on the task duration information and the task progress information, dynamically adjust and optimize the computing power resources in the process of computer processing tasks.
[0108] In addition, the logic instructions in the above-mentioned memory can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on such an understanding, the technical solution of the present application can be essentially or in other words, the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a disk or an optical disk.
[0109] Embodiment 4, the present application also provides a computer-readable storage medium, the present application provides a storage medium, on which a computer program is stored. When the computer program is executed by the processor, the steps in the computing power resource optimization method based on the AI large model are executed to achieve the following functions: collect task sample data of computer processing tasks, and perform data preprocessing on the task sample data to obtain first sample data; construct a task processing time prediction model based on the first sample data, predict the time required for the computer to process the task, and obtain task duration information; construct a task processing progress evaluation method, calculate and evaluate the progress of the computer processing task, and obtain task progress information; based on the task duration information and the task progress information, dynamically adjust and optimize the computing power resources during the computer processing task.
[0110] Through the description of the above implementation modes, the embodiments of the present invention can be provided as methods, systems or computer program products. Based on such understanding, the above technical solutions can be essentially or the part that contributes to the prior art can be embodied in the form of software products, which can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and include several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0111] In the embodiments provided in the present application, it should be understood that the disclosed system or method can be implemented in other ways. The embodiments described above are merely schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation. For example, multiple modules or units can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, and the indirect coupling or communication connection of systems, modules and units can be electrical, mechanical or other forms.
[0112] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A computing resource optimization method based on an AI big model, characterized in that: The steps include: Collecting task sample data of a computer processing task, and performing data preprocessing on the task sample data to obtain first sample data; Building a task processing time prediction model based on the first sample data, predicting the time required for the computer to process the task, and obtaining task duration information; Construct a task processing progress evaluation method to calculate and evaluate the progress of computer processing tasks and obtain task progress information; Based on task duration information and task progress information, the computing resources are dynamically adjusted and optimized during the computer processing task; After the computer starts processing the task, it obtains the corresponding task progress information at a first time interval, and records the number of times the corresponding task progress information is obtained, which is recorded as Q, and the first time interval is e; Based on the task duration information corresponding to the processing task, the task processing time deviation BC is calculated. The calculation formula for the task processing time deviation is as follows: , where A represents the predicted time required to complete the task, and P0 is the overall task processing progress; If BC is less than 0, the computing power resources of the corresponding processing task will be reduced based on the initial total resource demand of the corresponding processing task. If BC is greater than 0, the computing power resources of the corresponding processing task will be increased based on the initial total resource demand of the corresponding processing task. If BC is equal to 0, no adjustment will be made. The calculation formula for the reduction ratio and the increase ratio is as follows: , where BL is less than 0, which indicates a decrease in ratio, and BL is greater than 0, which indicates an increase in ratio.
2. The computing resource optimization method based on the AI big model according to claim 1 is characterized in that: Collecting task sample data of a computer processing task and performing data preprocessing on the task sample data to obtain first sample data includes the following sub-steps: Obtaining task sample data of each processing task from the computer task execution record, the task sample data including: CPU requirement, GPU requirement, memory requirement, task data size, and time required to complete the task; The total resource demand is calculated according to the CPU demand, GPU demand and memory demand of each processing task. The total resource demand calculation formula is as follows: total resource demand = CPU demand + GPU demand + memory demand. Then, the proportion of CPU demand, GPU demand and memory demand in the total resource demand is calculated respectively, and the proportions are compared to obtain the maximum proportion. The processing tasks are divided into task types according to the demand corresponding to the maximum proportion, and marked. The task types include: CPU demand type, GPU demand type and memory demand type. The CPU demand type is coded as 001, the GPU demand type is coded as 010, and the memory demand type is coded as 100. The task data size, time required to complete the task, and task type of each processing task are stored according to the corresponding processing task classification and marked as first sample data.
3. The computing resource optimization method based on AI big model according to claim 2 is characterized in that: Building a task processing time prediction model based on the first sample data includes the following sub-steps: The CPU demand, GPU demand, memory demand and task data size in the first sample data are normalized according to the data type, and the data size is scaled to [0, 1]. After completion, the second sample data is obtained; The CPU demand, GPU demand, memory demand, task data size and task type under the same processing task in the second sample data are combined into a feature vector, denoted as Ai, where i represents the i-th processing task; after completion, all feature vectors and the corresponding task completion time are stored together and marked as the third sample data; Constructing a task processing time prediction model includes: setting the number of input layer neurons to b1, setting the number of hidden layer neurons to b2, setting the number of output layer neurons to b3 based on a multi-layer perceptron, adding an attention mechanism layer between the input layer and the hidden layer, and obtaining a task processing time prediction model after completion.
4. The computing resource optimization method based on AI big model according to claim 3 is characterized in that: Predicting the time required for a computer to process a task and obtaining task duration information includes the following sub-steps: Divide the third sample data into a training sample set and a test sample set in a ratio of 3:1; Set the learning rate of the task processing time prediction model training to c1, the number of training iterations to c2, and the batch size to c3; set the model training loss function as follows: , where M is the model training loss function, n is the number of feature vectors input to the model, Xi is the actual time required to complete the task, and Yi is the time required to complete the task predicted by the model; The task processing time prediction model is trained using the training sample set. After completing c2 training iterations, the first time prediction model is obtained. The first time prediction model is tested using the test sample set, and the prediction fit index of the first time prediction model is calculated. The prediction fit index calculation formula is as follows: , where R1 is the prediction fit index, Xj is the actual time required to complete the task, Z is the average time required to complete the task corresponding to all feature vectors in the test sample set, Yj is the time required to complete the task predicted by the first-time prediction model, and m is the number of feature vectors input into the first-time prediction model.
5. The computing resource optimization method based on AI big model according to claim 4 is characterized in that: Predicting the time required for a computer to process a task and obtaining task duration information includes the following sub-steps: Set the prediction fitting threshold to R0. If R1 is less than R0, use the training sample set to train the first-time prediction model again. After completion, use the test sample set to test again and calculate the prediction fitting index R1 until R1 is not less than R0. If R1 is not less than R0, mark the first-time prediction model as the second-time prediction model. When the computer receives a processing task, it obtains the processing task and the task data size, and obtains the task type of the processing task based on the CPU demand, GPU demand and memory demand; then the task data size and the task type of the processing task are combined into a feature vector, which is input into the second time prediction model to obtain the predicted time required to complete the task, which is marked as the task duration information.
6. The computing resource optimization method based on AI big model according to claim 5 is characterized in that: Constructing a task processing progress evaluation method to calculate and evaluate the progress of computer processing tasks and obtain task progress information includes the following sub-steps: Split the processing task received by the computer into several processing subtasks, and record the subtask data size corresponding to the processing subtask, record the total number of processing subtasks as v, record any processing subtask as Wk, and record the task data size corresponding to the processing subtask Wk as Sk, where k represents the kth processing subtask; When the computer starts processing the processing task, it counts the number of completed processing subtasks in real time, denoted as u; and obtains the total size of the subtask data corresponding to the completed processing subtasks, denoted as Z1.
7. The computing resource optimization method based on AI big model according to claim 6 is characterized in that: Constructing a task processing progress evaluation method to calculate and evaluate the progress of computer processing tasks and obtain task progress information also includes the following sub-steps: Set the subtask progress weight to T1 and the data progress weight to T2; The subtask processing progress P1 is calculated based on the total number of subtasks v and the number of subtasks u that have been processed. The subtask processing progress formula is as follows: ; The data processing progress P2 is calculated based on the task data size Z0 of the processing task and the total size Z1 of the subtask data corresponding to the completed processing subtask. The data processing progress calculation formula is as follows: ; The overall task processing progress P0 is calculated based on the subtask processing progress P1 and the data processing progress P2. The overall task processing progress calculation formula is as follows: ; Mark the overall task processing progress P0 as task progress information.
8. A computing power resource optimization system based on an AI big model, applicable to the computing power resource optimization method based on an AI big model described in any one of claims 1 to 7, characterized in that: It includes data acquisition module, time prediction module, progress evaluation module and computing power adjustment module; The data acquisition module includes a collection unit and a processing unit, the collection unit is used to collect task sample data of the computer processing task, and the processing unit is used to perform data preprocessing on the task sample data to obtain first sample data; The time prediction module includes a model unit and a prediction unit. The model unit constructs a task processing time prediction model based on the first sample data; the prediction unit is used to predict the time required for the computer to process the task and obtain task time information; The progress evaluation module is used to construct a task processing progress evaluation method, calculate and evaluate the progress of computer processing tasks, and obtain task progress information; The computing power adjustment module dynamically adjusts and optimizes computing power resources during the computer processing task based on task duration information and task progress information.
9. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps in the method according to any one of claims 1 to 7 are executed.
Citation Information
Patent Citations
Computing power resource scheduling method and device, equipment and medium
CN117785487A
Optimization method and optimization device for computing power resource allocation, electronic equipment and medium
CN116541176A
Numerical simulation-oriented computing power scheduling method and system
CN118567840A