Dynamic resource optimization method and system based on large model incremental learning

By adopting a dynamic resource optimization method based on large model incremental learning in the big data processing system, dynamically adjusting server resource allocation, the problem of inefficient resource allocation in the existing technology is solved, and more efficient and reliable resource utilization is achieved.

CN120223553AActive Publication Date: 2025-06-27BEIJING MILLENNIUM VISION TECH CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510700464.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-06-27
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

The prior art fails to effectively adjust resource allocation dynamically based on the difference between the actual processing time and the predicted processing time of each slice, resulting in low resource allocation efficiency.

Method used

The dynamic resource optimization method based on incremental learning of large models is adopted, and the prediction model is updated to predict the prediction processing time of the pending data by obtaining the data amount of each server, the data amount completed and the actual processing time. According to the predicted evaluation value and actual processing time, adjust the server parameters, such as the quantity, network bandwidth or batch selection number of training data, to dynamically adjust resource allocation.

Benefits of technology

By updating the model in real time and adjusting the resource allocation plan dynamically, the efficiency and reliability of resource allocation are improved, resource waste and overload are avoided, and the overall performance of data processing is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120223553A_ABST
    Figure CN120223553A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of big data processing, in particular to a dynamic resource optimization method and system based on large model incremental learning, and the method comprises the steps: updating a model through training data based on a preset training period; judging whether the operation condition of the model is qualified or not, and when the operation condition of the model is judged to be abnormal, adjusting parameters of each server based on the actual processing duration of the to-be-processed data of the single slice, the number of the servers is adjusted to a corresponding value, the network bandwidth of the server corresponding to the to-be-processed data of the single slice is adjusted to a corresponding value, or the batch selection number of the selected training data used for updating the model is adjusted to a corresponding value; and judging that the operation condition of the model is qualified, and continuously using the current model to complete resource allocation. According to the method, the model is continuously updated according to real-time data, a resource allocation scheme is dynamically adjusted, and the resource allocation efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data processing, and particularly to a dynamic resource optimization method and system based on large model incremental learning. Background Art

[0002] In today's digital age, with the rapid development of technologies such as the Internet and the Internet of Things, the amount of data has shown an explosive growth. Enterprises and organizations need to process more and more various types of data (such as business transaction data, sensor data, user behavior data, etc.), which puts higher requirements on the efficiency of data processing and the rationality of resource utilization.

[0003] Traditional resource allocation methods are usually based on static rules or experience. For example, data is allocated to different servers according to a fixed ratio. However, this method cannot adapt to the dynamic changes in the amount of data and data characteristics, easily resulting in overloading of some servers while some server resources are idle, thus reducing the overall data processing efficiency and increasing the operating cost.

[0004] At the same time, with the development of deep learning technology, large models have shown powerful capabilities in data processing and prediction. However, the performance of large models gradually degrades as the data distribution changes. Therefore, how to use large models for dynamic resource optimization, improve resource utilization efficiency and data processing performance, has become a hot issue in current research.

[0005] Chinese Patent Application Publication No.: CN116614396A, discloses a cloud server resource allocation method, including: generating multiple groups of resource pre-allocation information in response to the resources requested by a user; calculating the supply cost information corresponding to the resource pre-allocation information based on the specification parameters, the quantity parameters, and the available zone parameters; using the resource opening record data of the target resources of the user to determine the evaluation information of the user; determining resource allocation information from the multiple groups of resource pre-allocation information according to the supply cost information and the evaluation information, and allocating cloud server resources to the user according to the resource allocation information; It can be seen that the above technical solution has the following problems: It does not consider dynamically adjusting resource allocation according to the difference between the actual processing duration and the predicted processing duration of the data to be processed in each slice, which affects the resource allocation efficiency. Summary of the Invention

[0006] Therefore, the present invention provides a dynamic resource optimization method and system based on large model incremental learning to overcome the problem in the prior art that the resource allocation is not dynamically adjusted according to the difference between the actual processing duration and the predicted processing duration of the data to be processed in each slice, which affects the resource allocation efficiency.

[0007] On the one hand, the present invention provides a dynamic resource optimization method based on large model incremental learning, including: Dividing the obtained data volume received by each server, the processed data volume, and the corresponding actual processing duration into several batches of training data according to the data volume; Updating the model for predicting the predicted processing duration of the data to be processed for each slice using the training data based on a preset training period; Obtaining the predicted processing duration of the data to be processed for each slice predicted by the model; Obtaining the actual processing duration of each server for actually processing the data to be processed for each slice; Determining the predicted evaluation value of the data to be processed for each slice based on the predicted processing duration and the actual processing duration; Judging whether the running status of the model is qualified one by one according to each predicted evaluation value, including: When it is determined that the running status of the model is abnormal, adjusting the parameters of each server based on the actual processing duration of the data to be processed for a single slice, including adjusting the number of servers to the corresponding value, adjusting the network bandwidth of the server corresponding to the data to be processed for a single slice to the corresponding value, or adjusting the number of batches of the training data selected for updating the model to the corresponding value; When it is determined that the running status of the model is qualified according to each predicted evaluation value one by one, continuously using the current model to complete the allocation of resources.

[0008] Further, for the data to be processed for a single slice, the process of determining the predicted evaluation value based on the predicted processing duration and the actual processing duration and judging whether the running status of the model is qualified based on the predicted evaluation value includes: Denoting the ratio of the calculated actual processing duration to the predicted processing duration as the predicted evaluation value; If the predicted evaluation value is less than or equal to the first preset evaluation value, it is determined that the running status of the model is qualified; If the predicted evaluation value is less than or equal to the second preset evaluation value and greater than the first preset evaluation value, then judge whether the running status of the model is qualified in combination with the historical call frequency of the data to be processed for a single slice; If the predicted evaluation value is greater than the second preset evaluation value, it is determined that the running status of the model is abnormal, and the parameters of the corresponding server are adjusted based on the actual processing duration of the data to be processed for a single slice.

[0009] Further, judging whether the running status of the model is qualified in combination with the historical call frequency of the data to be processed for a single slice includes: If the historical call frequency is less than or equal to the preset historical call frequency, it is determined that the running status of the model is qualified; If the historical call frequency is greater than the preset historical call frequency, it is determined that the operation status of the model is abnormal, and the parameters of the corresponding server are adjusted based on the actual processing duration of the data to be processed for a single slice.

[0010] Further, the process of adjusting the parameters of the corresponding server based on the actual processing duration of the data to be processed for a single slice includes: If the actual processing duration is less than or equal to the first preset actual processing duration, the prediction fluctuation parameter is determined based on the predicted evaluation values of the data to be processed for each slice, and the parameters of the corresponding server are adjusted based on the prediction fluctuation parameter; If the actual processing duration is less than or equal to the second preset actual processing duration and greater than the first preset actual processing duration, the number of servers is adjusted to the corresponding value based on the average evaluation value of the predicted evaluation values of the data to be processed for each slice; If the actual processing duration is greater than the second preset actual processing duration, the data duplication rate is determined based on the data content of the data to be processed for a single slice, and the parameters of the corresponding server are adjusted based on the data duplication rate.

[0011] Further, the process of adjusting the parameters of the corresponding server based on the prediction fluctuation parameter includes: Denote the variance of the predicted evaluation values of the data to be processed for each slice calculated as the prediction fluctuation parameter; If the prediction fluctuation parameter is less than or equal to the preset prediction fluctuation parameter, the network bandwidth of the server corresponding to the data to be processed for a single slice is adjusted to the corresponding value based on the actual processing duration; If the prediction fluctuation parameter is greater than the preset prediction fluctuation parameter, the number of batches of training data selected for updating the model is adjusted to the corresponding value based on the prediction fluctuation parameter.

[0012] Further, when adjusting the number of servers to the corresponding value based on the average evaluation value, where The increase amplitude of the number of servers is proportional to the average evaluation value.

[0013] Further, the process of determining the data duplication rate based on the data content of the data to be processed for a single slice and adjusting the parameters of each server based on the data duplication rate includes: Determine the data sender corresponding to the data to be processed for a single slice; Obtain the data sent by the data sender each time in history; Calculate the duplication rate between the data sent each time respectively; Solve the average value of each duplication rate to obtain the data duplication rate; If the data duplication rate is less than or equal to the preset data duplication rate, the number of servers is adjusted to the corresponding value based on the average evaluation value; If the data duplication rate is greater than the preset data duplication rate, mark a single data sender as an orderly sender.

[0014] Further, adjust the network bandwidth of the server corresponding to the data to be processed for a single slice to a corresponding value based on the actual processing duration, where The increase amplitude of the network bandwidth is proportional to the actual processing duration.

[0015] Further, adjust the number of batches of training data selected for model update based on the predicted fluctuation parameter, where The increase amplitude of the number of batches selected is proportional to the predicted fluctuation parameter.

[0016] On the other hand, the present invention also provides a dynamic resource optimization system for a dynamic resource optimization method based on large model incremental learning, including: A data output module, which includes several data senders for sending data to be processed; A model processing module, which is connected to the data output module, and is used to divide the data sent by each data sender into data to be processed in several slices, and determine the optimal resource allocation plan for each server to process the data to be processed and the processing duration of the server for processing each data to be processed; A data receiving module, which is respectively connected to the data output module and the model processing module, and includes several servers for receiving and processing each slice of data to be processed sent by each data sender according to the optimal resource allocation plan; A data statistics module, which is connected to the data receiving module, and is used to count the actual processing duration of each server for processing each slice of data to be processed; A data division module, which is respectively connected to the model processing module, the data receiving module and the data statistics module, and is used to divide the received data volume, the processed data volume and the corresponding actual processing duration of the server according to the data volume into several batches of training data; A model training module, which is respectively connected to the data division module and the model processing module, and is used to select several batches of training data to update the model; The data analysis module is respectively connected to the model processing module, the data statistics module, the data receiving module and the model training module, and is used to determine a prediction evaluation value based on the predicted processing duration and the actual processing duration, determine whether the operation status of the model is qualified based on the prediction evaluation value, and in the case of determining that the operation status of the model is abnormal, adjust the parameters of the corresponding server based on the actual processing duration of the data to be processed in a single slice, including adjusting the number of servers to a corresponding value, adjusting the network bandwidth of the server corresponding to the data to be processed in a single slice to a corresponding value, or adjusting the number of batches of training data selected for updating the model to a corresponding value.

[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: the model is updated using training data based on a preset training period; it is determined whether the operation status of the model is qualified, and when it is determined that the operation status of the model is abnormal, the parameters of each server are adjusted based on the actual processing duration of the data to be processed in a single slice, including adjusting the number of servers to a corresponding value, adjusting the network bandwidth of the server corresponding to the data to be processed in a single slice to a corresponding value, or adjusting the number of batches of training data selected for updating the model to a corresponding value; when it is determined that the operation status of the model is qualified, the current model is continuously used to complete the allocation of resources. The model is continuously updated according to real-time data, and the resource allocation scheme is dynamically adjusted, improving the resource allocation efficiency.

[0018] Further, when obtaining the predicted processing duration of each slice of data to be processed predicted by the model, the slice of data to be processed is a subset obtained by dividing the data to be processed, and the model predicts the processing duration of each slice of data on each server according to the input data to be processed. The ratio of the calculated actual processing duration to the predicted processing duration is recorded as the prediction evaluation value. The prediction evaluation value reflects the accuracy of the model prediction. The closer the ratio is to 1, the more accurate the model prediction is. When the prediction evaluation value is less than or equal to the first preset evaluation value, the accuracy of the model prediction is high, it is determined that the operation status of the model is qualified, and the current model is continuously used to complete the allocation of resources. When the prediction evaluation value is less than or equal to the second preset evaluation value and greater than the first preset evaluation value, the prediction accuracy of the model has decreased, but it cannot be directly determined that the model operation is abnormal, and it is necessary to judge in combination with the historical call frequency. When the prediction evaluation value is greater than the second preset evaluation value, the prediction error of the model is large, it is determined that the operation status of the model is abnormal, and the parameters of each server are adjusted based on the actual processing duration of the data to be processed in a single share. Dynamically adjusting the parameters and the number of servers according to the prediction evaluation value can perform adaptive adjustment according to the actual situation, improving the reliability of resource allocation and the resource allocation efficiency.

[0019] Furthermore, based on the historical call frequency evaluation model, the historical call frequency characterizes the usage frequency of data. For data with a higher usage frequency, the prediction accuracy requirement of the model is higher, and it is necessary to obtain the processed data more timely. When the historical call frequency is less than or equal to the preset historical call frequency, the usage frequency of this data is relatively low. Even if there are certain errors in the model prediction, the impact on data calls is small. It is determined that the running status of the model is qualified, and the current model is continuously used to complete the allocation of resources. When the historical call frequency is greater than the preset historical call frequency, the usage frequency of this data is relatively high, and the prediction error of the model will have a greater impact on data calls. In this case, it is determined that the running status of the model is abnormal, and the parameters of each server are adjusted based on the actual processing duration of a single share of the data to be processed. Determine the analysis criteria of the model according to the actual situation of the data, avoid excessive adjustment of the running parameters of the server, resulting in waste of resources, and improve the efficiency of resource allocation.

[0020] Furthermore, based on the actual processing duration adjustment strategy, the actual processing duration characterizes the deviation amplitude of the actual predicted processing duration in the case of a large predicted evaluation value. When the actual processing duration is less than or equal to the first preset actual processing duration, the data processing speed is relatively fast. In this case, the prediction fluctuation parameter is determined based on the predicted evaluation value of the data to be processed for each slice. The prediction fluctuation parameter characterizes the error difference situation of the model for the data to be processed for each slice. When the prediction fluctuation parameter is less than or equal to the preset prediction fluctuation parameter, the predictions of the model for the data to be processed for each slice have similar degrees of difference. In this case, the prediction differences are caused by abnormal network environment. In this case, the network bandwidth is adjusted to ensure stable data transmission and processing. When the prediction fluctuation parameter is greater than the preset prediction fluctuation parameter, in this case, the resource allocation is abnormal due to the abnormality of the model, resulting in different degrees of errors in the predictions of the data for each slice. In this case, the number of batches of training data used to update the model is increased to improve the update effect of the model and ensure the stable operation of the model. When the actual processing duration is less than or equal to the second preset actual processing duration and greater than the first preset actual processing duration, in this case, due to server abnormality, the sliced data requires a relatively high processing duration to complete data transmission and processing. In this case, the number of servers is adjusted to improve data processing performance. When the actual processing duration is greater than the second preset actual processing duration, in this case, the data processing speed is relatively slow. The data duplication rate is determined based on the data content of a single share of the data to be processed, and the parameters of each server are adjusted based on the data duplication rate. By dynamically adjusting the parameters of the server according to indicators such as the predicted evaluation value, actual processing duration, historical call frequency, and data duplication rate of the model, it can be adaptively adjusted according to the actual situation, improve the reliability of resource allocation, and thus improve the efficiency of resource allocation.

[0021] Further, when the data duplication rate is less than or equal to the preset data duplication rate, the data redundancy level is low, and the number of servers is increased to ensure data processing efficiency. When the data duplication rate is greater than the preset data duplication rate, the data redundancy level is high, and a single data sender is marked as an orderly sender for subsequent data processing. By predicting the resource allocation plan and processing duration of the data to be processed through a model, the data to be processed is allocated to each server, avoiding waste and overload of resources and improving the overall data processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 FIG. is a flowchart of the steps of the dynamic resource optimization method based on large model incremental learning according to an embodiment of the present invention; Figure 2 FIG. is a block diagram of the modules of the dynamic resource optimization system based on large model incremental learning according to an embodiment of the present invention; Figure 3 FIG. is a logical decision diagram for determining whether the operation status of the prediction evaluation value determination model according to an embodiment of the present invention is qualified; Figure 4 FIG. is a logical decision diagram for determining whether the operation status of the historical call frequency determination model according to an embodiment of the present invention is qualified. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] In order to make the objectives and advantages of the present invention clearer, the present invention will be further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0024] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present invention and do not limit the protection scope of the present invention.

[0025] It should be noted that in the description of the present invention, the terms indicating directions or positional relationships such as "upper", "lower", "left", "right", "inner", "outer", etc. are based on the directions or positional relationships shown in the drawings. This is only for convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.

[0026] In addition, it should be noted that in the description of the present invention, unless otherwise clearly specified and defined, the terms "installation", "connection", and "coupling" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, and it can be the communication inside two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0027] Please refer to Figure 1 、 Figure 2 、 Figure 3 and Figure 4 as shown, which are respectively the step flow chart of the dynamic resource optimization method based on large model incremental learning in the embodiment of the present invention, the module block diagram of the dynamic resource optimization system based on large model incremental learning, the logical decision diagram for determining whether the operation status of the prediction evaluation value determination model is qualified, and the logical decision diagram for determining whether the operation status of the historical call frequency determination model is qualified; An embodiment of the present invention provides a dynamic resource optimization method and system based on large model incremental learning, including: S1. Divide the data volume, processed data volume, and corresponding actual processing duration received by each server into several batches of training data according to the data volume. S2. Update the model for predicting the predicted processing duration of the data to be processed for each slice using the training data based on a preset training period. S3. Obtain the predicted processing duration of the data to be processed for each slice predicted by the model. S4. Obtain the actual processing duration of each server for actually processing the data to be processed for each slice. S5. Determine the predicted evaluation value of the data to be processed for each slice based on the predicted processing duration and the actual processing duration. S6. Determine whether the operation status of the model is qualified one by one according to each predicted evaluation value, including: When it is determined that the operation status of the model is abnormal, adjust the parameters of each server based on the actual processing duration of the data to be processed for a single slice, including adjusting the number of servers to the corresponding value, adjusting the network bandwidth of the server corresponding to the data to be processed for a single slice to the corresponding value, or adjusting the number of batches of the training data selected for updating the model to the corresponding value; S7. When it is determined that the operation status of the model is qualified according to each predicted evaluation value one by one, continuously use the current model to complete the allocation of resources.

[0028] Specifically, the process of dividing the data volume, processed data volume, and corresponding actual processing duration received by each server into several batches of training data according to the data volume includes: Package the data volume received by a single server, the processed data volume, and the corresponding actual processing duration into a single piece of data; collect each piece of data from each server and divide each piece of data into several batches of training data.

[0029] Specifically, update the model using the training data based on a preset training cycle; determine whether the running status of the model is qualified. When it is determined that the running status of the model is abnormal, adjust the parameters of each server based on the actual processing duration of the data to be processed in a single slice, including adjusting the number of servers to the corresponding value, adjusting the network bandwidth of the server corresponding to the data to be processed in a single slice to the corresponding value, or adjusting the number of selected batches of training data used to update the model to the corresponding value; when it is determined that the running status of the model is qualified, continue to use the current model to complete the resource allocation. Continuously updating the model according to real-time data and dynamically adjusting the resource allocation scheme improves the resource allocation efficiency.

[0030] Specifically, there is no limitation on the specific method of obtaining training data. The received data volume, the processed data volume, and the corresponding actual processing duration can be obtained from the log systems of each server; the training data can be collected through the built-in monitoring tools of the server, such as the Linux system sar, iostat, etc., or professional monitoring software such as Zabbix, Prometheus, etc.

[0031] Specifically, there is no limitation on the data volume division. In this embodiment, optionally, the training data can be divided into batches according to the size of the data volume.

[0032] Specifically, divide the training data into batches to load some batches of data into the memory for model training, so as to avoid training a large amount of data at one time, resulting in insufficient memory and causing a memory overflow error. By dividing the data into smaller batches and only loading some batches of data into the memory for processing each time, the memory usage can be effectively controlled to ensure the stability of the training process.

[0033] Specifically, train the model with each batch of training data to make the model more accurately predict the resource allocation scheme of the data to be processed.

[0034] Specifically, for the model, recurrent neural networks and their variants LSTM, GRU in deep learning, or models based on the Transformer architecture, such as BERT, can be selected as the basic model to achieve the functions of allocating and processing sequential data and time series prediction.

[0035] Specifically, for model training, the input data includes the amount of data received by the server, the amount of processed data, and the actual processing duration in each batch of training data. The Min - Max normalization method is used to normalize this data and map it to the interval [0, 1]. Stochastic Gradient Descent (SGD) and its optimization algorithms Adam and Adagrad are used for model training. In each preset training cycle, new training data is used for incremental learning of the model to update the model's parameters.

[0036] Specifically, for model acquisition, an open - source deep - learning framework such as TensorFlow or PyTorch can be used to build and train the model. After training, the model is saved as a file for subsequent use.

[0037] Specifically, the model is used to determine the optimal resource allocation plan for each server to process the data to be processed and the processing duration of the server for each piece of data to be processed.

[0038] Specifically, the resource optimization plan includes dividing the data to be processed into several slices, where the amount of data in each slice is different. It is determined to allocate the data to be processed in the corresponding slice to the corresponding server for processing, and the trained model predicts the processing duration of each data slice to obtain the predicted processing duration.

[0039] Specifically, the dynamic resources include several servers used to process the data to be processed.

[0040] Specifically, when each server processes each data slice, the actual processing duration is recorded, which can be achieved by adding time - recording code in the server processing program.

[0041] Specifically, when the adjustment of the parameters for each server is completed, the running status of the model is judged to be qualified one by one according to the re - obtained predicted evaluation values.

[0042] Specifically, for the data to be processed in a single slice, the process of determining the predicted evaluation value based on the predicted processing duration and the actual processing duration and judging whether the running status of the model is qualified based on the predicted evaluation value includes: Recording the ratio of the calculated actual processing duration to the predicted processing duration as the predicted evaluation value; If the predicted evaluation value is less than or equal to the first preset evaluation value, it is determined that the running status of the model is qualified; If the predicted evaluation value is less than or equal to the second preset evaluation value and greater than the first preset evaluation value, it is judged whether the running status of the model is qualified in combination with the historical call frequency of the data to be processed in a single slice; If the predicted evaluation value is greater than the second preset evaluation value, it is determined that the operation status of the model is abnormal, and the parameters of the corresponding server are adjusted based on the actual processing duration of the data to be processed for a single slice.

[0043] Specifically, the first preset evaluation value is selected within the range [0.9, 1.1], and the second preset evaluation value is selected within the range [1.2, 1.3].

[0044] Specifically, when obtaining the predicted processing time of the data to be processed for each slice predicted by the model, the data to be processed for a slice is a subset obtained by dividing the data to be processed. The model predicts the processing duration of each slice of data on each server according to the input data to be processed. The ratio of the calculated actual processing duration to the predicted processing duration is recorded as the predicted evaluation value. The predicted evaluation value reflects the accuracy of the model prediction. The closer the ratio is to 1, the more accurate the model prediction is. When the predicted evaluation value is less than or equal to the first preset evaluation value, the accuracy of the model prediction is high, and it is determined that the operation status of the model is qualified, and the current model is continuously used to complete the allocation of resources. When the predicted evaluation value is less than or equal to the second preset evaluation value and greater than the first preset evaluation value, the prediction accuracy of the model decreases, but it cannot be directly determined that the model operation is abnormal, and it is necessary to combine the historical call frequency for judgment. When the predicted evaluation value is greater than the second preset evaluation value, the prediction error of the model is large, and it is determined that the operation status of the model is abnormal, and the parameters of each server are adjusted based on the actual processing duration of the data to be processed for a single share. Dynamically adjusting the parameters and quantity of the server according to the predicted evaluation value can perform adaptive adjustment according to the actual situation, improve the reliability of resource allocation, and improve the resource allocation efficiency.

[0045] Specifically, determining whether the operation status of the model is qualified based on the historical call frequency of the data to be processed for a single slice includes: If the historical call frequency is less than or equal to the preset historical call frequency, it is determined that the operation status of the model is qualified; If the historical call frequency is greater than the preset historical call frequency, it is determined that the operation status of the model is abnormal, and the parameters of the corresponding server are adjusted based on the actual processing duration of the data to be processed for a single slice.

[0046] Specifically, the preset historical call frequency is selected within the range [0.1, 0.3], and the unit is times / day.

[0047] Specifically, obtain the data sender corresponding to the data to be processed for a single slice, and obtain the ratio of the number of times each data sent by the server from the data sender within the preset analysis duration to the preset analysis duration, to obtain the historical call frequency.

[0048] Specifically, based on the historical call frequency evaluation model, the historical call frequency characterizes the usage frequency of data. For data with a higher usage frequency, the model requires higher prediction accuracy and needs to obtain the processed data more timely. When the historical call frequency is less than or equal to the preset historical call frequency, the usage frequency of this data is relatively low. Even if there are certain errors in the model prediction, the impact on data calls is small, and it is determined that the operating condition of the model is qualified, and the current model continues to be used to complete the allocation of resources. When the historical call frequency is greater than the preset historical call frequency, the usage frequency of this data is relatively high, and the prediction error of the model will have a greater impact on data calls. In this case, it is determined that the operating condition of the model is abnormal, and the parameters of each server are adjusted based on the actual processing duration of a single share of the data to be processed. The analysis criteria of the model are determined according to the actual situation of the data, avoiding excessive adjustment of the operating parameters of the server, resulting in waste of resources, and improving the efficiency of resource allocation.

[0049] Specifically, the process of adjusting the parameters of the corresponding server based on the actual processing duration of the data to be processed for a single slice includes: If the actual processing duration is less than or equal to the first preset actual processing duration, the prediction fluctuation parameter is determined based on the prediction evaluation values of the data to be processed for each slice, and the parameters of the corresponding server are adjusted based on the prediction fluctuation parameter; If the actual processing duration is less than or equal to the second preset actual processing duration and greater than the first preset actual processing duration, the number of servers is adjusted to the corresponding value based on the average evaluation value of the prediction evaluation values of the data to be processed for each slice; If the actual processing duration is greater than the second preset actual processing duration, the data duplication rate is determined based on the data content of the data to be processed for a single slice, and the parameters of the corresponding server are adjusted based on the data duplication rate.

[0050] Specifically, the first preset actual processing duration T1 is selected within the range of [10, 20], and the second preset actual processing duration T2 is selected within the range of [40, 55], with the unit being seconds.

[0051] Specifically, the process of adjusting the parameters of the corresponding server based on the prediction fluctuation parameter includes: Denote the variance of the prediction evaluation values of the data to be processed for each slice calculated as the prediction fluctuation parameter; If the prediction fluctuation parameter is less than or equal to the preset prediction fluctuation parameter, the network bandwidth of the server corresponding to the data to be processed for a single slice is adjusted to the corresponding value based on the actual processing duration; Specifically, the single slice mentioned in adjusting the network bandwidth of the server corresponding to the data to be processed for a single slice to the corresponding value based on the actual processing duration is the single slice used to determine the prediction evaluation value in determining whether the operating condition of the model is qualified based on the prediction evaluation value.

[0052] If the predicted fluctuation parameter is greater than the preset predicted fluctuation parameter, the batch selection quantity of the training data selected for updating the model is adjusted to the corresponding value based on the predicted fluctuation parameter.

[0053] Specifically, the preset predicted fluctuation parameter Y0 is selected within the range [0.001, 0.005].

[0054] Specifically, based on the actual processing duration adjustment strategy, the actual processing duration characterizes the deviation amplitude of the actual predicted processing duration in the case of a relatively large predicted evaluation value. When the actual processing duration is less than or equal to the first preset actual processing duration, the data processing speed is relatively fast. In this case, the predicted fluctuation parameter is determined based on the predicted evaluation values of the to-be-processed data of each slice. The predicted fluctuation parameter characterizes the error difference situation of the model for the to-be-processed data of each slice. When the predicted fluctuation parameter is less than or equal to the preset predicted fluctuation parameter, the predictions of the model for the to-be-processed data of each slice have similar degrees of differences. In this case, the prediction differences are caused by an abnormal network environment. In this case, the network bandwidth of the server corresponding to the to-be-processed data of a single slice is adjusted to ensure stable data transmission and processing. When the predicted fluctuation parameter is greater than the preset predicted fluctuation parameter, in this case, the resource allocation is abnormal due to an abnormality of the model, resulting in different degrees of errors in the predictions of the data of each slice. In this case, the batch selection quantity of the training data selected for updating the model is increased to improve the updating effect of the model and ensure the stable operation of the model. When the actual processing duration is less than or equal to the second preset actual processing duration and greater than the first preset actual processing duration, in this case, the slice data requires a relatively high processing duration to complete data transmission and processing due to server abnormalities. In this case, the number of servers is adjusted to improve the data processing performance. When the actual processing duration is greater than the second preset actual processing duration, in this case, the data processing speed is relatively slow. The data repetition rate is determined based on the data content of the to-be-processed data of a single share, and the parameters of each server are adjusted based on the data repetition rate. By dynamically adjusting the parameters of the server according to indicators such as the predicted evaluation value, actual processing duration, historical call frequency, and data repetition rate of the model, it can be adaptively adjusted according to the actual situation, improving the reliability of resource allocation and further improving the efficiency of resource allocation.

[0055] Specifically, the number of servers is adjusted to the corresponding value based on the average evaluation value, where The increase amplitude of the number of servers is proportional to the average evaluation value.

[0056] In this embodiment, optionally, The average evaluation value is compared with the first preset average evaluation value and the second preset average evaluation value; If the average evaluation value is less than or equal to the first preset average evaluation value, adjust the number of servers to 1.1 times the initial number of servers; If the average evaluation value is less than or equal to the second preset average evaluation value and greater than the first preset average evaluation value, adjust the number of servers to 1.2 times the initial number of servers; If the average evaluation value is greater than the second preset average evaluation value, adjust the number of servers to 1.25 times the initial number of servers; The first preset average evaluation value is taken as 1.05, and the second preset average evaluation value is taken as 1.15.

[0057] Specifically, the process of determining the data duplication rate based on the data content of the data to be processed for a single slice and adjusting the parameters of each server based on the data duplication rate includes: Determine the data sender corresponding to the data to be processed for a single slice; Obtain the data sent by the data sender in each previous time; Calculate the duplication rate between the data sent each time respectively; Solve the average value of each duplication rate to obtain the data duplication rate; If the data duplication rate is less than or equal to the preset data duplication rate, adjust the number of servers to the corresponding value based on the average evaluation value; If the data duplication rate is greater than the preset data duplication rate, mark the single data sender as an orderly sender.

[0058] Specifically, the preset data duplication rate is selected within the interval [0.4, 0.6].

[0059] Specifically, for the data sent twice, record the ratio of the calculated duplicate data volume to the total data volume of the data sent twice as the duplication rate between the data sent twice.

[0060] Specifically, when a single server receives the data to be processed sent by the orderly sender, identify the difference area between the data to be processed and the sending template of the orderly sender, and only analyze and process the difference area.

[0061] Specifically, the specific method of analyzing and processing the difference area is not limited, and may include data standardization and inputting the standardized data into a pre-stored program for analysis, which will not be elaborated here.

[0062] Specifically, the difference area is the area where the identified data to be processed is different from the sending template.

[0063] Specifically, the determination method of the sending template is Calculate the sum of the repetition rates of the data sent in the historical single transmission and the data sent in each transmission to obtain the reference repetition rate for the data in the historical single transmission; Determine the data of the single transmission corresponding to the maximum value of the reference repetition rates of the data sent in each historical transmission as the transmission template.

[0064] Specifically, when the data repetition rate is less than or equal to the preset data repetition rate, the degree of data redundancy is low, and the number of servers is increased to ensure data processing efficiency. When the data repetition rate is greater than the preset data repetition rate, the degree of data redundancy is high, and a single data sender is marked as an orderly sender for subsequent data processing. Predict the resource allocation plan and processing duration of the data to be processed through a model, and allocate the data to be processed to each server to avoid waste and overload of resources and improve the overall data processing efficiency.

[0065] Specifically, based on the actual processing duration, adjust the network bandwidth of the server corresponding to the data to be processed for a single slice to the corresponding value, where The increase in network bandwidth is proportional to the actual processing duration.

[0066] In this embodiment, optionally, Compare the actual processing duration with the first duration comparison threshold and the second duration comparison threshold; If the actual processing duration is less than or equal to the first duration comparison threshold, adjust the network bandwidth of the server corresponding to the data to be processed for a single slice to 1.12 times the initial network bandwidth; If the actual processing duration is less than or equal to the second duration comparison threshold and greater than the first duration comparison threshold, adjust the network bandwidth of the server corresponding to the data to be processed for a single slice to 1.21 times the initial network bandwidth; If the actual processing duration is greater than the second duration comparison threshold, adjust the network bandwidth of the server corresponding to the data to be processed for a single slice to 1.28 times the initial network bandwidth; The first duration comparison threshold is taken as 0.5T1, and the second duration comparison threshold is taken as 0.7T1.

[0067] Specifically, based on the predicted fluctuation parameter, adjust the batch selection quantity of the training data for the updated model to be selected to the corresponding value, where The increase in the batch selection quantity is proportional to the predicted fluctuation parameter.

[0068] In this embodiment, optionally, Compare the predicted fluctuation parameter with the first fluctuation comparison threshold and the second fluctuation comparison threshold; If the predicted fluctuation parameter is less than or equal to the first fluctuation ratio threshold, the batch selection quantity of the training data selected for updating the model is adjusted to 1.15 times the initial batch selection quantity; If the predicted fluctuation parameter is less than or equal to the second fluctuation ratio threshold and greater than the first fluctuation ratio threshold, the batch selection quantity of the training data selected for updating the model is adjusted to 1.23 times the initial batch selection quantity; If the predicted fluctuation parameter is greater than the second fluctuation ratio threshold, the batch selection quantity of the training data selected for updating the model is adjusted to 1.3 times the initial batch selection quantity; The first fluctuation ratio threshold is 4Y0, and the second fluctuation ratio threshold is 6Y0.

[0069] Specifically, the data output module includes a number of data sending ends for sending data to be processed; The model processing module is connected to the data output module, and is used to divide the data sent by each data sending end into several slices of data to be processed, and determine the optimal resource allocation scheme for each server to process the data to be processed and the processing duration of the server for processing each data to be processed; The data receiving module is respectively connected to the data output module and the model processing module, and includes a number of servers for receiving and processing each slice of data to be processed sent by each data sending end according to the optimal resource allocation scheme; The data statistics module is connected to the data receiving module, and is used to count the actual processing duration of each server for processing each slice of data to be processed; The data division module is respectively connected to the model processing module, the data receiving module and the data statistics module, and is used to divide the counted data volume, processed data volume and corresponding actual processing duration received by the server into several batches of training data according to the data volume; The model training module is respectively connected to the data division module and the model processing module, and is used to select several batches of training data to update the model; The data analysis module is respectively connected to the model processing module, the data statistics module, the data receiving module and the model training module, and is used to determine the prediction evaluation value based on the predicted processing duration and the actual processing duration, determine whether the running status of the model is qualified based on the prediction evaluation value, and in the case of determining that the running status of the model is abnormal, adjust the parameters of the corresponding server based on the actual processing duration of the data to be processed for a single slice, including adjusting the number of servers to the corresponding value, adjusting the network bandwidth of the server corresponding to the data to be processed for a single slice to the corresponding value, or adjusting the batch selection quantity of the training data selected for updating the model to the corresponding value.

[0070] Specifically, the specific structure of the data sending end in the data output module is not limited and can be various data sources, such as sensors, business systems, user terminals, etc. These data sending ends can send the data to be processed to the model processing module and the data receiving module through network interfaces such as HTTP and TCP. To ensure the orderly sending of data and the stability of transmission, a buffer can be set at the data sending end to cache and queue the data to be processed.

[0071] Specifically, the specific method of dividing the data to be processed into dry slices is not limited. Slicing can be performed according to the size, type or business requirements of the data. For large datasets, they can be divided according to a fixed data volume; for time series data, they can be divided according to time windows.

[0072] Specifically, for resource allocation and processing duration prediction, a trained model can be used to determine the optimal resource allocation plan for each server to process the data to be processed and the processing duration of each server for processing the data to be processed. This model can be a deep learning-based model, such as LSTM and GRU. It can be understood that by learning historical data, it can predict the processing duration of different data slices on different servers, so as to achieve the optimal allocation of resources.

[0073] Specifically, the specific method of each server processing the data sent by each data sending end is not limited and can include data classification, data preprocessing, analyzing the data based on a preset program and outputting the analysis results, and data storage.

[0074] Specifically, the specific structure of the data statistics module is not limited and can be any logical component. It can be understood that it can only achieve the actual processing duration of each server processing each slice of data to be processed. This can be achieved by adding time recording code to the server processing program to record the time when the data starts to be received and the time when the processing is completed, and the difference between the two is the actual processing duration.

[0075] Specifically, the specific manner in which the server processes the data sent by each data sender can be as follows: data classification, that is, dividing the data in the data set according to certain rules or characteristics to form different categories. In the data for telecommunications equipment fault handling, the data can be classified according to characteristics such as equipment type, fault level, and professional type. The collected operation time, process status, changes in operation roles, warning feedback information, fault level information, fault category information, site information, equipment information, data update information, etc. can be used as the basis for data classification. Data preprocessing is to unify the data format to ensure the integrity and accuracy of the data. Based on a preset program, the data is analyzed and the analysis result is output for early warning analysis of the data. According to the set threshold, the analysis is carried out from three aspects: time, quantity, and management, the early warning level is determined, and an early warning report is output. Different servers process the data of different senders.

[0076] Specifically, the specific manner in which the model training module selects several batches of training data to update the model can be as follows: obtain the divided batches of training data from the data division module, and load the current model parameters. This model is a model file that has been trained and saved previously. The model is trained using Stochastic Gradient Descent (SGD) and its optimization algorithms Adam and Adagrad. During the training process, the training data is input into the model, the mean square error between the output of the model and the actual processing duration is calculated, and the model parameters are updated according to the error backpropagation. After the training is completed, the updated model parameters are saved to a file for subsequent use.

[0077] So far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present invention.

[0078] The above are only the preferred embodiments of the present invention and are not used to limit the present invention; for those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A dynamic resource optimization method based on large model incremental learning, characterized in that Including: Dividing the data volume received by each server, the processed data volume, and the corresponding actual processing duration into several batches of training data according to the data volume; Updating the model for predicting the predicted processing duration of the data to be processed for each slice based on the preset training period using the training data; Obtaining the predicted processing duration of the data to be processed for each slice predicted by the model; Obtaining the actual processing duration of each server for actually processing the data to be processed for each slice; Determining the predicted evaluation value of the data to be processed for each slice based on the predicted processing duration and the actual processing duration; Judging whether the running status of the model is qualified one by one according to each predicted evaluation value, including: When it is determined that the running status of the model is abnormal, adjusting the parameters of each server based on the actual processing duration of the data to be processed for a single slice, including adjusting the number of servers to the corresponding value, adjusting the network bandwidth of the server corresponding to the data to be processed for a single slice to the corresponding value, or adjusting the selected number of batches of training data used to update the model to the corresponding value; When it is determined that the running status of the model is qualified according to each predicted evaluation value one by one, continuously using the current model to complete the allocation of resources.

2. The dynamic resource optimization method based on large model incremental learning according to claim 1, characterized in that For the data to be processed for a single slice, the process of determining the predicted evaluation value based on the predicted processing duration and the actual processing duration, and judging whether the running status of the model is qualified based on the predicted evaluation value, including: Recording the ratio of the calculated actual processing duration to the predicted processing duration as the predicted evaluation value; If the predicted evaluation value is less than or equal to the first preset evaluation value, it is determined that the running status of the model is qualified; If the predicted evaluation value is less than or equal to the second preset evaluation value and greater than the first preset evaluation value, judging whether the running status of the model is qualified in combination with the historical call frequency of the data to be processed for a single slice; If the predicted evaluation value is greater than the second preset evaluation value, it is determined that the running status of the model is abnormal, and the parameters of the corresponding server are adjusted based on the actual processing duration of the data to be processed for a single slice.

3. The dynamic resource optimization method based on large model incremental learning according to claim 2, wherein Judging whether the running status of the model is qualified in combination with the historical call frequency of the data to be processed for a single slice, including: If the historical call frequency is less than or equal to the preset historical call frequency, it is determined that the running status of the model is qualified; If the historical call frequency is greater than the preset historical call frequency, it is determined that the running status of the model is abnormal, and the parameters of the corresponding server are adjusted based on the actual processing duration of the data to be processed for a single slice.

4. The dynamic resource optimization method based on large model incremental learning according to claim 3, characterized in that The process of adjusting the parameters of the corresponding server based on the actual processing duration of the data to be processed for a single slice, including: If the actual processing duration is less than or equal to the first preset actual processing duration, determining the predicted fluctuation parameter based on the predicted evaluation value of the data to be processed for each slice, and adjusting the parameters of the corresponding server based on the predicted fluctuation parameter; If the actual processing duration is less than or equal to the second preset actual processing duration and greater than the first preset actual processing duration, adjusting the number of servers to the corresponding value based on the average evaluation value of the predicted evaluation values of the data to be processed for each slice; If the actual processing duration is greater than the second preset actual processing duration, determining the data repetition rate based on the data content of the data to be processed for a single slice, and adjusting the parameters of the corresponding server based on the data repetition rate.

5. The dynamic resource optimization method based on large model incremental learning according to claim 4, wherein The process of adjusting the parameters of the corresponding server based on the predicted fluctuation parameter includes: Denote the variance of the predicted evaluation values of the data to be processed for each slice calculated as the predicted fluctuation parameter; If the predicted fluctuation parameter is less than or equal to the preset predicted fluctuation parameter, adjust the network bandwidth of the server corresponding to the data to be processed for a single slice to the corresponding value based on the actual processing duration; If the predicted fluctuation parameter is greater than the preset predicted fluctuation parameter, adjust the batch selection quantity of the training data used to update the model selected based on the predicted fluctuation parameter to the corresponding value.

6. The dynamic resource optimization method based on large model incremental learning according to claim 5, characterized in that, Adjust the number of servers to the corresponding value based on the average evaluation value, where The increase amplitude of the number of servers is proportional to the average evaluation value.

7. The dynamic resource optimization method based on large model incremental learning according to claim 6, wherein The process of determining the data repetition rate based on the data content of the data to be processed for a single slice and adjusting the parameters of each server based on the data repetition rate includes: Determine the data sender corresponding to the data to be processed for a single slice; Obtain the data sent by the data sender in each previous transmission; Calculate the repetition rate between the data sent in each transmission respectively; Solve the average value of each repetition rate to obtain the data repetition rate; If the data repetition rate is less than or equal to the preset data repetition rate, adjust the number of servers to the corresponding value based on the average evaluation value; If the data repetition rate is greater than the preset data repetition rate, mark the single data sender as an orderly sender.

8. The dynamic resource optimization method based on large model incremental learning according to claim 7, wherein Adjust the network bandwidth of the server corresponding to the data to be processed for a single slice to the corresponding value based on the actual processing duration, where The increase amplitude of the network bandwidth is proportional to the actual processing duration.

9. The dynamic resource optimization method based on large model incremental learning according to claim 8, wherein, Adjust the batch selection quantity of the training data used for model update selected based on the predicted fluctuation parameter to the corresponding value, where The increase amplitude of the batch selection quantity is proportional to the predicted fluctuation parameter.

10. A dynamic resource optimization system using the dynamic resource optimization method based on large model incremental learning according to any one of claims 1-9, characterized in that, It includes: A data output module, which includes several data senders for sending data to be processed; A model processing module, which is connected to the data output module and is used to divide the data sent by each data sender into data to be processed for several slices, and determine the optimal resource allocation scheme for each server to process the data to be processed and the processing duration of the server for processing each data to be processed; A data receiving module, which is respectively connected to the data output module and the model processing module, and includes several servers for receiving and processing the data to be processed for each slice sent by each data sender according to the optimal resource allocation scheme; A data statistics module, which is connected to the data receiving module and is used to count the actual processing duration of each server for processing the data to be processed for each slice; A data division module, which is respectively connected to the model processing module, the data receiving module and the data statistics module, and is used to divide the received data volume, the processed data volume and the corresponding actual processing duration of the server counted according to the data volume into several batches of training data; A model training module, which is respectively connected to the data division module and the model processing module, and is used to select several batches of training data to update the model; The data analysis module is respectively connected to the model processing module, the data statistics module, the data receiving module, and the model training module, and is used to determine a prediction evaluation value based on the predicted processing duration and the actual processing duration, determine whether the operating condition of the model is qualified based on the prediction evaluation value, and in the case of determining that the operating condition of the model is abnormal, adjust the parameters of the corresponding server based on the actual processing duration of the data to be processed in a single slice, including adjusting the number of servers to the corresponding value, adjusting the network bandwidth of the server corresponding to the data to be processed in a single slice to the corresponding value, or adjusting the selected number of batches of training data used to update the model to the corresponding value.

Citation Information

Patent Citations

  • Cloud server resource allocation method

    CN116614396A

  • Model operation method, device, and service server

    CN109242135A

  • Battery performance test method based on data analysis

    CN118275903A

  • Server resource utilization rate improving method, system, device, medium and product

    CN119292776A

  • Intelligent work order management method, system, device and medium

    CN119784005A