AI model efficient training method and system based on dynamic resource allocation

Through real-time monitoring and dynamic resource allocation, the problem of AI systems being unable to optimize resource allocation is solved, automated performance management and fault prediction are realized, and operation and maintenance costs are reduced.

CN120336024AInactive Publication Date: 2025-07-18WUHAN TUOQIANHAI NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510476748.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art cannot allocate resources to AI systems based on the real-time processing volume, resulting in performance being unoptimized and increasing complexity and cost.

Method used

By acquiring and preprocessing training resource data, monitoring CPU, memory and storage values in real time, using performance calculations and fault matching formulas to predict system failures, and dynamic resource allocation based on failure rate to ensure optimal system performance.

Benefits of technology

It realizes automated and intelligent performance management, reduces manual intervention, improves resource utilization, reduces operation and maintenance costs, and ensures that the system automatically adjusts under performance pressure to maintain optimal condition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336024A_ABST
    Figure CN120336024A_ABST
Patent Text Reader

Abstract

The invention discloses an AI model efficient training method and system based on dynamic resource allocation. The method comprises the following steps: acquiring training resource data required by an AI model; according to the method, numerical values of the CPU, the memory and the storage are monitored in real time, fluctuation or abnormity of system performance can be found in time, service interruption or quality reduction caused by performance reduction is reduced, similarity calculation can be carried out on real-time performance and performance data before historical faults through a preset performance calculation formula and a fault matching degree calculation formula, and the fault matching degree is calculated. Therefore, the system fault can be accurately predicted, measures can be taken in advance, the influence of the fault on the system can be avoided or reduced, dynamic resource allocation is carried out according to a preset resource scheduling model based on the fault rate, automatic and intelligent performance management and fault prediction capability are realized, the requirement of manual intervention is reduced, and the system fault prediction efficiency is improved. Meanwhile, by optimizing the resource allocation, the resource utilization rate can be improved, and the cost can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of dynamic resource allocation, and specifically relates to an efficient training method and system for an AI model based on dynamic resource allocation. Background Art

[0002] In the current digital age, the application of artificial intelligence (AI) technology has penetrated into all fields, bringing great potential and opportunities for social development. Traditional AI construction methods often rely on professional teams of data scientists and engineers, requiring manual steps such as data collection, feature extraction, model selection, and tuning, which consume a large amount of time and labor costs.

[0003] During the model training and optimization process, it is necessary to continuously adjust parameters and algorithms, and there are often problems of high trial-and-error costs and low efficiency. In addition, once the model is built, its operation and management also require professional technicians for monitoring and maintenance, increasing the complexity and cost of the intelligent AI system. The prior art cannot allocate resources according to the real-time processing volume of the AI system to ensure its best performance. Summary of the Invention

[0004] To solve the above technical problems, an efficient training method and system for an AI model based on dynamic resource allocation are provided. The technical solution of the present invention solves the problem that the prior art cannot allocate resources according to the real-time processing volume of the AI system to ensure its best performance as proposed in the above background art.

[0005] To achieve the above object, the technical solution adopted by the present invention is as follows: In the first aspect of the present invention, an efficient training method for an AI model based on dynamic resource allocation is provided, including: Obtain the training resource data required by the AI model, and preprocess the training resource data required by the AI model. The preprocessing includes data deduplication, outlier removal, and data filling to obtain the standard data required for training the AI model; Retrieve the historical performance data of the system from the database, divide the historical performance data into a training set, a validation set, and a test set, and based on the neural network model, obtain the best performance state of the system. The historical performance data includes CPU, memory, and storage values; Input the standard data required for training the AI model into the system for training, and monitor the values of CPU, memory, and storage in real time; Input the real-time values of CPU, memory, and storage into a preset performance calculation formula to obtain the real-time performance; Based on the real-time performance and the real-time values of CPU, memory, and storage, use a fault matching degree calculation formula to calculate the similarity with the performance data before the historical fault to determine the failure rate; Based on the magnitude of the failure rate, dynamic resource allocation is performed according to a pre-set resource scheduling model to ensure the optimal performance state of the system.

[0006] Preferably, the data deduplication of the training resource data required by the AI model specifically includes the following steps: Obtain the training resource data required by the AI model, and sort it to obtain sorted data; Obtain the first piece of data in the sorted data and use it as the target data; Traverse the target data, and judge one by one whether it is repeated with the target data. If so, delete the data that is repeated with the target data in sequence, update the sorting plan data, and redefine the target data; Continue to traverse the newly defined target data and repeat the operation until the training resource data required by the AI model without duplicates is obtained.

[0007] Preferably, the outlier removal of the training resource data required by the AI model specifically includes the following steps: Arrange all the training resource data required by the AI model in the form of coordinates (x, y) as a column; Calculate the deviation value k between the coordinates (x, y) of the training resource data required by the AI model and the coordinates (a, b) and (c, d) of the adjacent training resource data required by the AI model; When the deviation value k is greater than 100%, the coordinates (x, y) of the training resource data required by the AI model are outliers, and the coordinates (x, y) of the training resource data required by the AI model are replaced by the mean value of the adjacent training resource data; The deviation value formula is: .

[0008] Preferably, the data filling of the training resource data required by the AI model specifically includes the following steps: Arrange all the training resource data required by the AI model in the form of coordinates (x, y) as a column; Retrieve all data coordinates (x, y). If x or y is missing, fill in the data at that position; The calculation formula for the mean value is: ; In the formula, Y is the missing value to be filled in the retrieved sequence, is the data of the i-th non-missing value in the sequence, and n is the number of non-missing values in the sequence.

[0009] Preferably, the pre-set performance calculation formula is: ; In the formula, is a performance value; is the CPU performance value; is the memory performance value; is the storage performance value; are all coefficients of green performance.

[0010] Preferably, the calculation method is: Classify the historical data according to whether the system finally needs to allocate, and obtain several groups of historical data that the system resources finally need to allocate and several groups of historical data that the system resources do not need to allocate; According to several groups of historical data that the system resources finally need to allocate and several groups of historical data that the system resources do not need to allocate, perform calculation by the maximum likelihood method; Verify the significance of the parameters of green performance, and judge whether it meets the significance requirements.

[0011] Preferably, the fault matching degree calculation formula is:

[0012] In the formula, is the similarity of the real-time values of CPU, memory and storage to the i-th feature of the performance data before the historical fault, is the j-th feature index value of the i-th feature of the real-time values of CPU, memory and storage, is the j-th feature index value of the i-th feature of the performance data before the historical fault, is the total number of the i-th feature index values; If the calculated value is greater than or equal to the preset threshold, allocation is required, otherwise it is not.

[0013] Preferably, the resource scheduling model is:

[0014] In the formula, is the resource to be allocated currently, is the number of several required resources, is the predicted idle time of the i-th resource, is the total number of resources, is the linear regression equation function of the predicted idle time of the k-th resource for the i-th resource, is the scheduling time.

[0015] In a second aspect of the present invention, an efficient training system for an AI model based on dynamic resource allocation is further provided, including: A preprocessing module, which is used to obtain the training resource data required by the AI model and preprocess the training resource data required by the AI model. The preprocessing includes data deduplication, outlier removal, and data filling to obtain the standard data required for training the AI model; A determination module, which is used to retrieve the historical performance data of the system from the database, divide the historical performance data into a training set, a validation set, and a test set, and based on a neural network model, obtain the optimal performance state of the system. The historical performance data includes CPU, memory, and storage values; A monitoring module, which is used to input the standard data required for training the AI model into the system for training and to monitor the values of CPU, memory, and storage in real time; A first calculation module, which is used to input the real-time values of CPU, memory, and storage into a preset performance calculation formula to obtain the real-time performance; A second calculation module, which is used to calculate the similarity between the real-time performance and the real-time values of CPU, memory, and storage using a fault matching degree calculation formula and compare it with the performance data before the historical fault to determine the failure rate; An allocation module, which is used to allocate dynamic resources according to a preset resource scheduling model based on the size of the failure rate to ensure the optimal performance state of the system.

[0016] In a third aspect of the present invention, an electronic device is further provided. The electronic device has at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method of the first aspect of the present invention.

[0017] Compared with the prior art, the present invention provides an efficient training method and system for an AI model based on dynamic resource allocation, which has the following beneficial effects: By preprocessing the training resource data required for the AI model, the present invention ensures that the data input into the AI model is accurate and reliable, thereby improving the training effect and prediction accuracy of the model. By retrieving the historical performance data of the system from the database and obtaining the optimal performance state of the system based on the neural network model, the normal operating state of the system can be understood more precisely, which helps to formulate more reasonable performance management strategies. By real-time monitoring the values of CPU, memory, and storage, fluctuations or anomalies in system performance can be detected in a timely manner, reducing service interruptions or quality degradation caused by performance decline. By using the preset performance calculation formula and fault matching degree calculation formula, the similarity between the real-time performance and the performance data before historical faults can be calculated, thereby accurately predicting system faults, which helps to take measures in advance to avoid or reduce the impact of faults on the system. Based on the size of the failure rate, dynamic resource allocation is performed according to the preset resource scheduling model, which can ensure that the system can automatically adjust resources when facing performance pressure to maintain the optimal performance state, improving the flexibility and scalability of the system, thereby realizing the capabilities of automated and intelligent performance management and fault prediction, reducing the need for manual intervention, and lowering the operation and maintenance costs. At the same time, by optimizing resource allocation, resource utilization can also be improved, further reducing costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 FIG. is a schematic diagram of an efficient training method for an AI model based on dynamic resource allocation in the present invention; Figure 2 FIG. is a schematic diagram of a method for removing duplicate data from the training resource data required for the AI model in the present invention; Figure 3 FIG. is a schematic diagram of a method for removing outliers from the training resource data required for the AI model in the present invention; Figure 4 FIG. is a schematic diagram of a method for filling in data for the training resource data required for the AI model in the present invention; Figure 5 In the present invention Calculation method schematic diagram. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] The following description is used to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments described below are only examples, and those skilled in the art can think of other obvious variations.

[0020] Embodiment 1 Please refer to Figure 1 As shown, in the first aspect of the present invention, an efficient training method for an AI model based on dynamic resource allocation is provided, including: S101. Obtain the training resource data required for the AI model, and preprocess the training resource data required for the AI model. The preprocessing includes data de-duplication, outlier removal, and data filling, so as to obtain the standard data required for training the AI model. S102. Retrieve the historical performance data of the system from the database, divide the historical performance data into a training set, a validation set, and a test set, and based on the neural network model, obtain the optimal performance state of the system. The historical performance data includes CPU, memory, and storage values. S104. Input the real-time values of CPU, memory, and storage into the pre-set performance calculation formula to obtain the real-time performance. S104. Input the real-time values of CPU, memory, and storage into the pre-set performance calculation formula to obtain the real-time performance. S105. Based on the real-time performance and the real-time values of CPU, memory, and storage, use the failure matching degree calculation formula to calculate the similarity with the performance data before the historical failure, and determine the failure rate. S106. Based on the size of the failure rate, perform dynamic resource allocation according to the pre-set resource scheduling model to ensure the optimal performance state of the system.

[0021] Those skilled in the art can understand that through the preprocessing of the training resource data required for the AI model, the data input into the AI model is ensured to be accurate and reliable, thereby improving the training effect and prediction accuracy of the model. By retrieving the historical performance data of the system from the database and obtaining the optimal performance state of the system based on the neural network model, the normal operating state of the system can be understood more accurately, which helps to formulate more reasonable performance management strategies. By real-time monitoring the values of CPU, memory, and storage, fluctuations or anomalies in the system performance can be discovered in time, reducing service interruptions or quality degradation caused by performance degradation. Through the pre-set performance calculation formula and failure matching degree calculation formula, the real-time performance can be calculated for similarity with the performance data before the historical failure, thereby accurately predicting system failures, which helps to take measures in advance to avoid or reduce the impact of failures on the system. Based on the size of the failure rate, perform dynamic resource allocation according to the pre-set resource scheduling model, which can ensure that the system can automatically adjust resources when facing performance pressure to maintain the optimal performance state, improving the flexibility and scalability of the system, thereby realizing the capabilities of automated and intelligent performance management and failure prediction, reducing the need for manual intervention, and reducing the operation and maintenance costs. At the same time, by optimizing resource allocation, the resource utilization rate can also be improved, further reducing costs.

[0022] Please refer to Figure 2 As shown, the specific steps for data de-duplication of the training resource data required for the AI model are as follows: S201. Obtain the training resource data required for the AI model and sort it to obtain sorted data; S202. Obtain the first piece of data in the sorted data and use it as the target data; S203. Traverse the target data and judge one by one whether it is repeated with the target data. If so, delete the data that is repeated with the target data in turn, update the sorted planning data, and re-define the target data; S204. Continue to traverse the newly defined target data and repeat the operation until the training resource data required for the AI model without duplicates is obtained.

[0023] Please refer to Figure 3 As shown in the figure, the specific steps for removing outliers from the training resource data required for the AI model are as follows: S301. Arrange all the training resource data required for the AI model in the form of coordinates (x, y) as a column; S302. Calculate the deviation value k between the coordinates (x, y) of the training resource data required for the AI model and the coordinates (a, b) and (c, d) of the adjacent training resource data required for the AI model; S303. When the deviation value k is greater than 100%, the coordinates (x, y) of the training resource data required for the AI model are outliers, and the coordinates (x, y) of the training resource data required for the AI model are replaced by the mean value of the adjacent training resource data; The deviation value formula is: .

[0024] Please refer to Figure 4 As shown in the figure, the specific steps for data filling of the training resource data required for the AI model are as follows: S401. Arrange all the training resource data required for the AI model in the form of coordinates (x, y) as a column; S402. Retrieve all data coordinates (x, y). If x or y is missing, fill in the data at that position; The formula for calculating the mean value is: ; In the formula, Y is the missing value to be filled in the retrieved sequence, is the data of the i-th non-missing value in the sequence, and n is the number of non-missing values in the sequence.

[0025] The pre-set performance calculation formula is: ; In the formula, is the performance value; is the CPU performance value; is the memory performance value; is the storage performance value; are all coefficients of green performance.

[0026] Please refer to Figure 5 as shown, The calculation method of S501. Classify historical data according to whether the system finally needs to allocate, and obtain several groups of historical data that the system resources finally need to allocate and several groups of historical data that the system resources do not need to allocate; S502. According to several groups of historical data that the system resources finally need to allocate and several groups of historical data that the system resources do not need to allocate, calculate by the maximum likelihood method; S503. Test the significance of the parameters of green performance, and judge whether it meets the significance requirement.

[0027] The failure matching degree calculation formula is:

[0028] In the formula, is the similarity of the real-time values of CPU, memory and storage and the i-th feature of the performance data before historical faults, is the j-th feature index value of the i-th feature of the real-time values of CPU, memory and storage, is the j-th feature index value of the i-th feature of the performance data before historical faults, is the total number of the i-th feature index values; If the calculated value is greater than or equal to the preset threshold, allocation is required, otherwise it is not.

[0029] The resource scheduling model is:

[0030] In the formula, is the resource to be allocated currently, is the number of several required resources, is the predicted idle time of the i-th resource, is the total number of resources, is the linear regression equation function of the predicted idle time of the k-th resource for the i-th resource, is the scheduling time.

[0031] In the second aspect of the present invention, an efficient training system for an AI model based on dynamic resource allocation is further provided, including: A preprocessing module, which is used to obtain the training resource data required by the AI model and preprocess the training resource data required by the AI model. The preprocessing includes data deduplication, outlier removal, and data filling to obtain the standard data required for training the AI model; A determination module, which is used to retrieve the historical performance data of the system from the database, divide the historical performance data into a training set, a validation set, and a test set, and obtain the optimal performance state of the system based on a neural network model. The historical performance data includes CPU, memory, and storage values; A monitoring module, which is used to input the standard data required for training the AI model into the system for training and monitor the values of CPU, memory, and storage in real time; A first calculation module, which is used to input the real-time values of CPU, memory, and storage into a preset performance calculation formula to obtain the real-time performance; A second calculation module, which is used to calculate the similarity between the real-time performance and the real-time values of CPU, memory, and storage using a fault matching degree calculation formula and compare it with the performance data before the historical fault to determine the failure rate; An allocation module, which is used to dynamically allocate resources according to a preset resource scheduling model based on the size of the failure rate to ensure the optimal performance state of the system.

[0032] In the third aspect of the present invention, an electronic device is further provided.

[0033] The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described herein and / or claimed.

[0034] The electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0035] Multiple components in the electronic device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, an optical disc, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the electronic device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0036] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above, such as methods S101 - S106. For example, in some embodiments, methods S101 - S106 can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of methods S101 - S106 described above can be executed. Alternatively, in other embodiments, the computing unit 601 can be configured to execute methods S101 - S106 in any other appropriate way (e.g., by means of firmware).

[0037] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0038] The program code for implementing the methods of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on the remote machine or server.

[0039] In the context of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0040] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0041] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0042] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, can also be a server of a distributed system, or a server incorporating a blockchain.

[0043] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and what is described in the above embodiments and the specification is only the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will also have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. An efficient training method for AI models based on dynamic resource allocation, characterized in that, Including: Obtain the training resource data required for the AI model, and preprocess the training resource data required for the AI model. The preprocessing includes data deduplication, outlier removal, and data filling, and obtain the standard data required for training the AI model; Retrieve the historical performance data of the system from the database, divide the historical performance data into a training set, a validation set, and a test set, and based on the neural network model, obtain the optimal performance state of the system. The historical performance data includes CPU, memory, and storage values; Input the standard data required for training the AI model into the system for training, and monitor the values of CPU, memory, and storage in real time; Input the real-time values of CPU, memory, and storage into a preset performance calculation formula to obtain the real-time performance; Based on the real-time performance and the real-time values of CPU, memory, and storage, use the failure matching degree calculation formula to calculate the similarity with the performance data before the historical failure, and determine the failure rate; Based on the magnitude of the failure rate, perform dynamic resource allocation according to the preset resource scheduling model to ensure the optimal performance state of the system.

2. The efficient training method for an AI model based on dynamic resource allocation according to claim 1, wherein, The specific steps for data deduplication of the training resource data required for the AI model are as follows: Obtain the training resource data required for the AI model, and sort it to obtain sorted data; Obtain the first piece of data in the sorted data and use it as the target data; Traverse the target data, and judge one by one whether it is repeated with the target data. If so, delete the data that is repeated with the target data in sequence, update the sorted planning data, and redefine the target data; Continue to traverse the newly defined target data, repeat the operation until the training resource data required for the AI model without duplicates is obtained.

3. The efficient training method of the AI model based on dynamic resource allocation according to claim 2, wherein, The specific steps for outlier removal of the training resource data required for the AI model are as follows: Arrange all the training resource data required for the AI model in the form of coordinates (x, y) as a column; Calculate the deviation value k between the coordinates (x, y) of the training resource data required for the AI model and the coordinates (a, b) and (c, d) of the adjacent training resource data required for the AI model; When the deviation value k is greater than 100%, the coordinates (x, y) of the training resource data required for the AI model are outliers, and the coordinates (x, y) of the training resource data required for the AI model are replaced by the mean value of the adjacent training resource data; The deviation value formula is: 。 4. The efficient training method of the AI model based on dynamic resource allocation according to claim 3, wherein The specific steps for data filling of the training resource data required for the AI model are as follows: Arrange all the training resource data required for the AI model in the form of coordinates (x, y) as a column; Retrieve all data coordinates (x, y), and if x or y is missing, fill in the data at that position; The calculation formula for the mean value is: ; Where Y is the missing value to be filled in the retrieved sequence, is the data of the i-th non-missing value in the sequence, and n is the number of non-missing values in the sequence.

5. The method for efficiently training an AI model based on dynamic resource allocation according to claim 4, wherein The preset performance calculation formula is: ; In the formula, is the performance value; is the CPU performance value; is the memory performance value; For the storage performance value; All are coefficients of green performance.

6. The method for efficiently training an AI model based on dynamic resource allocation according to claim 5, wherein The said is calculated as follows: Classify the historical data according to whether the system finally needs to allocate, and obtain several groups of historical data for which the system resources finally need to be allocated and several groups of historical data for which the system resources do not need to be allocated; Based on several groups of historical data of the system resources that ultimately need to be allocated and several groups of historical data of the system resources that do not need to be allocated, perform the calculation; Inspection Judge the significance of the parameters of the green performance Whether it meets the significance requirements 7. The method for efficiently training an AI model based on dynamic resource allocation according to claim 6, wherein The failure matching degree calculation formula is: ; Wherein, is the similarity of the real-time values of the CPU, memory, and storage to the i-th feature of the performance data before the historical failure, is the j-th feature index value of the i-th feature of the real-time values of the CPU, memory, and storage, is the j-th feature index value of the i-th feature of the performance data before the historical failure, is the total number of the i-th feature index values; If the calculated value is greater than or equal to the preset threshold, allocation is required; otherwise, it is not.

8. The method for efficiently training an AI model based on dynamic resource allocation according to claim 7, wherein The resource scheduling model is: ; Wherein, is the resource to be currently allocated, is the number of several required resources, is the predicted idle time of the i-th resource, is the total number of resources, is the linear regression equation function of the predicted idle time of the k-th resource for the i-th resource, is the scheduling time.

9. An efficient training system for an AI model based on dynamic resource allocation, which is used to implement the efficient training method for an AI model based on dynamic resource allocation according to any one of claims 1-8, characterized in that, Including: A preprocessing module, which is used to obtain the training resource data required by the AI model and preprocess the training resource data required by the AI model. The preprocessing includes data deduplication, outlier removal, and data filling to obtain the standard data required for training the AI model. A determination module, which is used to retrieve the historical performance data of the system from the database, divide the historical performance data into a training set, a validation set, and a test set, and obtain the optimal performance state of the system based on a neural network model. The historical performance data includes CPU, memory, and storage values. A monitoring module, which is used to input the standard data required for training the AI model into the system for training and monitor the values of CPU, memory, and storage in real time. A first calculation module, which is used to input the real-time values of CPU, memory, and storage into a preset performance calculation formula to obtain the real-time performance. A second calculation module, which is used to calculate the similarity between the real-time performance and the real-time values of CPU, memory, and storage using a fault matching degree calculation formula and compare it with the performance data before the historical fault to determine the failure rate. An allocation module, which is used to dynamically allocate resources according to a preset resource scheduling model based on the size of the failure rate to ensure the optimal performance state of the system.

10. An electronic device, comprising at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-8.