Method and device for intelligent computing center cloud platform to adjust large model training task based on computing power use state

By acquiring and adjusting the computing power usage status of the intelligent computing center cloud platform, the problems of waste of computing power and low efficiency during the training process are solved, and more efficient training task optimization is achieved.

CN120407211AActive Publication Date: 2025-08-01DATACANVAS LTD

Patent Information

Application Number
CN202510916796.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-08-01
Estimated Expiration
2045-07-03

AI Technical Summary

Technical Problem

During the training of large models by the intelligent computing center cloud platform, when the computing power usage state does not reach the expected state, it leads to waste of computing power and low training efficiency.

Method used

By obtaining the computing power usage status, determining whether it matches the expected state, and adjusting the training tasks of the big model when it does not match, including adjusting the core utilization of the computing card, bandwidth of data exchanged between memory and computing card, memory occupancy, memory throughput, and throughput between different computing cards, etc., to optimize the training tasks.

Benefits of technology

Reduces waste of computing power and improves the training efficiency of large models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407211A_ABST
    Figure CN120407211A_ABST
Patent Text Reader

Abstract

The invention provides a method and a device for adjusting a large model training task by an intelligent computing center cloud platform based on a computing power use state, and relates to the technical field of intelligent computing centers, intelligent computing centers, computing power infrastructures and intelligent computing clouds, in particular to the method for adjusting the large model training task by the intelligent computing center cloud platform based on the computing power use state. Comprising the steps that S1, a computing power use state is obtained, and the computing power use state is used for representing the use state of a training task on computing power in an intelligent computing center cloud platform in the large model training process of the intelligent computing center cloud platform; s2, determining whether the computing power use state is matched with an expected state or not; and S3, under the condition that the computing power use state is not matched with the expected state, adjusting a training task of the large model according to the computing power use state. In this way, the phenomenon that computing power is wasted can be reduced, and the training efficiency of a large model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of intelligent computing centers, intelligent computing centers, computing power infrastructure, and intelligent computing clouds, and particularly relates to a method and device for adjusting large model training tasks based on the usage status of computing power in an intelligent computing center cloud platform. Background Art

[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged as the times require.

[0003] An "intelligent computing center" refers to a facility that provides the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios like artificial intelligence deep learning model development, model training, and model inference) by using large-scale heterogeneous computing power resources, including general computing power and intelligent computing power. An intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.

[0004] The "intelligent computing center" includes but is not limited to the "intelligent computing center".

[0005] An "intelligent computing center", that is, an artificial intelligence computing center, is a type of computing power infrastructure that provides computing power services, data services, and algorithm services required for artificial intelligence applications based on artificial intelligence theory and using an artificial intelligence computing architecture.

[0006] "Computing power" is the core of "intelligent computing centers" and "intelligent computing centers", which is the ability of computer devices or computing / data centers to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to achieve the output of target results by processing information data, and a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.

[0007] When currently training a large model on an intelligent computing center cloud platform, it is usually necessary to use the computing power of the intelligent computing center cloud platform to train the large model. However, during the process of training a large model on the intelligent computing center cloud platform currently, when the usage status of the computing power does not reach the expected status, it is easy to waste the computing power and make the training efficiency of the large model very low. It can be seen that since the emergence of intelligent computing centers, when the usage status of the computing power does not reach the expected status, it is easy to waste the computing power and make the training efficiency of the large model very low, which is an urgent problem to be solved. Summary of the Invention

[0008] The present invention provides a method and device for adjusting large model training tasks based on the usage status of computing power in an intelligent computing center cloud platform, which is used to solve the problem that when the usage status of the computing power does not reach the expected status, it is easy to waste the computing power and make the training efficiency of the large model very low.

[0009] To solve the above technical problems, the present invention is implemented as follows: In a first aspect, the present invention provides a method for adjusting a large model training task based on the computing power usage status of an intelligent computing center cloud platform, including: Step S1: Obtain the computing power usage status, which is used to represent the usage status of the computing power in the intelligent computing center cloud platform by the training task during the training of the large model; Step S2: Determine whether the computing power usage status matches the expected status; Step S3: In the case where the computing power usage status does not match the expected status, adjust the training task of the large model according to the computing power usage status.

[0010] Optionally, the computing power usage status includes at least one of the following: the core utilization rate of the computing card, the bandwidth for exchanging data between the memory and the computing card, the video memory occupancy rate, the video memory throughput, and the throughput between different computing cards.

[0011] Optionally, the computing power usage status includes the core utilization rate of the computing card, and step S3 includes: Step S31: In the case where the calculated volatility of the core utilization rate of the computing card within a preset period is greater than the expected volatility, determine that the computing power usage status does not match the expected status; Step S32: In the case where the computing power usage status does not match the expected status, adjust the storage level of the training data in the training task according to the fluctuation status of the core utilization rate of the computing card, and the fluctuation status includes the volatility.

[0012] Optionally, the computing power usage status includes the throughput between different computing cards, and step S3 includes: Step S31': In the case where the calculated throughput between different computing cards is greater than the first throughput, determine that the computing power usage status does not match the expected status; Step S32': In the case where the computing power usage status does not match the expected status, adjust the batch sample number of a single computing card according to the throughput between different computing cards.

[0013] Optionally, the computing power usage status includes the video memory throughput, and step S3 includes: Step S31'': In the case where the calculated video memory throughput is greater than the second throughput, determine that the computing power usage status does not match the expected status; Step S32'': In the case where the computing power usage status does not match the expected status, adjust the parallelism of the computing power according to the video memory throughput.

[0014] Optionally, the computing power usage status includes the bandwidth of data exchange between the memory and the computing card and the video memory occupancy rate, and the step S3 includes: Step S31’’’: When the video memory occupancy rate does not reach the preset occupancy rate and the bandwidth of data exchange between the memory and the computing card reaches the preset bandwidth, it is determined that the computing power usage status does not match the expected status; Step S32’’’: Adjust the proportion of the model weight data unloaded to the memory according to the bandwidth of data exchange between the memory and the computing card and the video memory occupancy rate.

[0015] In a second aspect, the present invention provides an apparatus for adjusting a large model training task based on the computing power usage status in an intelligent computing center cloud platform, including: An acquisition module, configured to acquire the computing power usage status, where the computing power usage status is used to represent the usage status of the computing power in the intelligent computing center cloud platform by the training task during the training of the large model; A determination module, configured to determine whether the computing power usage status matches the expected status; An adjustment module, configured to adjust the training task of the large model according to the computing power usage status when the computing power usage status does not match the expected status.

[0016] In a third aspect, the present invention provides an electronic device, including: a processor, a memory, and a program stored on the memory and executable on the processor, where when the program is executed by the processor, the steps of the method for adjusting the large model training task based on the computing power usage status in the intelligent computing center cloud platform as described in the first aspect above are implemented.

[0017] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method for adjusting the large model training task based on the computing power usage status in the intelligent computing center cloud platform as described in the first aspect above are implemented.

[0018] In a fifth aspect, the present invention provides a computer program product, including computer instructions, and when the computer instructions are executed by a processor, the steps of the method for adjusting the large model training task based on the computing power usage status in the intelligent computing center cloud platform as described in the first aspect above are implemented.

[0019] In the present invention, the computing power usage status is obtained, and the computing power usage status is used to represent the usage status of the intelligent computing center cloud platform for computing power during the process of training a large model; it is determined whether the computing power usage status matches the expected status; in the case where the computing power usage status does not match the expected status, the training task of the large model is adjusted according to the computing power usage status. In this way, when it is determined that the computing power usage status does not match the expected status, the training task of the large model can be adjusted according to the computing power usage status, thereby reducing the occurrence of the phenomenon of wasted computing power and improving the training efficiency of the large model. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings: Figure 1 It is a flowchart of a method for an intelligent computing center cloud platform to adjust the training task of a large model based on the computing power usage status provided by the present invention; Figure 2 It is a flowchart of a method for training a large model provided by the present invention; Figure 3 It is a schematic structural diagram of a device for an intelligent computing center cloud platform to adjust the training task of a large model based on the computing power usage status provided by the present invention; Figure 4 It is a schematic structural diagram of an electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] The technical solutions in the present invention will be clearly and completely described below with reference to the drawings in the present invention. Obviously, the described content is a part of the present invention, not all of the content. Based on the content in the present invention, all other content obtained by those of ordinary skill in the art without creative efforts belongs to the scope of protection of the present invention.

[0022] The "computing power" referred to in the present invention means: the ability of a computer device or a computing / data center to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to achieve the output of a target result by processing information data, a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly providing services to society through computing power infrastructure.

[0023] The "Computational Power (CP)" as described in the present invention refers to: an ability of the data center server to process data and output results, which is a comprehensive indicator for measuring the computing power of the data center and includes general computing power, supercomputing power, and intelligent computing power. The commonly used measurement unit is the number of floating-point operations per second (FLOPS, 1 EFLOPS = 10^18 FLOPS). The larger the value, the stronger the comprehensive computing power. It is estimated that 1 EFLOPS is approximately the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream laptops. The calculation formula is: CP = CP 通用 + CP 智能 + CP 超级 。

[0024] The "Network Power (NP)" as described in the present invention refers to: an expression of the data transmission ability of the computing power facilities, which is a comprehensive ability including network architecture, network bandwidth, transmission delay, intelligent management and scheduling, etc., and involves network transmission within and between data centers, and is a comprehensive indicator for measuring network transmission scheduling ability.

[0025] The "Storage Power (SP)" as described in the present invention refers to: a comprehensive ability of the data center in four aspects of data storage capacity, performance, security and reliability, and green and low-carbon, which is a comprehensive indicator for measuring the data storage ability of the data center and includes external storage devices such as storage arrays and server internal storage devices. The commonly used measurement unit for storage capacity is exabyte (EB, 1 EB = 2^60 bytes), and the commonly used measurement unit for performance is the number of read and write operations per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB). The disaster recovery ratio is an important manifestation of security and reliability.

[0026] The "computing power infrastructure" as described in the present invention refers to: a new type of information infrastructure integrating information computing power, network carrying power, and data storage power, which can realize centralized computing, storage, transmission, and application of information.

[0027] The "new type of information infrastructure" as described in the present invention mainly includes network infrastructures such as 5G networks, fiber broadband networks, backbone networks, international communication networks, and satellite Internet, computing power infrastructures such as data centers, general computing power centers, intelligent computing centers, and supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing.

[0028] The "computing power" as described in the present invention includes: general computing power, intelligent computing power, and supercomputing power.

[0029] The "general computing power" described in the present invention refers to the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.

[0030] The "intelligent computing power" described in the present invention refers to the computing platforms deployed on a large scale based on dedicated chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit) for various artificial intelligence innovation applications, such as natural language processing, machine vision, etc.

[0031] The "super computing power" described in the present invention mainly refers to the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and processes extremely complex or data-intensive problems through a dedicated operating system. It is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, gene analysis, etc.

[0032] The "intelligent computing center" described in the present invention refers to a facility that provides the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios of artificial intelligence deep learning model development, model training, and model inference) by using large-scale heterogeneous computing power resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.

[0033] The "intelligent computing center cloud platform" referred to as "intelligent computing cloud" in the present invention refers to a cloud computing platform that comprehensively serves based on the hardware resources and software resources of the intelligent computing center.

[0034] The "intelligent computing center" described in the present invention includes, but is not limited to, the "intelligent computing center".

[0035] The "intelligent computing center" described in the present invention, namely the artificial intelligence computing center, is a type of computing power infrastructure that provides computing power services, data services, and algorithm services required for artificial intelligence applications based on artificial intelligence theory and using an artificial intelligence computing architecture.

[0036] The "computing power center" described in the present invention refers to a facility mainly composed of infrastructure such as wind, fire, water, and electricity and IT software and hardware devices, which has computing power, carrying capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.

[0037] The "supercomputing center" described in the present invention refers to: a supercomputing data center, which is a data center based on supercomputers or large-scale computing clusters and can provide functions such as large-scale computing, storage, and network services, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling, and genome sequencing.

[0038] The "computing power resources" described in the present invention refers to: technologies and facilities with information computing, transmission, storage, and application capabilities required for the development of the digital society, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and support and guarantee resources such as wind, fire, water, and electricity.

[0039] The "models" described in the present invention include but are not limited to "large language models" and "multimodal large models". The "large models" described in the present invention can be understood as the above-mentioned "models".

[0040] The "large language model" described in the present invention refers to a large language model (LLM), which is a language model with a relatively large number of parameters, aiming to understand and generate human language, trained through a large amount of text data, and can perform a wide range of tasks including text summarization, translation, sentiment analysis, etc.

[0041] The "multimodal large models" (Multimodal Large Models) described in the present invention refers to: models that jointly train multimodal information such as text, images, videos, and audio, including but not limited to multimodal large language models.

[0042] The "computing power usage status" described in the present invention refers to: the usage status of the computing power in the intelligent computing center cloud platform by the training task during the process of training large models in the intelligent computing center cloud platform.

[0043] Please refer to Figure 1 , Figure 1 which is a flowchart of a method for an intelligent computing center cloud platform to adjust large model training tasks based on the computing power usage status provided by the present invention. As Figure 1 shown, it includes the following steps: Step S1: Obtain the computing power usage status, which is used to represent the usage status of the computing power in the intelligent computing center cloud platform by the training task during the process of training large models in the intelligent computing center cloud platform.

[0044] Among them, the specific types of status parameters included in the computing power usage status are not limited here. Optionally, the computing power usage status includes at least one of the following: the core utilization rate of the computing card, the bandwidth for data exchange between the memory and the computing card, the video memory occupancy rate, the video memory throughput, and the throughput between different computing cards. The above-mentioned core utilization rate of the computing card, the bandwidth for data exchange between the memory and the computing card, the video memory occupancy rate, the video memory throughput, and the throughput between different computing cards can all be regarded as the status parameters included in the computing power usage status. In this way, the diversity and flexibility of the types of status parameters included in the computing power usage status can be increased.

[0045] It should be noted that each type of status parameter included in the computing power usage status has a corresponding expected status, that is, the status parameter and the expected status can be in one-to-one correspondence.

[0046] In addition, optionally, the status parameter, the training parameter, and the expected status can also be in one-to-one correspondence, and the types included in the training parameter are not limited here. For example, the training parameter can include at least one of the following: the training time period, the training process, the training samples, the training method. That is, when the training time period is different, the training process is different, the training samples input into the large model are different, and the training method for the large model is different, it may also lead to different corresponding expected statuses.

[0047] For example: when the status parameter included in the computing power usage status is the core utilization rate of the computing card, and the training parameter includes the training time period, the expected status corresponding to the core utilization rate of the computing card in the first time period can be that the utilization rate is greater than 70%, and the expected status corresponding to the core utilization rate of the computing card in the second time period can be that the utilization rate is greater than 80%. It can be seen that when the status parameters included in the computing power usage status are the same and the training parameters are different, the corresponding expected statuses are also different.

[0048] It should be noted that the more types of status parameters included in the computing power usage status, the more accurate the result of determining whether the computing power usage status matches the expected status.

[0049] Step S2: Determine whether the computing power usage status matches the expected status.

[0050] Among them, when there is one type of status parameter included in the computing power usage status, the status parameter included in the computing power usage status can be directly matched with the corresponding expected status; when there are multiple types of status parameters included in the computing power usage status, each type of status parameter can be matched with the corresponding expected status. When all status parameters match the corresponding expected statuses, it can be determined that the computing power usage status matches the expected status; when there is at least one type of status parameter that does not match the corresponding expected status, it can be determined that the computing power usage status does not match the expected status.

[0051] Among them, matching each state parameter with the corresponding expected state can be understood as: calculating the difference between the state parameter and the corresponding expected state. If the difference is within the preset range, it can be determined that the state parameter matches the corresponding expected state; if the difference is not within the preset range, it can be determined that the state parameter does not match the corresponding expected state.

[0052] For example: the state parameter can be the core utilization rate of the computing card, and the expected state is a utilization rate of 70%. If the core utilization rate of the computing card is 69.5%, and the preset range is less than 1%, then 70% - 69.5% = 0.5%, which is obviously less than 1%. Therefore, it can be determined that the core utilization rate of the computing card being 69.5% matches the expected state; if the core utilization rate of the computing card is 68.5%, and the preset range is less than 1%, then 70% - 68.5% = 1.5%, which is obviously greater than 1%. Therefore, it can be determined that the core utilization rate of the computing card being 69.5% does not match the expected state.

[0053] Step S3: In the case where the computing power usage state does not match the expected state, adjust the training task of the large model according to the computing power usage state.

[0054] In the case where the computing power usage state does not match the expected state, it indicates that there is a bottleneck in the computing power usage state at this time. It is necessary to adjust the training task of the large model, thereby improving the utilization rate of the computing power, reducing the waste of the computing power, and improving the training efficiency of the large model.

[0055] In the present invention, through steps S1 to S3, the computing power usage state is obtained. The computing power usage state is used to represent the state of the intelligent computing center cloud platform's use of computing power during the training of the large model; it is determined whether the computing power usage state matches the expected state; in the case where the computing power usage state does not match the expected state, the training task of the large model is adjusted according to the computing power usage state. In this way, when it is determined that the computing power usage state does not match the expected state, the training task of the large model can be adjusted according to the computing power usage state, thereby reducing the occurrence of the phenomenon of wasted computing power and improving the training efficiency of the large model.

[0056] It should be noted that if the types of state parameters included in the computing power usage state are different, the adjustment method and adjustment parameters of the training task of the large model are also different.

[0057] Optionally, the computing power usage state includes the core utilization rate of the computing card, and step S3 includes: Step S31: In the case where it is calculated that the volatility of the core utilization rate of the computing card within the preset period is greater than the expected volatility, determine that the computing power usage state does not match the expected state; Step S32: In the case where the computing power usage status does not match the expected status, adjust the storage level of the training data in the training task according to the fluctuation status of the core utilization rate of the computing card, where the fluctuation status includes the volatility rate.

[0058] Among them, the volatility rate of the core utilization rate of the computing card within a preset period can be obtained according to the difference between the maximum utilization rate and the minimum utilization rate of the computing card within the preset period. And the fluctuation status of the core utilization rate of the computing card can include the volatility rate of the core utilization rate of the computing card within the preset period. In addition, the fluctuation status of the core utilization rate of the computing card can also include the fluctuation duration and the fluctuation time period of the core utilization rate of the computing card within the preset period.

[0059] In the present invention, when the volatility rate of the core utilization rate of the computing card within a preset period is greater than the expected volatility rate, it indicates that there is a problem with the data reading speed of the computing card. Therefore, the storage level of the training data in the training task can be adjusted according to the fluctuation status of the core utilization rate of the computing card to improve the data reading speed. And the storage level of the training data in the above training task can be understood as the priority of storage and reading. The higher the storage level, the higher the priority of reading and storing the training data.

[0060] In addition, optionally, when the volatility rate of the core utilization rate of the computing card within a preset period is greater than the expected volatility rate, it can also indicate that there is a problem with the computing power allocation of the computing card. Therefore, the computing power allocated to the computing card can be adjusted. By adjusting the computing power allocated to the computing card, the volatility rate of the core utilization rate of the computing card can be reduced, so that the core utilization rate of the computing card always reaches the expectation, that is, the waste of computing power is reduced, and at the same time, the training efficiency of the large model can be improved.

[0061] Optionally, the computing power usage status includes the throughput between different computing cards, and step S3 includes: Step S31': In the case where the throughput between different computing cards calculated is greater than the first throughput, determine that the computing power usage status does not match the expected status; Step S32': In the case where the computing power usage status does not match the expected status, adjust the number of batch samples of a single computing card according to the throughput between different computing cards.

[0062] Among them, the number of batch samples of a single computing card can be referred to as batchsize, and the total number of samples can be understood as the number of samples input into the large model. Since the training task of the large model is executed on the intelligent computing center cloud platform, and the computing power of the intelligent computing center cloud platform can include multiple computing cards, the training task of the large model can be split into multiple subtasks, and each computing card can execute part of the subtasks. In this way, samples corresponding to the subtasks to be executed can be input on each computing card (the number of these samples can be referred to as the number of batch samples of a single computing card) to achieve the training of the subtasks.

[0063] Among them, the throughput between different computing cards can be understood as the data communication volume between different computing cards.

[0064] In the present invention, when the throughput between different computing cards is greater than the first throughput, it can be understood that the communication between different computing cards is too frequent, that is, it may be that the number of samples allocated on the above-mentioned computing cards is unreasonable, resulting in the need to communicate with other computing cards frequently. Therefore, the number of batch samples of a single computing card can be adjusted according to the throughput between different computing cards to improve the utilization rate of the computing power of each computing card, thereby enhancing the training efficiency of the subtasks and further improving the training efficiency of the large model.

[0065] Optionally, the computing power usage status includes the video memory throughput, and step S3 includes: Step S31'': In the case where it is calculated that the video memory throughput is greater than the second throughput, it is determined that the computing power usage status does not match the expected status; Step S32'': In the case where the computing power usage status does not match the expected status, adjust the parallelism of the computing power according to the video memory throughput.

[0066] In the present invention, the video memory throughput can be understood as the data communication volume between the computing card and the memory. When the video memory throughput is greater than the second throughput, it means that the load of this computing card has reached the upper limit at this time. At this time, the parallelism of the computing power can be adjusted, such as increasing the number of computing cards included in the computing power, so as to further improve the training efficiency of the large model.

[0067] Optionally, the computing power usage status includes the bandwidth for data exchange between the memory and the computing card and the video memory occupancy rate, and step S3 includes: Step S31''': In the case where the video memory occupancy rate does not reach the preset occupancy rate and the bandwidth for data exchange between the memory and the computing card reaches the preset bandwidth, it is determined that the computing power usage status does not match the expected status; Step S32''': Adjust the proportion of the model weight data unloaded to the memory according to the bandwidth for data exchange between the memory and the computing card and the video memory occupancy rate.

[0068] Among them, the video memory occupancy rate can be understood as the usage rate of the computing card and the memory, and the bandwidth for exchanging data between the video memory and the computing card can be understood as the bandwidth for exchanging data between the memory and the computing card, and the above-mentioned data exchange can also be understood as data communication.

[0069] Among them, the model weight data can also be referred to as the model weight coefficient or the model weight. The model weight data can represent the parameters of the neuron connections in each large model. These weights are continuously adjusted during the training process of the large model so that the large model can predict the output more accurately. And the model weight data determines how the input data of the large model is processed and transformed by the model.

[0070] In the present invention, when the video memory occupancy rate does not reach the preset occupancy rate and the bandwidth for exchanging data between the memory and the computing card reaches the preset bandwidth, it indicates that the data exchange between the memory and the computing card is too frequent at this time. Therefore, the proportion of the model weight data unloaded to the memory can be reduced to reduce the communication between the computing card and the memory, thereby improving the training efficiency of the subtasks allocated on the computing card and further improving the training efficiency of the large model.

[0071] It should be noted that the above-mentioned bandwidth for exchanging data between the memory and the computing card can be calculated through the video memory occupancy rate or the communication volume between the computing card and the memory. When the video memory occupancy rate or the communication volume between the computing card and the memory is larger, the bandwidth for exchanging data between the memory and the computing card is larger, and the three can be linearly correlated.

[0072] Among them, the types included in the computing power are not specifically limited herein. Optionally, the computing power can include the corresponding computing power of a central processing unit (CPU), a graphics processing unit (GPU), and a memory, etc.

[0073] It should be noted that the above-mentioned adjustment of the storage level of the training data in the training task, the adjustment of the batch sample quantity of a single computing card, the adjustment of the parallelism of the computing power, and the adjustment of the proportion of the model weight data unloaded to the memory are all for improving the training efficiency of the training task of the large model.

[0074] It should be noted that the process of the training task of the large model can include the following steps: 1. Set training parameters, and the training parameters can include the learning rate of the large model and the number of samples; 2. Prepare training samples, and the above-mentioned training samples can be referred to as samples, training data, or dataloaders; 3. Perform forward propagation on the large model. The above forward propagation can be understood as a process of inputting samples into the large model for iterative training and the large model outputting prediction results. 4. Calculate the loss. Input the prediction results output by the above large model and the actual labels of the samples into the loss function to calculate the loss. The specific type of the above loss function is not limited here. Optionally, the above loss function may include mean square error function, cross - entropy function, etc. 5. Backward propagation. Calculate the gradient of the loss function with respect to the parameters of the large model and update the model parameters according to the above gradient. 6. Verification and adjustment. After the training of the large model with each batch of samples is completed, the validation set of the samples can be used to evaluate the performance of the large model. If the performance of the large model does not improve or starts to decline during the evaluation of the validation set, it may be that the large model is overfitting. At this time, the hyperparameters of the large model can be adjusted, or regularization techniques can be used to adjust the large model. The above regularization techniques may include: dropout, L1 / L2 regularization, or early stopping method, etc.

[0075] See Figure 2 , Figure 2 For the schematic diagram of the training method of the large model, as Figure 2 shown, the training method of the large model may include the following steps: Step S1: Load training data; the training data can be understood as samples or sample data. Step S2: Calculate the loss; the loss can be calculated using a loss function, and the specific calculation method can refer to the above corresponding description. Step S3: Subsequent calculations. The subsequent calculations may include the calculation processes in the above backward propagation, verification and adjustment.

[0076] See Figure 3 , Figure 3 For the schematic structural diagram of a device for adjusting the large - model training task based on the computing power usage status in an intelligent computing center cloud platform provided by the present invention, as Figure 3 shown, the device 300 for adjusting the large - model training task based on the computing power usage status in the intelligent computing center cloud platform includes: An acquisition module 301, configured to acquire the computing power usage status, where the computing power usage status is used to represent the usage status of the computing power in the intelligent computing center cloud platform by the training task during the training of the large model. A determination module 302, configured to determine whether the computing power usage status matches the expected status. An adjustment module 303, configured to adjust the training task of the large model according to the computing power usage status when the computing power usage status does not match the expected status.

[0077] Optionally, the computing power usage status includes at least one of the following: the core utilization rate of the computing card, the bandwidth for data exchange between the memory and the computing card, the video memory throughput, and the throughput between different computing cards.

[0078] Optionally, the computing power usage status includes the core utilization rate of the computing card. The adjustment module 303 includes: A first calculation sub-module, configured to determine that the computing power usage status does not match the expected status when it is calculated that the volatility of the core utilization rate of the computing card within a preset period is greater than the expected volatility; A first adjustment sub-module, configured to adjust the storage level of the training data in the training task according to the fluctuation status of the core utilization rate of the computing card when the computing power usage status does not match the expected status, where the fluctuation status includes the volatility.

[0079] Optionally, the computing power usage status includes the throughput between different computing cards. The adjustment module 303 includes: A second determination sub-module, configured to determine that the computing power usage status does not match the expected status when it is calculated that the throughput between different computing cards is greater than the first throughput; A second adjustment sub-module, configured to adjust the number of batch samples of a single computing card according to the throughput between different computing cards when the computing power usage status does not match the expected status.

[0080] Optionally, the computing power usage status includes the video memory throughput. The adjustment module 303 includes: A third determination sub-module, configured to determine that the computing power usage status does not match the expected status when it is calculated that the video memory throughput is greater than the second throughput; A third adjustment sub-module, configured to adjust the parallelism of the computing power according to the video memory throughput when the computing power usage status does not match the expected status.

[0081] Optionally, the computing power usage status includes the bandwidth for data exchange between the memory and the computing card and the video memory occupancy rate. The adjustment module 303 includes: A fourth determination sub-module, configured to determine that the computing power usage status does not match the expected status when the video memory occupancy rate does not reach the preset occupancy rate and the bandwidth for data exchange between the memory and the computing card reaches the preset bandwidth; A fourth adjustment sub-module, configured to adjust the proportion of the model weight data unloaded to the memory according to the bandwidth for data exchange between the memory and the computing card and the video memory occupancy rate.

[0082] The device 300 for adjusting the large model training task based on the computing power usage status provided by the intelligent computing center cloud platform of the present invention can execute each step in the method for adjusting the large model training task based on the computing power usage status of the intelligent computing center cloud platform, and thus has the same beneficial technical effects as the method for adjusting the large model training task based on the computing power usage status of the intelligent computing center cloud platform, which will not be elaborated herein for the sake of brevity.

[0083] Please refer to Figure 4 , the present invention also provides an electronic device 40, including a processor 41, a memory 42, and a computer program stored on the memory 42 and executable on the processor 41. When the computer program is executed by the processor 41, it realizes each process shown in the method for adjusting the large model training task based on the computing power usage status of the intelligent computing center cloud platform, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0084] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes each process of the method for adjusting the large model training task based on the computing power usage status of the intelligent computing center cloud platform, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here. Among them, the computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0085] The present invention also provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, they realize each process of the method for adjusting the large model training task based on the computing power usage status of the intelligent computing center cloud platform shown above Figure 1 , and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0086] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article, or device. Without further limitations, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article, or device including that element.

[0087] Through the description of the above embodiments, those skilled in the art can clearly understand that the method provided by the above invention can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to enable a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the various methods provided by the present invention.

[0088] The present invention has been described above with reference to the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the purpose of the present invention and the scope protected by the claims, and all of them belong to the protection scope of the present invention.

Claims

1. A method for an intelligent computing center cloud platform to adjust large model training tasks based on the computing power usage status, characterized in that, Including: Step S1: Obtain the computing power usage status, which is used to represent the usage status of the computing power in the intelligent computing center cloud platform by the training task during the training of the large model; Step S2: Determine whether the computing power usage status matches the expected status; Step S3: In the case where the computing power usage status does not match the expected status, adjust the training task of the large model according to the computing power usage status.

2. The method according to claim 1, wherein The computing power usage status includes at least one of the following: the core utilization rate of the computing card, the bandwidth for data exchange between the memory and the computing card, the video memory occupancy rate, the video memory throughput, and the throughput between different computing cards.

3. The method according to claim 2, wherein The computing power usage status includes the core utilization rate of the computing card, and Step S3 includes: Step S31: In the case where the calculated volatility of the core utilization rate of the computing card within a preset period is greater than the expected volatility, determine that the computing power usage status does not match the expected status; Step S32: In the case where the computing power usage status does not match the expected status, adjust the storage level of the training data in the training task according to the fluctuation status of the core utilization rate of the computing card, and the fluctuation status includes the volatility.

4. The method according to claim 2, characterized in that The computing power usage status includes the throughput between different computing cards, and Step S3 includes: Step S31’: In the case where the calculated throughput between different computing cards is greater than the first throughput, determine that the computing power usage status does not match the expected status; Step S32’: In the case where the computing power usage status does not match the expected status, adjust the batch sample number of a single computing card according to the throughput between different computing cards.

5. The method according to claim 2, wherein The computing power usage status includes the video memory throughput, and Step S3 includes: Step S31’’: In the case where the calculated video memory throughput is greater than the second throughput, determine that the computing power usage status does not match the expected status; Step S32’’: In the case where the computing power usage status does not match the expected status, adjust the parallelism of the computing power according to the video memory throughput.

6. The method according to claim 2, wherein The computing power usage status includes the bandwidth for data exchange between the memory and the computing card and the video memory occupancy rate, and Step S3 includes: Step S31’’’: In the case where the video memory occupancy rate does not reach the preset occupancy rate and the bandwidth for data exchange between the memory and the computing card reaches the preset bandwidth, determine that the computing power usage status does not match the expected status; Step S32’’’: Adjust the proportion of the model weight data unloaded to the memory according to the bandwidth for data exchange between the memory and the computing card and the video memory occupancy rate.

7. An apparatus for adjusting large model training tasks based on the computing power usage status in an intelligent computing center cloud platform, characterized in that, Including: An acquisition module for acquiring the computing power usage status, which is used to represent the usage status of the computing power in the intelligent computing center cloud platform by the training task during the training of the large model; A determination module for determining whether the computing power usage status matches the expected status; An adjustment module for adjusting the training task of the large model according to the computing power usage status in the case where the computing power usage status does not match the expected status.

8. An electronic device, characterized in that, Including: A processor, a memory, and a program stored on the memory and executable on the processor, the program, when executed by the processor, implementing the steps of the method for adjusting the large model training task of the intelligent computing center cloud platform according to any one of claims 1 to 6 based on the computing power usage status.

9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the method for adjusting the large model training task of the intelligent computing center cloud platform according to any one of claims 1 to 6 based on the computing power usage status are implemented.

10. A computer program product, characterized in that, It includes computer instructions, and when the computer instructions are executed by a processor, the steps of the method for adjusting the large model training task of the intelligent computing center cloud platform according to any one of claims 1 to 6 based on the computing power usage status are implemented.

Citation Information

Patent Citations

  • Target tracking model training method and device and target tracking method and device

    CN114169425A

  • Large model computing power distribution and scheduling system oriented to edge computing

    CN119166369A

  • Self-adaptive computing power scheduling system for large model training

    CN119322682A

  • Intelligent computing power scheduling method and device for intelligent common computing power computing center

    CN119938340A

  • Model training method and apparatus, system, and storage medium

    WO2024055979A1

Cited By

  • Method and device for automatically detecting injection-writing consistency of automobile drawings by intelligent computing cloud platform through computing power

    CN121392886A