Model asynchronous training method and device of intelligent computing center for providing computing power resources
Through the asynchronous training method, the cache pool and multi-GPU architecture are used to realize asynchronous data preparation and model training, solving the problem of low model training efficiency and improving training speed and resource utilization efficiency.
Patent Information
- Application Number
- CN202510639268.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-19
AI Technical Summary
In the prior art, there is idle computing resources caused by incomplete preparation of training data during model training, resulting in inefficient model training, especially in large-scale parameters.
Using an asynchronous training method, by obtaining multiple batches of original training data for preprocessing and storing it in the cache pool, the GPU uses to retrieve data from the cache pool for model training, and at the same time uses multiple GPUs for evaluation and testing, to realize asynchronous data preparation, model evaluation and training.
This improves the efficiency of model training, avoids training waiting time, optimizes the utilization of computing resources, and significantly improves the training speed under large-scale parameters.
Smart Images

Figure CN120508828A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent computing centers, smart computing centers and computing power infrastructure, and specifically to a model asynchronous training method and device for an intelligent computing center that provides computing power resources. Background Art
[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged.
[0003] An "Intelligent Computing Center" is a facility that uses large-scale heterogeneous computing resources, including general-purpose and intelligent computing power, to provide the computing power, data, and algorithms required for AI applications (such as AI deep learning model development, model training, and model inference). The Intelligent Computing Center encompasses facilities, hardware, and software, and provides a full stack of capabilities, from bottom-level computing power to top-level application enablement.
[0004] “Intelligent Computing Center” includes but is not limited to “Smart Computing Center”.
[0005] "Intelligent Computing Center" refers to an artificial intelligence computing center. It is a type of computing power infrastructure that is based on artificial intelligence theory, adopts artificial intelligence computing architecture, and provides computing power services, data services, and algorithm services required for artificial intelligence applications.
[0006] "Computing power" is the core of "intelligent computing center" and "intelligent computing center". It is the ability of computer equipment or computing / data center to process parameters. It is the ability of computer hardware and software to work together to perform certain computing needs. It is the computing power to achieve target result output by processing parameter data. It is a new type of productivity that integrates parameter computing power, network carrying capacity, and data storage capacity. It mainly provides services to society through computing power infrastructure.
[0007] In intelligent computing centers, GPUs (Graphics Processing Units) are typically used to provide computing resources for model training. In existing technologies, model training typically involves multiple rounds of training, each of which requires training data preparation before the prepared data is used for model training. However, in existing technologies, after model training has processed the training data, the preparation of the next round of training data is not yet complete. Model training needs to wait for the training data to be prepared, resulting in idle computing resources for model training and low model training efficiency. This phenomenon is particularly serious when the model parameters are large.
[0008] It can be seen that the existing technology has the problem of low efficiency of model training. Summary of the Invention
[0009] The present invention provides a model asynchronous training method and device for an intelligent computing center that provides computing resources, so as to solve the problem of low efficiency of model training in the prior art.
[0010] To solve the above problems, the present invention is achieved as follows:
[0011] In a first aspect, the present invention provides a model asynchronous training method for an intelligent computing center that provides computing resources, comprising:
[0012] Step S1: Acquire multiple batches of original training data corresponding to a first working component, where the first working component is a component deployed in a node in the intelligent computing center;
[0013] Step S2: preprocessing the multiple batches of original training data to obtain multiple batches of training data;
[0014] Step S3: storing the plurality of batches of training data in a first buffer pool;
[0015] Step S4: using the first GPU to sequentially retrieve the multiple batches of training data from the first buffer pool to perform multiple rounds of training on the first model in the first working component;
[0016] Among them, obtaining the multiple batches of original training data, preprocessing the multiple batches of original training data, and storing the multiple batches of training data are all performed asynchronously with the multiple rounds of training of the first model.
[0017] In one embodiment, the method further comprises:
[0018] Step S5: When the number of batches of training data stored in the first buffer pool reaches a set number threshold, stop preprocessing the remaining batches of original training data in the multiple batches of original training data;
[0019] Step S6: If the number of batches of training data stored in the first buffer pool does not reach a set number threshold, preprocess the remaining batches of original training data;
[0020] The preprocessing of the remaining batches of original training data is performed asynchronously with the multiple rounds of training of the first model.
[0021] In one embodiment, the method further comprises:
[0022] Step S7: When the first model is trained for N rounds, evaluating and / or testing a first intermediate model using a second GPU of the intelligent computing center, where the first intermediate model is a model obtained after training the first model for N rounds, where N is a positive integer greater than 1;
[0023] The evaluation and / or testing of the first intermediate model is performed asynchronously with the multiple rounds of training of the first model.
[0024] In one embodiment, step S7 includes:
[0025] Step S71: After training the first model for N rounds, transfer the first intermediate model to the video memory of the second GPU via the video memory of the first GPU;
[0026] Step S72: Retrieving evaluation sample data and / or test sample data through the second GPU;
[0027] Step S73: Evaluate the first intermediate model based on the evaluation sample data through the second GPU, and / or test the first intermediate model based on the test sample data through the second GPU.
[0028] In one embodiment, step S72 includes:
[0029] Step S721: Obtain original evaluation data and / or original test data;
[0030] Step S722: pre-process the original evaluation data to obtain the evaluation sample data; and / or pre-process the original test data to obtain the test sample data;
[0031] Step S723: Storing the evaluation sample data in the second buffer pool, and / or storing the test sample data in the third buffer pool;
[0032] The step S73 includes:
[0033] Step S731: Retrieve the evaluation sample data from the second cache pool through the second GPU to evaluate the first intermediate model, and / or retrieve the test sample data from the third cache pool through the second GPU to test the first intermediate model.
[0034] In one embodiment, the method further comprises:
[0035] Step S8: When the first model is trained for M rounds, the CPU of the intelligent computing center saves a second intermediate model to a storage space, where the second intermediate model is a model obtained after training the first model for M rounds, where M is a positive integer greater than 1.
[0036] Saving the second intermediate model to the storage space is performed asynchronously with the multiple rounds of training of the first model.
[0037] In one embodiment, the first working component is a component included in a target node in the intelligent computing center, and the target node includes multiple components. Before step S1, the method further includes:
[0038] Step S0: split the initial model into multiple sub-models, and assign the multiple sub-models to different components, wherein the multiple sub-models include the first model.
[0039] In a second aspect, the present invention further provides a model asynchronous training device for an intelligent computing center that provides computing resources, comprising:
[0040] an acquisition module, configured to acquire multiple batches of original training data corresponding to a first working component, where the first working component is a component deployed in a node in the intelligent computing center;
[0041] A processing module, configured to preprocess the multiple batches of original training data to obtain multiple batches of training data;
[0042] A storage module, configured to store the plurality of batches of training data into a first cache pool;
[0043] A calling module, configured to sequentially call the plurality of batches of training data from the first cache pool through the first GPU to perform multiple rounds of training on the first model in the first working component;
[0044] Wherein, obtaining the multiple batches of original training data, preprocessing the multiple batches of original training data, and storing the multiple batches of training data are all performed asynchronously with the multiple rounds of training of the first model.
[0045] In a third aspect, the present invention also provides an electronic device comprising a processor, a memory, and a computer program stored on the memory and runnable on the processor. When the computer program is executed by the processor, the steps in the asynchronous model training method of the intelligent computing center providing computing resources as described in the first aspect above are implemented.
[0046] In a fourth aspect, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the asynchronous model training method of an intelligent computing center that provides computing power resources as described in the first aspect above.
[0047] In a fifth aspect, the present invention also provides a computer program product comprising computer instructions, which, when executed by a processor, implement the steps in the model asynchronous training method of the intelligent computing center that provides computing power resources as described in the first aspect above.
[0048] In the present invention, multiple batches of original training data corresponding to a first working component are obtained, and the first working component is a component deployed in a node in the intelligent computing center; the multiple batches of original training data are preprocessed to obtain multiple batches of training data; the multiple batches of training data are stored in a first cache pool; the multiple batches of training data are sequentially retrieved from the first cache pool by a first GPU to perform multiple rounds of training on the first model in the first working component; wherein, the acquisition of the multiple batches of original training data, the preprocessing of the multiple batches of original training data, and the storage of the multiple batches of training data are all performed asynchronously with the multiple rounds of training of the first model. In this way, after the model has trained a batch of training data, there is no need to wait, and the preprocessed training data can be directly retrieved from the first cache pool, thereby achieving the goal of acquiring multiple batches of original training data, preprocessing multiple batches of original training data, and storing multiple batches of training data asynchronously with the multiple rounds of training of the first model, thereby avoiding the problem of waiting for model training, and greatly improving the efficiency of model training. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solution of the present invention, the following is a brief introduction to the drawings required for the description of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0050] Figure 1 This is a flow chart of a method for asynchronous training of a model in an intelligent computing center providing computing resources provided by the present invention;
[0051] Figure 2 It is a structural diagram of the first working component provided by the present invention;
[0052] Figure 3 It is a schematic diagram of a node of the intelligent computing center provided by the present invention;
[0053] Figure 4 This is a structural diagram of a model asynchronous training device for an intelligent computing center that provides computing resources provided by the present invention;
[0054] Figure 5 This is a structural diagram of an electronic device provided by the present invention. DETAILED DESCRIPTION
[0055] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0056] The "computing power" mentioned in the present invention refers to: the ability of computer equipment or computing / data centers to process information, the ability of computer hardware and software to work together to execute certain computing requirements, and the computing power to achieve target result output by processing information data. It is a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.
[0057] The "computing power" (Computational Power, CP) mentioned in the present invention refers to: the ability of a data center server to process data and output results. It is a comprehensive indicator to measure the computing power of a data center, including general computing power, super computing power and intelligent computing power. The commonly used unit of measurement is the number of floating-point operations performed per second (FLOPS, 1EFLOPS=10^18FLOPS). The larger the value, the stronger the comprehensive computing power. According to calculations, 1EFLOPS is approximately the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream notebooks. The calculation formula is: CP=CP 通用 +CP 智能 +CP 超级 .
[0058] The "carrying capacity" (Network Power, NP) mentioned in the present invention refers to: it is the performance of the data transmission capability of the computing power facility, which includes comprehensive capabilities such as network architecture, network bandwidth, transmission latency, intelligent management and scheduling, etc. It involves network transmission within and between data centers, and is a comprehensive indicator for measuring network transmission scheduling capabilities.
[0059] The "Storage Power" (SP) described in this invention refers to the comprehensive capabilities of a data center in terms of data storage capacity, performance, security and reliability, and environmental friendliness. It is a comprehensive indicator for measuring a data center's data storage capacity, encompassing both external storage devices such as storage arrays and internal server storage. Storage capacity is commonly measured in exabytes (EB, 1EB = 2^60 bytes), while performance is commonly measured in IOPS / TB (Input / Output Operations Per Second / TB). Disaster recovery ratio is a key indicator of security and reliability.
[0060] The "computing power infrastructure" mentioned in the present invention refers to a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity, and can realize the centralized calculation, storage, transmission and application of information.
[0061] The "new information infrastructure" mentioned in the present invention refers to: mainly including network infrastructure such as 5G networks, fiber-optic broadband networks, backbone networks, international communication networks, satellite Internet, computing power infrastructure such as data centers, general computing power centers, intelligent computing centers, supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing.
[0062] The "computing power" mentioned in the present invention includes: general computing power, intelligent computing power and super computing power.
[0063] The "general computing power" mentioned in the present invention refers to the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.
[0064] The "intelligent computing power" mentioned in this invention refers to: a computing platform based on specialized chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit) for various innovative artificial intelligence applications, such as natural language processing and machine vision.
[0065] The "supercomputing power" mentioned in the present invention refers to the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and uses a dedicated operating system to handle extremely complex or data-intensive problems. It is mainly used for calculations in cutting-edge scientific fields, such as planetary simulation, drug molecule design, genetic analysis, etc.
[0066] The "intelligent computing center" described in this article refers to a facility that provides the computing power, data, and algorithms required for artificial intelligence applications (such as AI deep learning model development, model training, and model inference) by utilizing large-scale heterogeneous computing resources, including general-purpose computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center encompasses facilities, hardware, and software, and can provide a full stack of capabilities, from bottom-level computing power to top-level application enablement.
[0067] The "intelligent computing center" mentioned in the present invention includes but is not limited to the "intelligent computing center".
[0068] The "intelligent computing center" mentioned in the present invention is an artificial intelligence computing center, which is a type of computing power infrastructure based on artificial intelligence theory, adopts artificial intelligence computing architecture, and provides computing power services, data services and algorithm services required for artificial intelligence applications.
[0069] The "computing power center" mentioned in the present invention refers to: a facility that is mainly composed of infrastructure such as wind, fire, water, electricity, and IT hardware and software equipment, and has computing power, transportation capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.
[0070] The "supercomputing center" mentioned in the present invention refers to: a supercomputing data center, which is a data center based on a supercomputer or a large-scale computing cluster, which can provide large-scale computing, storage and network services and other functions, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling and genome sequencing.
[0071] The "computing resources" mentioned in the present invention refer to: technologies and facilities with information computing, transmission, storage and application capabilities required for the development of a digital society, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and supporting and guarantee resources such as wind, fire, water and electricity.
[0072] The “model” mentioned in the present invention includes but is not limited to a “large language model” and a “multimodal large model”.
[0073] The "large language model" mentioned in the present invention refers to a large language model (LLM), which is a language model with a large parameter scale. It is designed to understand and generate human language. It is trained with a large amount of text data and can perform a wide range of tasks including text summarization, translation, sentiment analysis, etc.
[0074] The "Multimodal Large Models" mentioned in the present invention refer to models that combine multimodal information such as text, images, video, and audio for training, including but not limited to multimodal large language models.
[0075] The "asynchronous training" mentioned in the present invention means that during the model training process, at least one of the data preparation link, model testing link, model evaluation link and model preservation link is carried out asynchronously with the model training link, so that the model training link is not affected by the data preparation link, model testing link, model evaluation link, model preservation link and other links, and more computing power resources can be used in the model training link, thereby improving the efficiency of model training.
[0076] See Figure 1 , Figure 1This is a flowchart of a method for asynchronous training of a model of an intelligent computing center providing computing resources provided by the present invention. Figure 1 As shown, the following steps are included:
[0077] Step S1: Acquire multiple batches of original training data corresponding to a first working component, where the first working component is a component deployed in a node in the intelligent computing center.
[0078] Step S2: preprocess the multiple batches of original training data to obtain multiple batches of training data.
[0079] The aforementioned first working component is a component deployed in a node of the intelligent computing center and is used to train the model using the computing resources provided by the intelligent computing center. It should be noted that the intelligent computing center can have multiple nodes, and each node can have multiple working components, thereby enabling distributed model training or training of different models.
[0080] Before the first working component performs model training, environment information for model training is pre-created in the first working component, so that the first working component can run the model in the configured environment information and train the model.
[0081] The above original training data is used to train the model in the first working component. It should be noted that the training data used for model training is the data pre-processed from the original training data. The training data is obtained by processing the original training data, and then the model is trained using the training data.
[0082] Step S3: storing the plurality of batches of training data in a first buffer pool;
[0083] It should be noted that before model training, multiple batches of original training data need to be preprocessed to obtain preprocessed training data, which can then be used for model training. Figure 2 As shown, in the prior art, a first working component simultaneously preprocesses the raw training data and performs model training. That is, while the model is being trained on a batch of training data, the unprocessed raw training data is also preprocessed. Specifically, a batch of raw training data is preprocessed to obtain a batch of training data, which is then retrieved by the model training. After the model training retrieves the batch of training data, the raw training data is preprocessed again.
[0084] In the present invention, the original training data can be continuously preprocessed to obtain multiple batches of training data and cache them in the first cache pool. After the model has trained a batch of training data, there is no need to wait, and the preprocessed training data can be directly retrieved from the first cache pool, thereby avoiding the problem of waiting for model training and greatly improving the efficiency of model training.
[0085] Furthermore, a workflow of training data may be set in the first working component, and a first buffer pool may be set in the workflow of training data, so that the workflow of training data can be used to pre-process multiple batches of original training data.
[0086] In some implementations, all batches of original training data may be stored in the first cache pool after preprocessing.
[0087] Step S4: using the first GPU to sequentially retrieve the multiple batches of training data from the first buffer pool to perform multiple rounds of training on the first model in the first working component;
[0088] Wherein, obtaining the multiple batches of original training data, preprocessing the multiple batches of original training data, and storing the multiple batches of training data are all performed asynchronously with the multiple rounds of training of the first model.
[0089] The first GPU described above is a GPU deployed in an intelligent computing center. It is understood that multiple GPUs can be deployed in an intelligent computing center, and different working components can use the computing power resources provided by different GPUs for model training. In the present invention, the first working component uses the computing power resources provided by the first GPU to train the first model.
[0090] The model training process requires multiple rounds of training. Each round of training requires the input of at least one batch of training data. After a certain number of rounds of model training, the model is saved, evaluated and tested to determine the effect of the model training.
[0091] In the present invention, multiple batches of original training data corresponding to a first working component are obtained, and the first working component is a component deployed in a node in the intelligent computing center; the multiple batches of original training data are preprocessed to obtain multiple batches of training data; the multiple batches of training data are stored in a first cache pool; the multiple batches of training data are sequentially retrieved from the first cache pool by a first GPU to perform multiple rounds of training on the first model in the first working component; wherein, the acquisition of the multiple batches of original training data, the preprocessing of the multiple batches of original training data, and the storage of the multiple batches of training data are all performed asynchronously with the multiple rounds of training of the first model. In this way, after the model has trained a batch of training data, there is no need to wait, and the preprocessed training data can be directly retrieved from the first cache pool, thereby achieving the goal of acquiring multiple batches of original training data, preprocessing multiple batches of original training data, and storing multiple batches of training data asynchronously with the multiple rounds of training of the first model, thereby avoiding the problem of waiting for model training, and greatly improving the efficiency of model training.
[0092] In some implementations, part of the original training data may be pre-processed and then stored in the first buffer pool. Specifically, in one embodiment, the method further includes:
[0093] Step S5: When the number of batches of training data stored in the first buffer pool reaches a set number threshold, stop preprocessing the remaining batches of original training data in the multiple batches of original training data;
[0094] Step S6: If the number of batches of training data stored in the first buffer pool does not reach a set number threshold, preprocess the remaining batches of original training data;
[0095] Among them, stopping the preprocessing of the remaining batches of original training data is asynchronously performed with the multiple rounds of training of the first model, and preprocessing the remaining batches of original training data is asynchronously performed with the multiple rounds of training of the first model.
[0096] It should be noted that the first cache pool still requires some computing resources. If all batches of raw training data are preprocessed and stored in the first cache pool, the first cache pool will occupy too much computing resources. Step S15 ensures that a certain number of batches of training data are continuously stored in the first cache pool. This ensures that the first cache pool consumes less computing resources while eliminating the need for model training to wait, thereby improving model training efficiency.
[0097] In the present invention, if the number of batches of training data stored in the first buffer pool reaches a set threshold, preprocessing of the remaining batches of original training data in the multiple batches of original training data is stopped; if the number of batches of training data stored in the first buffer pool does not reach the set threshold, preprocessing of the remaining batches of original training data is performed. This eliminates the need to preprocess all the original training data at once. The first buffer pool can be configured to retain fewer batches of training data, allowing model training and original data preprocessing to proceed asynchronously. This eliminates the need for model training to wait while the first buffer pool consumes less computing resources, thereby improving model training efficiency.
[0098] In some embodiments, in addition to the above-mentioned step S6, after step S5, pre-processing may be performed by the following steps:
[0099] Step S6′: When all the training data stored in the first buffer pool are retrieved, pre-process the remaining batches of original training data.
[0100] In one embodiment, the method further comprises:
[0101] Step S7: When the first model is trained for N rounds, evaluating and / or testing a first intermediate model using a second GPU of the intelligent computing center, where the first intermediate model is a model obtained after training the first model for N rounds, where N is a positive integer greater than 1;
[0102] The evaluation and / or testing of the first intermediate model is performed asynchronously with the multiple rounds of training of the first model.
[0103] It should be noted that in the existing technology, data preparation, model evaluation and model testing are usually required during the model training process. Data preparation may cause the GPU to wait, while model evaluation and model testing require part of the GPU's computing power resources, resulting in a reduction in computing power resources used for model training, making the model training efficiency very low. This phenomenon is particularly serious when data processing is heavy and the model parameters are large.
[0104] To solve the above problems, in the present invention, model training is performed by the first GPU, and model evaluation and / or testing is performed by the second GPU, so that the evaluation and / or testing of the first intermediate model is performed asynchronously with the multiple rounds of training of the first model. This can avoid the situation where the first GPU provides part of the computing power resources for model evaluation and / or testing. The first GPU only needs to provide computing power resources for model training, thereby greatly improving the efficiency of model training.
[0105] The above-mentioned second GPU is a GPU in addition to the first GPU in the intelligent computing center. The second GPU is used to perform evaluation and / or testing during the model training process, so that the first GPU only needs to bear the model training task, and does not need to bear the evaluation task and / or testing task, thereby realizing the asynchronous progress of model evaluation and / or testing and model training. Compared with using the first GPU to bear the model training task, evaluation task and testing task at the same time, the efficiency of model training can be greatly improved.
[0106] N can be set as needed. The number of training rounds required for model evaluation on the second GPU can be the same as or different from the number of training rounds required for model testing on the second GPU. For example, model evaluation and model testing can be performed after every 10 training rounds; or, model evaluation can be performed after every 5 training rounds and model testing can be performed after every 8 training rounds.
[0107] The step S7 comprises:
[0108] Step S71: After training the first model for N rounds, transfer the first intermediate model to the video memory of the second GPU via the video memory of the first GPU;
[0109] Step S72: Retrieving evaluation sample data and / or test sample data through the second GPU;
[0110] Step S73: Evaluate the first intermediate model based on the evaluation sample data through the second GPU, and / or test the first intermediate model based on the test sample data through the second GPU.
[0111] In the present invention, when the first model is trained for N rounds, the first intermediate model is transferred to the video memory of the second GPU via the video memory of the first GPU; evaluation sample data and / or test sample data are retrieved via the second GPU; the first intermediate model is evaluated based on the evaluation sample data via the second GPU, and / or the first intermediate model is tested based on the test sample data via the second GPU. Step S72 is performed asynchronously with step S73; step S73 is performed asynchronously with model training. In this way, the first intermediate model is quickly transferred via the video memory of the first GPU and the video memory of the second GPU, and the first intermediate model can then be evaluated and / or tested via the second GPU.
[0112] In one embodiment, step S72 includes:
[0113] Step S721: Obtain original evaluation data and / or original test data;
[0114] Step S722: pre-process the original evaluation data to obtain the evaluation sample data; and / or pre-process the original test data to obtain the test sample data;
[0115] Step S723: Storing the evaluation sample data in the second buffer pool, and / or storing the test sample data in the third buffer pool;
[0116] The step S73 includes:
[0117] Step S731: Retrieve the evaluation sample data from the second cache pool through the second GPU to evaluate the first intermediate model, and / or retrieve the test sample data from the third cache pool through the second GPU to test the first intermediate model.
[0118] It should be noted that the evaluation sample data for model evaluation and the test sample data for model testing also need to be preprocessed before they can be evaluated and / or tested. Similar to model training, in the present invention, original evaluation data and / or original test data are obtained; the original evaluation data are preprocessed to obtain the evaluation sample data; and / or the original test data are preprocessed to obtain the test sample data; the evaluation sample data are stored in the second buffer pool, and / or the test sample data are stored in the third buffer pool. Step S322 is performed asynchronously with step S331, the evaluation data preprocessing is performed asynchronously with the model evaluation, and the test data preprocessing is performed asynchronously with the model testing. In this way, when evaluating the model, there is no need to wait for the original evaluation data to be preprocessed, but the evaluation sample data can be directly obtained from the second buffer pool for model evaluation, which greatly improves the efficiency of model evaluation; similarly, when testing the model, there is no need to wait for the original test data to be preprocessed, but the test sample data can be directly obtained from the third buffer pool for model testing, which greatly improves the efficiency of model testing.
[0119] In some implementations, a workflow for evaluating data may be set within a first working component, and a second buffer pool may be set within the workflow for evaluating data, so that multiple batches of original evaluation data can be preprocessed through the workflow for evaluating data.
[0120] In some implementations, a workflow for test data may be set within the first working component, and a third buffer pool may be set within the workflow for test data, so that multiple batches of original test data can be preprocessed through the workflow for test data.
[0121] In one embodiment, the method further comprises:
[0122] Step S4: When the first model is trained for M rounds, the CPU of the intelligent computing center saves a second intermediate model to a storage space, where the second intermediate model is a model obtained after training the first model for M rounds, where M is a positive integer greater than 1.
[0123] Saving the second intermediate model to the storage space is performed asynchronously with the multiple rounds of training of the first model.
[0124] In the present invention, when the first model is trained for M rounds, the second intermediate model is saved to the storage space by the CPU of the intelligent computing center. The second intermediate model is the model obtained after training the first model for M rounds, where M is a positive integer greater than 1. In this way, the second intermediate model is saved to the storage space by the CPU of the intelligent computing center, and the saving of the second intermediate model is performed asynchronously with the training of the first model. There is no need to save the second intermediate model to the storage space by the first GPU, which reduces the computing power resource occupation of the first GPU and further improves the efficiency of model training.
[0125] In one embodiment, the first working component is a component included in a target node in the intelligent computing center, and the target node includes multiple components. Before step S1, the method further includes:
[0126] Step S0: split the initial model into multiple sub-models, and assign the multiple sub-models to different components, wherein the multiple sub-models include the first model.
[0127] It is understandable that if Figure 3 As shown, the intelligent computing center can have multiple nodes, and each node can have multiple work components, thereby enabling distributed training of models or training of different models. For example, the intelligent computing center has nodes 1 and 2. Work components 1, 2, 3, and 4 are deployed in node 1, and work components 5, 6, 7, and 8 are deployed in node 2. Different work components can work together.
[0128] In the present invention, the initial model is split into multiple sub-models, and the multiple sub-models are assigned to different components. The multiple sub-models include the first model. In this way, the initial model can be distributedly trained through different working components, thereby greatly improving the efficiency of model training.
[0129] See Figure 4 , Figure 4 This is a structural diagram of a model asynchronous training device for an intelligent computing center that provides computing resources provided by the present invention, such as Figure 4As shown, the model asynchronous training device 400 of the intelligent computing center providing computing resources includes:
[0130] An acquisition module 401 is configured to acquire multiple batches of original training data corresponding to a first working component, where the first working component is a component deployed in a node in the intelligent computing center;
[0131] A processing module 402 is configured to preprocess the multiple batches of original training data to obtain multiple batches of training data;
[0132] Storage module 403, configured to store the plurality of batches of training data into a first buffer pool;
[0133] A calling module 404 is configured to sequentially call the plurality of batches of training data from the first buffer pool through the first GPU to perform multiple rounds of training on the first model in the first working component;
[0134] Wherein, obtaining the multiple batches of original training data, preprocessing the multiple batches of original training data, and storing the multiple batches of training data are all performed asynchronously with the multiple rounds of training of the first model.
[0135] In one embodiment, the model asynchronous training device 400 of the intelligent computing center providing computing resources further includes:
[0136] a stopping module, configured to stop preprocessing the remaining batches of original training data in the plurality of batches of original training data when the number of batches of training data stored in the first buffer pool reaches a set number threshold;
[0137] a continuing module, configured to pre-process the remaining batches of original training data when the number of batches of training data stored in the first buffer pool does not reach a set number threshold;
[0138] The preprocessing of the remaining batches of original training data is performed asynchronously with the multiple rounds of training of the first model.
[0139] In one embodiment, the model asynchronous training device 400 of the intelligent computing center providing computing resources further includes:
[0140] an evaluation and testing module, configured to evaluate and / or test a first intermediate model using a second GPU of the intelligent computing center when the first model is trained for N rounds, where the first intermediate model is a model obtained after training the first model for N rounds, where N is a positive integer greater than 1;
[0141] The evaluation and / or testing of the first intermediate model is performed asynchronously with the multiple rounds of training of the first model.
[0142] In one embodiment, the evaluation test module further includes:
[0143] a transmitting unit, configured to transmit the first intermediate model to the video memory of the second GPU through the video memory of the first GPU when the first model is trained for N rounds;
[0144] a retrieving unit, configured to retrieve evaluation sample data and / or test sample data through the second GPU;
[0145] A processing unit is configured to evaluate the first intermediate model based on the evaluation sample data via the second GPU, and / or to test the first intermediate model based on the test sample data via the second GPU.
[0146] In one embodiment, the calling unit includes:
[0147] an acquisition subunit, configured to acquire original evaluation data and / or original test data;
[0148] a preprocessing subunit, configured to preprocess the original evaluation data to obtain the evaluation sample data; and / or preprocess the original test data to obtain the test sample data;
[0149] a cache subunit, configured to store the evaluation sample data in a second cache pool, and / or store the test sample data in a third cache pool;
[0150] The processing unit includes:
[0151] A processing sub-unit is used to retrieve the evaluation sample data from the second cache pool through the second GPU to evaluate the first intermediate model, and / or to retrieve the test sample data from the third cache pool through the second GPU to test the first intermediate model.
[0152] In one embodiment, the model asynchronous training device 400 of the intelligent computing center providing computing resources further includes:
[0153] a saving module, configured to save, by the CPU of the intelligent computing center, a second intermediate model to a storage space when the first model is trained for M rounds, where the second intermediate model is a model obtained after training the first model for M rounds, where M is a positive integer greater than 1;
[0154] Saving the second intermediate model to the storage space is performed asynchronously with the multiple rounds of training of the first model.
[0155] In one embodiment, the model asynchronous training device 400 of the intelligent computing center providing computing resources further includes:
[0156] An allocation module is used to split the initial model into multiple sub-models and allocate the multiple sub-models to different components, wherein the multiple sub-models include the first model.
[0157] The model asynchronous training device for the intelligent computing center that provides computing power resources provided by the present invention is capable of implementing the various processes of each embodiment of the above-mentioned model asynchronous training method for the intelligent computing center that provides computing power resources. The technical features correspond one to one and can achieve the same technical effects. To avoid repetition, they will not be described here.
[0158] It should be noted that the model asynchronous training device of the intelligent computing center that provides computing resources in the present invention can be a device, or a component, integrated circuit, or chip in an electronic device.
[0159] The present invention also provides an electronic device, see Figure 5 , Figure 5 This is a schematic diagram of the structure of an electronic device provided by the present invention. The electronic device includes a memory 501, a processor 502, and a program or instruction stored in the memory 501 and running on the memory 501. When the program or instruction is executed by the processor 502, Figure 1 Any steps in the corresponding embodiment of the model asynchronous training method of the intelligent computing center that provides computing resources and the same beneficial effects are achieved will not be repeated here.
[0160] The processor 502 may be a CPU, an ASIC, an FPGA, or a GPU.
[0161] Those skilled in the art will appreciate that all or part of the steps of the embodiment of the asynchronous model training method for an intelligent computing center providing computing resources can be accomplished through hardware related to program instructions, and the program can be stored in a readable medium.
[0162] The present invention also provides a readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above Figure 1 The corresponding steps in the embodiment of the asynchronous training method of the model of the intelligent computing center providing computing resources can achieve the same technical effect. To avoid repetition, they are not repeated here. The storage medium is such as a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk.
[0163] The present invention also provides a computer program product, comprising computer instructions, which, when executed by a processor, implement the above Figure 1 The various processes of the embodiment of the model asynchronous training method of the intelligent computing center that provides computing resources correspond to each other and can achieve the same technical effect. In order to avoid repetition, they will not be repeated here.
[0164] The terms "first", "second" and the like in the present invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. In addition, the terms "comprise" and "have" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or that are inherent to these processes, methods, products or devices. In addition, "and / or" is used in this application to represent at least one of the connected objects, for example A and / or B and / or C, which means comprising seven situations including single A, single B, single C, and both A and B exist, both B and C exist, both A and C exist, and both A, B and C exist.
[0165] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0166] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or second terminal device, etc.) to execute the methods of each embodiment of the present application.
[0167] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.
Claims
1. A model asynchronous training method for an intelligent computing center that provides computing resources, characterized in that: include: Step S1: Acquire multiple batches of original training data corresponding to a first working component, where the first working component is a component deployed in a node in the intelligent computing center; Step S2: preprocessing the multiple batches of original training data to obtain multiple batches of training data; Step S3: storing the plurality of batches of training data in a first buffer pool; Step S4: using the first GPU to sequentially retrieve the multiple batches of training data from the first buffer pool to perform multiple rounds of training on the first model in the first working component; Among them, obtaining the multiple batches of original training data, preprocessing the multiple batches of original training data, and storing the multiple batches of training data are all performed asynchronously with the multiple rounds of training of the first model.
2. The method according to claim 1, wherein The method further comprises: Step S5: When the number of batches of training data stored in the first buffer pool reaches a set number threshold, stop preprocessing the remaining batches of original training data in the multiple batches of original training data; Step S6: If the number of batches of training data stored in the first buffer pool does not reach a set number threshold, preprocess the remaining batches of original training data; The preprocessing of the remaining batches of original training data is performed asynchronously with the multiple rounds of training of the first model.
3. The method according to claim 1 or 2, wherein: The method further comprises: Step S7: When the first model is trained for N rounds, evaluating and / or testing a first intermediate model using a second GPU of the intelligent computing center, where the first intermediate model is a model obtained after training the first model for N rounds, where N is a positive integer greater than 1; The evaluation and / or testing of the first intermediate model is performed asynchronously with the multiple rounds of training of the first model.
4. The method according to claim 3, wherein The step S7 comprises: Step S71: After training the first model for N rounds, transfer the first intermediate model to the video memory of the second GPU via the video memory of the first GPU; Step S72: Retrieving evaluation sample data and / or test sample data through the second GPU; Step S73: Evaluate the first intermediate model based on the evaluation sample data through the second GPU, and / or test the first intermediate model based on the test sample data through the second GPU.
5. The method according to claim 4, wherein The step S72 includes: Step S721: Obtain original evaluation data and / or original test data; Step S722: pre-process the original evaluation data to obtain the evaluation sample data; and / or pre-process the original test data to obtain the test sample data; Step S723: Storing the evaluation sample data in the second buffer pool, and / or storing the test sample data in the third buffer pool; The step S73 includes: Step S731: Retrieve the evaluation sample data from the second cache pool through the second GPU to evaluate the first intermediate model, and / or retrieve the test sample data from the third cache pool through the second GPU to test the first intermediate model.
6. The method according to claim 1 or 2, wherein: The method further comprises: Step S8: When the first model is trained for M rounds, the CPU of the intelligent computing center saves a second intermediate model to a storage space, where the second intermediate model is a model obtained after training the first model for M rounds, where M is a positive integer greater than 1. Saving the second intermediate model to the storage space is performed asynchronously with the multiple rounds of training of the first model.
7. The method according to claim 1 or 2, wherein: The first working component is a component included in the target node in the intelligent computing center, and the target node includes multiple components. Before step S1, the method further includes: Step S0: split the initial model into multiple sub-models, and assign the multiple sub-models to different components, wherein the multiple sub-models include the first model.
8. A model asynchronous training device for an intelligent computing center providing computing resources, characterized in that: include: an acquisition module, configured to acquire multiple batches of original training data corresponding to a first working component, where the first working component is a component deployed in a node in the intelligent computing center; A processing module, configured to preprocess the multiple batches of original training data to obtain multiple batches of training data; A storage module, configured to store the plurality of batches of training data into a first cache pool; A calling module, configured to sequentially call the plurality of batches of training data from the first cache pool through the first GPU to perform multiple rounds of training on the first model in the first working component; Among them, obtaining the multiple batches of original training data, preprocessing the multiple batches of original training data, and storing the multiple batches of training data are all performed asynchronously with the multiple rounds of training of the first model.
9. An electronic device, characterized in that: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, the steps of the method for asynchronous training of a model of an intelligent computing center providing computing resources as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the model asynchronous training method of the intelligent computing center providing computing power resources as described in any one of claims 1 to 7.
11. A computer program product, characterized in that It includes computer instructions, which, when executed by a processor, implement the steps of the model asynchronous training method of the intelligent computing center providing computing power resources as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Distributed model training system, method and device, equipment and storage medium
CN112508188A
Visual model training method, vehicle identification method and device
CN113177497A
Training method and device of iron tower steel index prediction model and readable storage medium
CN113468816A
Training method, training device, equipment, system and storage medium
CN114492834A
Model training method and device
CN114818863A