Model training method, data processing method, equipment and program product
By deploying training and inference frameworks on different nodes and configuring communication domains to achieve communication between them, the computing power shortage in the model training and inference stages is solved, parallel computing is realized, and the data processing speed and performance of the model are improved.
Patent Information
- Application Number
- CN202510713889.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, due to the limitations of computing power resources, the model training and inference stages need to reduce the scale of the model or shorten the training time, resulting in a decline in model performance and the large-scale data processing capability cannot be effectively utilized.
The training framework and inference framework of the pre-trained model are deployed on different nodes, and communication between them is realized by configuring the communication domain. The training framework sends the model weight parameters to the inference framework during the training process, and the inference framework is updated and verified to realize parallel computing of training and inference.
Accelerate the data processing speed of model training and inference within a limited training time, alleviate the problem of computing power shortage, ensure model performance, support the interaction of isomorphic and heterogeneous devices, and improve overall computing efficiency.
Smart Images

Figure CN120258094A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly relates to a model training method, a data processing method, a device, and a program product. Background Art
[0002] With the continuous deepening of the application of neural network models, the model parameters have increased exponentially. The number of parameters of the latest generation of large models has exceeded one trillion, and computing power resources have become the bottleneck restricting model technology. The computing power required for training many models doubles every few months. Moreover, training a basic large model often requires thousands of high-end GPUs (Graphics Processing Units) to run continuously for several weeks, and the energy consumption cost is as high as millions of dollars. In related technologies, limited by computing power resources, developers can only reduce the model scale or shorten the training duration, but this will sacrifice the model performance and reduce the data processing ability of the model. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to provide a model training method, a data processing method, a device, and a program product, which can accelerate the data processing speed of model training and inference, alleviate the problem of computing power shortage in the training and inference stages, and ensure the performance of the model within a limited training duration. The specific solutions are as follows: In the first aspect, the present application discloses a model training method. The training framework and the inference framework of the pre-trained model are deployed on different nodes, and the communication domain between the training framework and the inference framework has been configured, including: The training framework performs training according to the training data, and sends the model weight parameters of the training framework to the inference framework through the communication domain during the training process; The inference framework updates according to the model weight parameters, performs inference verification after the update, and sends the inference result to the training framework; Based on the final outputs of the training framework and the inference framework, a trained model is obtained.
[0004] In the second aspect, the present application discloses a data processing method, which is applied to the trained model described above, including: Obtain the data to be processed, and input the data to be processed into the trained model; According to the output of the trained model, obtain the processing result corresponding to the data to be processed.
[0005] In the third aspect, the present application discloses an electronic device, including: A memory for storing a computer program; A processor for executing the computer program to implement the model training method described above, or the data processing method described above.
[0006] Fourthly, the present application discloses a computer-readable storage medium for storing a computer program; wherein when the computer program is executed by a processor, the foregoing model training method or the foregoing data processing method is implemented.
[0007] Fifthly, the present application discloses a computer program product, including a computer program, which implements the foregoing model training method or the foregoing data processing method when executed by a processor.
[0008] In the present application, the training framework and the inference framework of the pre-trained model are deployed on different nodes, and a communication domain between the training framework and the inference framework has been configured. The training framework performs training according to training data, and sends the model weight parameters of the training framework to the inference framework through the communication domain during the training process; the inference framework is updated according to the model weight parameters, performs inference verification after the update, and sends the inference result to the training framework; based on the final outputs of the training framework and the inference framework, a trained model is obtained.
[0009] By separately deploying the training framework and the inference framework on different nodes and realizing the communication between the training framework and the inference framework by configuring the communication domain, the training framework can send the model weight parameters to the inference framework during the training process so that the inference framework can perform inference verification, realizing parallel computing of training and inference, accelerating the data processing speed during model training and inference, alleviating the problem of shortage of computing power in the training and inference stages, and being able to ensure the performance of the model within a limited training duration. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings according to the provided drawings without creative efforts.
[0011] Figure 1 It is a flowchart of a model training method provided by the present application; Figure 2 It is a flowchart of a specific model training method provided by the present application; Figure 3 It is a flowchart of a specific model training method provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0012] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0013] In the related art, limited by computing power resources, developers can only reduce the model size or shorten the training duration, but this will sacrifice the model performance and reduce the data processing ability of the model. To overcome the above technical problems, this application proposes a model training method, which can accelerate the data processing speed of model training and inference, alleviate the problem of computing power shortage in the training and inference stages, and ensure the performance of the model within a limited training duration.
[0014] An embodiment of this application discloses a model training method. The training framework and inference framework of the pre-trained model are deployed on different nodes, and the communication domain between the above training framework and the above inference framework has been configured. See Figure 1 As shown, the method may include the following steps: Step S11: The training framework performs training according to the training data, and sends the model weight parameters of the training framework to the inference framework through the communication domain during the training process.
[0015] The pre-trained model is a neural network model, specifically it can be a large model. A large model refers to a machine learning model with a very large number of parameters, and a large model can process and analyze large-scale data sets. The large model can be a large language model, a vision large model, a multi-modal model, or a basic science large model, etc. The training framework is a tool and library for constructing, training, and optimizing machine learning or deep learning models. These frameworks provide rich APIs (Application Programming Interfaces), optimizers, loss functions, and data processing tools to help developers efficiently design and train models. The inference framework is a tool and library for deploying the trained model into a production environment for actual prediction or decision-making. These frameworks usually optimize the model loading, inference speed, and resource occupancy to meet the requirements of real-time performance and efficiency.
[0016] The above training framework can use efficient training frameworks such as DeepSpeed and Megatron, but the inference code needs to be deleted. This embodiment does not limit the specific framework. The inference framework can use inference frameworks such as vLLM (Virtual Large Language Model), and this embodiment does not limit the specific framework.
[0017] In this embodiment, the training framework and the inference framework of the pre-trained model are deployed on different computing nodes. Since inference verification is based on training, the inference framework needs to use the model weights of the training framework. To enable communication between the training framework and the inference framework, a communication domain between the training framework and the inference framework needs to be pre-constructed to transmit the weight parameters of the model in the training framework to the inference framework service. The inference framework and the training framework use the same model weight parameters to ensure the correctness of the reverse update process of the training framework; the computing nodes of the inference framework can be dynamically expanded, such as using multiple computing nodes to accelerate the inference process.
[0018] Specifically, the configuration of the communication domain between the training framework and the above-mentioned inference framework includes: constructing the communication domain between the training framework and the inference framework according to the communication domain configuration; the communication domain is configured according to the correspondence between the acceleration cards in the training framework (i.e., AI (Artificial Intelligence) acceleration cards) and the acceleration cards in the inference framework, and is determined by taking the acceleration cards in the training framework as the dimension; communication is supported between the same communication domains, and communication is not possible between different communication domains. It can be understood that, for example, the training framework uses 8 cards (acceleration card 0 - acceleration card 7), and each card contains part of the model weight parameters, and the inference framework uses 2 cards (acceleration card a, acceleration card b); since the training framework and the inference framework have inconsistent splitting methods for splitting and storing the weight parameters in the cards, there is a storage correspondence between the acceleration cards. For example, acceleration card a and acceleration card b need to use the weight parameters stored on acceleration card 0, so a communication domain between acceleration card 0 and acceleration card a and acceleration card b is constructed.
[0019] The communication domain configuration specifically includes the training framework node address (such as the training framework node IP), the inference framework node IP, the acceleration card number, and the acceleration card port number. To construct the communication domain between the training framework and the inference framework, first load the communication domain configuration; assume that the training framework allocates 2 AI acceleration cards, the inference service starts 2 inference services, and each inference service allocates 2 AI acceleration cards. According to the tasks assigned to each computing node, configure the IP addresses, port numbers, device numbers, and communication domain numbers of the sender and receiver as follows: "comm1": "master":"addr_ip:port"; "devices": {"addr_ip:port":device_id}; {"addr_ip:port":device_id}; "comm2": "master":"addr_ip: port "; "devices": {"addr_ip:port":device_id}; {"addr_ip:port":device_id}.
[0020] Among them, comm1 and comm2 respectively represent the construction of two communication domains. Communication can be carried out between the same communication domains, while communication between different communication domains is not allowed. By constructing the communication domains in the above way, the interaction between irrelevant acceleration cards can be avoided, the accuracy of data transmission objects can be guaranteed, and the data transmission efficiency can be improved. master represents the sending node within this communication domain, and the sending node is the node of the training framework; devices represents the receiving node within this communication domain, that is, the node of the inference framework; addr_ip is the IP of the node, which is configured according to the deployed training node and inference node; port represents the allocated port number, and different port numbers need to be set for the same node with different ports; device_id represents the allocated AI acceleration card number. After configuring the communication domain, start the training framework and the inference framework, and construct the communication domains in the training framework and the inference framework respectively.
[0021] Sending the model weight parameters of the training framework to the above-mentioned inference framework may specifically include: the target acceleration card within the training framework sends the weight parameters stored in the target acceleration card to the corresponding acceleration card within the inference framework through the communication domain corresponding to the acceleration card; the inference framework obtains the model weight parameters based on all the received weight parameters. That is, the model weight parameters are specifically sent by each acceleration card under the training framework using the configured communication domain.
[0022] Among them, sending the weight parameters stored in the target acceleration card to the corresponding acceleration card within the above-mentioned inference framework includes: the target acceleration card combines various weight parameters stored, as well as the weight parameter names and weight parameter sizes corresponding to each weight parameter, to generate a weight parameter package, and sends the weight parameter package to the inference framework. Correspondingly, before the inference framework updates according to the above-mentioned model weight parameters, it also includes: reorganizing all the above-mentioned weight parameters according to the above-mentioned weight parameter names and the above-mentioned weight parameter sizes to obtain the model weight parameters adapted to the inference framework. It can be understood that the weight splitting and storage methods adopted by the training framework and the inference framework may be different. For example, the training framework adopts model parallelization with stream splitting, and the inference framework adopts model parallelization with tensor splitting. Of course, there may also be other differential methods; model parallelization is to allocate different parts of the model to different acceleration cards, and stream splitting is to allocate different layers of the model to different acceleration cards. Under stream splitting, the acceleration card will combine the weight parameters of its corresponding layer, and combine the weight parameter names and weight parameter sizes corresponding to each weight parameter to generate a weight parameter package. Tensor splitting is to split a large tensor (such as a weight matrix) into multiple smaller sub-tensors and store them on different acceleration cards respectively. At this time, the data of a certain layer may be distributed and stored on different acceleration cards.
[0023] Therefore, to ensure that the inference framework uses the model weight parameters, it is necessary to add the weight parameter names and weight parameter sizes as tags to the weight parameter package at the same time. After receiving the data packet of the weight parameters, the inference framework parses it, splits and stores the model weight parameters according to the weight parameter names and weight parameter sizes for the inference framework to use.
[0024] For the above-mentioned reorganization of all weight parameters, the inference framework can specifically reorganize all weight parameters in a parallel manner. That is, the parsing and reorganization process of the weight parameters by the inference framework does not delay the inference process.
[0025] Step S12: The inference framework updates according to the model weight parameters, performs inference verification after the update, and sends the inference result to the training framework.
[0026] The inference framework updates itself according to the model weight parameters and performs inference verification after the update. It should be noted that the process of the training framework sending the model weight parameters to the inference framework does not affect the training of the training framework itself. That is, the training framework is training while the inference framework is also performing inference; moreover, the training framework can specifically send the model weight parameters multiple times at multiple time points during the training process. After the inference model performs inference verification, it feeds back the inference result to the training framework so that the training framework can perform subsequent training according to the inference result.
[0027] Step S13: Obtain the trained model based on the final outputs of the training framework and the inference framework.
[0028] The main outputs of the training framework include the architecture and weight parameters of the model. The main output of the inference framework is the optimized result. The inference framework is used to perform optimization operations such as quantizing, pruning, or operation fusion on the model. Finally, after achieving the training objective, the trained model is obtained based on the final outputs of the training framework and the inference framework, so as to process the data to be processed using the trained model.
[0029] In the related art, in the face of the shortage of AI computing power, technologies such as model quantization and distillation are used to improve the computing power utilization efficiency. However, model quantization will cause a certain decrease in accuracy, and the quantization scheme cannot be used for tasks with relatively high accuracy requirements. And the distillation technology often only solves the computing power problem in deployment and cannot solve the problem of computing power shortage in the training stage. This application proposes to separate training and inference, realizing parallel computing in the training stage and the inference stage, and alleviating the problem of computing power shortage in the training stage.
[0030] As can be seen from the above, in this embodiment, the training framework and the inference framework of the pre-trained model are deployed on different nodes, and the communication domain between the above training framework and the above inference framework has been configured. The above training framework performs training according to the training data, and sends the model weight parameters of the above training framework to the above inference framework through the above communication domain during the training process. The above inference framework is updated according to the above model weight parameters, performs inference verification after the update, and sends the inference result to the above training framework. Based on the final outputs of the training framework and the inference framework, the trained model is obtained. By separating and deploying the training framework and the inference framework on different nodes and realizing the communication between the training framework and the inference framework by configuring the communication domain, the training framework can send the model weight parameters to the inference framework during the training process for the inference framework to perform inference verification, realizing parallel computing of training and inference, making full use of the hardware computing resources on different nodes, accelerating the data processing speed during the training of the training framework and the inference framework, alleviating the problem of computing power shortage in the training and inference stages, and being able to ensure the performance of the model within a limited training duration.
[0031] In some embodiments, sending the model weight parameters of the above training framework to the above inference framework may specifically include: sending the model weight parameters to the inference framework in batches through a communication cache block. Considering the time consumption of weight sending in the training framework and weight receiving in the inference framework, a certain amount of weight parameters are cached in the communication cache block before sending to avoid excessive resource occupation caused by real-time sending. Further, the size of the communication cache block may be determined according to the remaining video memory space after the deployment of the inference framework. For example, the size of the remaining video memory space after the deployment of the inference framework is used as the size of the communication cache block. By such design, the inference framework can have sufficient ability to process after receiving the weight parameters, avoiding the accumulation of weight parameters in the inference framework. Of course, the default size (such as 1GB) can also be used when the cache block size is not configured. It can be seen that on the basis of separating the training of the training framework and the inference framework, the cache block is further used to implement the sending of weight parameters, reducing the time consumption of weight parameter transmission and improving the overall training speed of the training framework and the inference framework.
[0032] In some embodiments, during the training process, sending the model weight parameters of the above training framework to the above inference framework through the above communication domain includes: during the training process, the training framework sends the current model weight parameters to the inference framework every n steps. Considering the time consumption of weight parameter sending in the model training stage and weight parameter receiving in the model inference stage, inference is performed once after multiple steps of training. It can be understood that if the training framework sends the currently updated model weight parameters to the inference framework after each step (step), facing hundreds or even thousands of steps in the whole training, it will cause serious time consumption and prolong the training time of the model. Therefore, the model weights are sent to the inference framework once after training for n steps. The value of n can be customized according to the actual size of the model, communication-related resources, and model training requirements, etc.
[0033] In some embodiments, sending the model weight parameters of the above training framework to the above inference framework may specifically include: the training framework sends the model weight parameters to the inference framework asynchronously. That is, the interaction between the training framework and the inference framework is realized by means of asynchronous communication. In asynchronous communication, the sender and the receiver do not need to be online or operate synchronously at the same time. The sender can send messages, while the receiver can receive and process these messages at any time without real-time interaction, that is, the sender and the receiver can operate independently, reducing the waiting time. Through asynchronous communication, the efficiency and flexibility of the interaction between the training framework and the inference framework can be improved.
[0034] Specifically, sending the model weight parameters of the above training framework to the above inference framework may specifically include: the training framework sending a weight update request to the inference framework; sending the model weight parameters of the training framework to the inference framework; after the sending of the model weight parameters is completed, sending an inference verification request to the inference framework, and the inference framework performing inference verification according to the inference verification request. That is, the training framework first sends a weight update request to the inference framework, after the sending of the weight update request is completed, sends the weight parameters to the inference framework, and then sends an inference verification request to the inference framework after the sending of the current model parameter weights is completed, so that the inference framework can start receiving the weight parameters after receiving the weight update request, and can clarify that all the current model parameter weights have been sent after receiving the inference verification request, and perform inference verification according to all the received model weight parameters.
[0035] In some embodiments, the accelerator cards of the training framework and the inference framework are homogeneous devices, that is, the training framework and the inference framework use the same type of accelerator cards. At this time, sending the model weight parameters of the training framework to the inference framework through the above communication domain may include: the accelerator cards of the training framework and the inference framework directly communicating through the above communication domain to send the model weight parameters to the inference framework. That is, in the case of homogeneous devices, direct communication is adopted between the accelerator cards, such as using NVLink (a bus and its communication protocol) communication technology, etc. Such communication does not require occupying CPU resources and further improves the communication efficiency and speed.
[0036] In related technologies, there are also computing power optimization methods for homogeneous devices. For example, by constructing a distributed computing network and a shared computing power platform, resource optimization configuration is achieved. However, such shared computing power platforms are only for homogeneous AI accelerator card devices and cannot be used for heterogeneous accelerator cards. Based on the solution of separating the training framework and the inference framework in this application and constructing a communication domain between the training framework and the above inference framework, this application can not only support the interaction of homogeneous devices but also support the interaction of heterogeneous devices.
[0037] Specifically, when the acceleration card of the training framework and the acceleration card of the above-mentioned inference framework are heterogeneous devices; the sending of the model weight parameters of the training framework to the inference framework through the above-mentioned communication domain may specifically include: sending the model weight parameters of the training framework to the inference framework through the above-mentioned communication domain using a processor-based remote access method. That is to say, in the case of heterogeneous devices, the training framework and the inference framework use different types of cards, so the communication between the cards needs to use technologies such as RPC (Remote Procedure Call) or RDMA (Remote Direct Memory Access) for communication, or a combination of multiple communication technologies can also be used. Central Processing Unit (CPU) resources are required in the case of heterogeneity; that is, the acceleration card first sends the data to the processor of the local node, and then the processor sends it to the other node.
[0038] That is to say, the training framework and the inference framework are deployed using different devices. Due to the use of heterogeneous devices, the existing communication libraries cannot be directly used. This application conducts data communication for heterogeneous device communication based on the processor-based remote access method. Moreover, the model inference framework needs to update the weights dynamically online. Taking the vLLM as an example of the inference framework, vLLM currently only supports offline loading of the model weight deployment service and does not support online dynamic weight update. This application uses GRPC and RDMA data communication to achieve online dynamic weight update of vLLM to dynamically receive the weight parameters sent by the model training framework; in the static graph mode of vLLM, the weight parameters can still be updated normally, and the static graph does not need to be recompiled, greatly improving the inference performance.
[0039] It can be seen that by supporting the training framework and the inference framework to use heterogeneous acceleration cards, the advantages of different acceleration cards in different scenarios are fully utilized to make full use of the computing power resources under different architectures; and the overlapping calculation is used to hide the time of the inference verification stage to accelerate the convergence speed of the model training stage and solve the problem of insufficient computing power resources in model training.
[0040] Based on the above embodiments, the embodiments of this application also disclose a specific model training method. For example Figure 2 As shown, the model training includes the following steps: S21: Initialize the training framework, initialize the cache, initialize the communication domain, construct the sending end of the communication domain (i.e., the training framework), initialize the inference framework, and construct the receiving end of the communication domain (i.e., the inference framework).
[0041] S22: The training framework loads the training data, executes the training program, sends a weight update request after training for n steps, and sends the weight parameters through communication based on GRPC and RDMA.
[0042] When sending weight parameters, in order to improve communication efficiency and reduce the number of communications, multiple weight parameters are combined so that their size does not exceed the size of the communication buffer block. At the same time, the weight parameter names and the sizes of the weight parameters are combined and added to the communication data. The sending of weight parameters is processed asynchronously by multiple threads without affecting the progress of the main training process.
[0043] S23: After the training framework has sent the weight parameters, an inference verification request is sent so that the inference framework can perform inference on the validation set.
[0044] S24: After the inference framework receives the weight update request, it receives the weight parameters through communication based on GRPC and RDMA.
[0045] The specific weight update process of the inference framework is as Figure 3 shown. First, the inference framework parses the weight parameter names and the sizes of the weight parameters. According to the parameter names and the sizes of the weight parameters, it parses the received weight parameters and updates the weight parameters in the inference framework. Since the inference framework has fusion operators that combine multiple weight parameters, the weight parameter names and the sizes of the weight parameters may not exactly match those in the training framework. Before updating the parameters, reorganization processing needs to be performed according to the code for loading weight parameters in the inference framework to ensure that the model weight parameters are adapted to the inference framework, thereby ensuring the correctness of the inference results.
[0046] S25: The inference framework receives the inference request and processes the inference request. It determines whether the inference verification of all the weight parameters received this time has been completed. If not, it continues the inference. If completed, it returns the inference results to the training framework.
[0047] S26: The training framework obtains the inference results returned by the inference framework and completes the verification of the validation set.
[0048] S26: It is determined whether the training requirements are met. If so, the training ends. Otherwise, the training of the training framework and the inference framework continues.
[0049] By separating the calculations of the training framework and the inference framework, the advantages of the training framework and the inference framework are effectively utilized to accelerate the performance of the inference stage, thereby improving the overall training performance. Moreover, to address the problem of insufficient computing acceleration card resources, this application supports the training framework and the inference framework to use heterogeneous devices for computing. Thus, different types of acceleration cards can be dynamically extended for training, effectively utilizing the diverse computing power of the computing cluster, solving the problem of coordinating diverse computing power, and enabling the training of models with larger numbers of parameters.
[0050] Correspondingly, an embodiment of this application also discloses a data processing method applied to the aforementioned trained model, which specifically includes the following steps: S31: Obtain the data to be processed, and input the data to be processed into the trained model above.
[0051] S32: Obtain the processing result corresponding to the data to be processed according to the output of the trained model above.
[0052] The pre-trained model is a neural network model, specifically it can be a large model. A large model refers to a machine learning model with a very large number of parameters, and a large model can process and analyze large-scale data sets. The trained large model above can be a large language model, a vision large model or a multi-modal large model, etc.; if it is a large language model, the data to be processed is a text sequence, and according to the input text sequence (such as a sentence, a paragraph, a question, etc.), the output processing result of the trained model is also a text sequence, such as an answer, a generated text, a translated text, etc. If it is a vision large model, the data to be processed is image or video data, and according to the input image or video data, the trained model outputs an image classification result, a target detection box or an image segmentation mask, etc. If it is a multi-modal large model, the data to be processed can be data of multiple modalities (such as text, image, audio, video), and according to the input data of multiple modalities, the trained model outputs corresponding processing results, such as text generation, image generation, speech generation, etc.
[0053] Correspondingly, an embodiment of the present application also discloses a model training device. The training framework and the inference framework of the pre-trained model are deployed on different nodes, and the communication domain between the training framework and the inference framework has been configured. The device includes: A training framework training module, configured to enable the training framework to perform training according to training data, and send the model weight parameters of the training framework to the inference framework through the communication domain during the training process; An inference framework inference module, configured to enable the inference framework to update according to the model weight parameters, perform inference verification after the update, and send the inference result to the training framework; A trained model determination module, configured to obtain the trained model based on the final outputs of the training framework and the inference framework.
[0054] As can be seen from the above, in this embodiment, the training framework and the inference framework of the pre-trained model are deployed on different nodes, and the communication domain between the training framework and the inference framework has been configured. The training framework performs training according to the training data, and during the training process, sends the model weight parameters of the training framework to the inference framework through the communication domain; the inference framework updates according to the model weight parameters, performs inference verification after the update, and sends the inference result to the training framework; based on the final outputs of the training framework and the inference framework, the trained model is obtained. By separately deploying the training framework and the inference framework on different nodes and configuring the communication domain to enable communication between the training framework and the inference framework, the training framework can send the model weight parameters to the inference framework during the training process so that the inference framework can perform inference verification, realizing parallel computing for training and inference, accelerating the data processing speed of model training and inference, alleviating the problem of shortage of computing power during the training and inference stages, and being able to ensure the performance of the model within a limited training duration.
[0055] In some specific embodiments, the model training device may specifically include: A communication domain construction unit, configured to construct a communication domain between the training framework and the inference framework according to the communication domain configuration; The communication domain is configured according to the corresponding relationship between the acceleration cards in the training framework and the acceleration cards in the inference framework, and is determined with the acceleration cards in the training framework as the dimension; communication is supported between the same communication domains, and communication is not possible between different communication domains.
[0056] In some specific embodiments, the communication domain configuration includes the training framework node address, the inference framework node address, the acceleration card number, and the acceleration card port number.
[0057] In some specific embodiments, the training framework training module may specifically include: A weight sending unit, configured to send the weight parameters stored in the target acceleration card in the training framework to the corresponding acceleration card in the inference framework through the communication domain corresponding to the acceleration card; the inference framework obtains the model weight parameters according to all the received weight parameters.
[0058] In some specific embodiments, the training framework training module may specifically include: A merging unit, configured to merge the multiple weight parameters stored in the target acceleration card, as well as the weight parameter name and the weight parameter size corresponding to each weight parameter, into a weight parameter package, and send the weight parameter package to the inference framework; The inference framework inference module may specifically include: A weight reorganization unit, configured to reorganize all the weight parameters according to the weight parameter names and the weight parameter sizes before the inference framework updates according to the model weight parameters, so as to obtain the model weight parameters adapted to the inference framework.
[0059] In some specific embodiments, the weight reorganization unit is configured to reorganize all the weight parameters in a parallel manner by the inference framework.
[0060] In some specific embodiments, the training framework training module may specifically include: A weight sending unit, configured to send the model weight parameters to the inference framework in batches through a communication cache block.
[0061] In some specific embodiments, the size of the communication cache block is determined according to the remaining video memory space after the inference framework is deployed.
[0062] In some specific embodiments, the training framework training module may specifically include: A weight sending unit, configured to send the current model weight parameters to the inference framework every n steps during the training process by the training framework.
[0063] In some specific embodiments, the training framework training module may specifically include: A weight sending unit, configured to send the model weight parameters to the inference framework in an asynchronous manner by the training framework.
[0064] In some specific embodiments, the weight sending unit may specifically include: A weight update request sending unit, configured to send a weight update request from the training framework to the inference framework; A weight parameter sending unit, configured to send the model weight parameters of the training framework to the inference framework; An inference verification request sending unit, configured to send an inference verification request to the inference framework after the model weight parameters are sent, and the inference framework performs inference verification according to the inference verification request.
[0065] In some specific embodiments, the acceleration card of the training framework and the acceleration card of the inference framework are heterogeneous devices; The training framework training module may specifically include: A remote sending unit, configured to send the model weight parameters of the training framework to the inference framework by using a processor-based remote access method through the communication domain.
[0066] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute any of the above model training methods or the steps in the foregoing data processing method embodiments.
[0067] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute any of the above model training methods or the steps in the foregoing data processing method embodiments when running.
[0068] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs that can store computer programs.
[0069] An embodiment of the present application further provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements any of the above model training methods or the steps in the foregoing data processing method embodiments.
[0070] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements any of the above model training methods or the steps in the foregoing data processing method embodiments.
[0071] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.
[0072] The above has introduced in detail a model training method provided by this application. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of this application.
Claims
1. A model training method, characterized in that, The training framework and the inference framework of the pre-trained model are deployed on different nodes, and the communication domain between the training framework and the inference framework has been configured, including: The training framework performs training based on training data, and during the training process, sends the model weight parameters of the training framework to the inference framework through the communication domain; The inference framework is updated according to the model weight parameters, performs inference verification after the update, and sends the inference result to the training framework; Based on the final outputs of the training framework and the inference framework, a trained model is obtained.
2. The model training method according to claim 1, wherein The configuration of the communication domain between the training framework and the inference framework includes: Construct the communication domain between the training framework and the inference framework according to the communication domain configuration; The communication domain is configured according to the correspondence between the acceleration cards in the training framework and the acceleration cards in the inference framework, and is determined with the acceleration cards in the training framework as the dimension; communication is supported between the same communication domains, and communication is not possible between different communication domains.
3. The model training method according to claim 2, wherein The communication domain configuration includes the training framework node address, the inference framework node address, the acceleration card number, and the acceleration card port number.
4. The model training method according to claim 2, wherein Sending the model weight parameters of the training framework to the inference framework includes: The target acceleration card in the training framework sends the weight parameters stored in the target acceleration card to the corresponding acceleration card in the inference framework through the communication domain corresponding to the acceleration card; The inference framework obtains the model weight parameters according to all the received weight parameters.
5. The model training method according to claim 4, wherein Sending the weight parameters stored in the target acceleration card to the corresponding acceleration card in the inference framework includes: The target acceleration card combines the stored multiple weight parameters, as well as the weight parameter name and the weight parameter size corresponding to each weight parameter, to generate a weight parameter package, and sends the weight parameter package to the inference framework; Before the inference framework is updated according to the model weight parameters, it also includes: According to the weight parameter name and the weight parameter size, all the weight parameters are reorganized to obtain the model weight parameters adapted to the inference framework.
6. The model training method according to claim 5, wherein Reorganizing all the weight parameters includes: The inference framework reorganizes all the weight parameters in a parallel manner.
7. The model training method according to claim 1, wherein Sending the model weight parameters of the training framework to the inference framework includes: Sending the model weight parameters to the inference framework in batches through the communication cache block.
8. The model training method according to claim 7, characterized in that The size of the communication cache block is determined according to the remaining video memory space after the deployment of the inference framework.
9. The model training method according to claim 1, wherein During the training process, sending the model weight parameters of the training framework to the inference framework through the communication domain includes: During the training process, the training framework sends the current model weight parameters to the inference framework every n steps.
10. The model training method according to claim 1, wherein Sending the model weight parameters of the training framework to the inference framework includes: The training framework sends the model weight parameters to the inference framework in an asynchronous manner.
11. The model training method according to claim 10, characterized in that, Sending the model weight parameters of the training framework to the inference framework includes: The training framework sends a weight update request to the inference framework; Sending the model weight parameters of the training framework to the inference framework; After the transmission of the model weight parameters is completed, an inference verification request is sent to the inference framework, and the inference framework performs inference verification according to the inference verification request.
12. The model training method according to any one of claims 1 to 11, characterized in that, The acceleration card of the training framework and the acceleration card of the inference framework are heterogeneous devices; Sending the model weight parameters of the training framework to the inference framework through the communication domain includes: Sending the model weight parameters of the training framework to the inference framework through the communication domain by using a processor-based remote access method.
13. A data processing method, characterized in that, Applied to the trained model according to any one of claims 1 to 12, including: Obtaining data to be processed and inputting the data to be processed into the trained model; Obtaining a processing result corresponding to the data to be processed according to the output of the trained model.
14. An electronic device, characterized in that, Including: A memory for storing a computer program; A processor for executing the computer program to implement the model training method according to any one of claims 1 to 12 or the data processing method according to claim 13.
15. A computer program product, characterized in that, Including a computer program which, when executed by a processor, implements the model training method according to any one of claims 1 to 12 or the data processing method according to claim 13.
Citation Information
Patent Citations
Distributed training and reasoning method, system and device based on artificial intelligence, and readable storage medium
CN114035937A
Construction method and system of heterogeneous computing system
CN114116236A
Distributed training method and device, computer equipment, storage medium and product
CN114327399A
Computing system, model training method, device and product
CN116541338A
Model training method and device, electronic equipment and storage medium
CN117556921A
Cited By
Content generation method and device
CN120450940A