Parameter updating method of ai model, ai service system, and inference device
Patent Information
- Application Number
- CN202510386354.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2026-09-29
AI Technical Summary
[0004]然而,由于模型参数量巨大,推理系统每次部署更新后的模型参数时均会消耗较长时间,导致推理业务长时间暂停,造成资源浪费
Smart Images

Figure CN122840217A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a parameter update method for an AI model, an AI service system, and an inference device. Background Technology
[0002] With the rapid development of large-scale artificial intelligence models, integrated model training and inference platforms have emerged, enabling the entire process from model training to model inference to be completed within a single platform.
[0003] In related technologies, the integrated model training and inference platform includes a training system and an inference system. The training system trains the model based on the samples generated by the inference system at regular intervals to update the model parameters and broadcasts the updated model parameters to the inference system. The inference system needs to pause its inference operations to deploy the updated model parameters to the underlying hardware computing unit and restart its inference operations after deployment is complete.
[0004] However, due to the large number of model parameters, the inference system takes a long time to deploy updated model parameters each time, resulting in prolonged pauses in inference operations and wasted resources. Summary of the Invention
[0005] This application provides a method for updating parameters of an AI model, an AI service system, and an inference device, which can realize online updates of model parameters without interrupting inference operations and effectively avoid resource waste.
[0006] Firstly, a parameter update method for an AI model is provided, applied to an AI service system. The AI service system includes a training system and an inference system. The training system is used to train the AI model, and the inference system is used to perform inference tasks based on the AI model. The AI service system can also be understood as an integrated model training and inference platform. The method includes:
[0007] The inference system obtains notification messages and model layering information. The notification messages indicate that the training system has trained the AI model, and the model layering information indicates multiple parts of the AI model, each part including at least one layer of the AI model.
[0008] During the inference process, the inference system performs multiple parameter updates of the AI model based on the model's hierarchical information. Each time, it retrieves the model parameters of a portion of the multiple parts from the training system and updates the corresponding part of the model parameters in the inference system.
[0009] The AI model in this application is, for example, a large language model (LLM) based on the Transformer architecture, which can be applied to fields such as speech processing, text processing, video processing, or audio processing. Specific application scenarios include, but are not limited to, intelligent writing assistants, automatic translation systems, intelligent customer service robots, entertainment, transportation, healthcare, education, finance, law, information recommendation, and so on.
[0010] In the above method, when the inference system knows that the training system has trained the AI model, it can perform multiple parameter updates of the AI model according to the multiple parts of the AI model indicated by the model hierarchical information during the inference task. Each time, it obtains the model parameters of one part of the multiple parts from the training system and updates the model parameters of the corresponding part. In this way, the model parameters are gradually updated during the inference process, realizing online model updates for inference business without interruption and effectively avoiding resource waste.
[0011] In some embodiments, during the inference system's execution of an inference task based on the AI model, the system performs multiple parameter updates of the AI model based on the model's hierarchical information, including: the inference system obtaining model parameters of a first part of multiple parts from the training system and updating the model parameters of the first part in the inference system; and during the inference system's execution of an inference task based on the updated first part, obtaining model parameters of a second part of multiple parts from the training system and updating the model parameters of the second part in the inference system.
[0012] Using the above method, the process of the inference system executing inference tasks and the process of the inference system updating parameters are executed in parallel, realizing uninterrupted online model updates for inference operations and effectively avoiding resource waste.
[0013] In some embodiments, the duration for which the inference system performs an inference task based on the updated first part is greater than or equal to the duration for which the inference system obtains the model parameters of the second part from the training system; the method further includes: after performing an inference task based on the updated first part, the inference system performs an inference task based on the updated second part.
[0014] By using the above method, since the time taken for the inference system to perform the inference task based on the first part is greater than or equal to the time taken for the inference system to obtain the model parameters of the second part, after the inference system performs the inference task based on the first part, the inference system has already obtained the model parameters of the second part. Therefore, the inference system can continue to perform the inference task based on the updated second part, avoiding the inference task waiting due to excessive time spent obtaining parameters. This improves the parameter update efficiency while ensuring the inference efficiency.
[0015] In some embodiments, after the inference system obtains the notification message and model hierarchical information, the method further includes: the inference system establishing a parameter acquisition link with the training system based on the parameter storage address of the AI model in the training system; the inference system obtaining model parameters of a portion of multiple parts from the training system, including: the inference system obtaining model parameters of a portion of multiple parts from the training system through the parameter acquisition link.
[0016] In the above method, the parameter storage address is, for example, a network IP address. The inference system establishes a parameter acquisition link with the training system based on the parameter storage address of the AI model in the training system, providing a channel for subsequent transmission of model parameters.
[0017] In some embodiments, the AI service system further includes a scheduling system for scheduling the training system and the inference system. The method further includes: the scheduling system sending a parameter storage address to the inference system; the scheduling system receiving a notification message sent by the training system and forwarding the notification message to the inference system.
[0018] In the above method, the scheduling system acts as a bridge between the training system and the inference system. It can coordinate the work of each link in the entire process of model training and inference, avoid information transmission chaos, and ensure that the parameters of the inference system can be updated in a timely manner after the training system is completed, so as to achieve seamless connection between training and inference.
[0019] In some embodiments, the method further includes: the inference system determining model layering information based on the number of layers in the AI model, the inference time of the AI model, the inference time of each layer in the AI model, and the transmission time of model parameters of each layer in the AI model.
[0020] Using the above methods, the inference system comprehensively considers multiple dimensions to determine the model layering information, making the model layering more consistent with the actual operating state, dynamically adapting to different model structures and performance requirements, optimizing the rationality of the layering strategy, and further improving the efficiency of parameter updates.
[0021] In some embodiments, the AI model is a first AI model or a second AI model, and the method further includes: an inference system performing an inference task based on the first AI model to generate an inference result, and evaluating the inference result based on the second AI model to generate an evaluation result, and sending the inference result and the evaluation result to a training system; the training system training the first AI model and the second AI model based on the inference result and the evaluation result.
[0022] By deploying the first AI model and the second AI model in the AI service system, a self-reinforcing closed loop of "training-inference-data feedback" can be formed, continuously improving the system's performance in the inference generation and evaluation stages.
[0023] Secondly, an AI service system is provided, which includes a training system and an inference system. The training system is used to train an AI model, the inference system is used to perform inference tasks based on the AI model, and the AI service system is used to implement the parameter update method of the AI model provided by the first aspect or any possible implementation of the first aspect.
[0024] Thirdly, a reasoning apparatus is provided, the reasoning apparatus comprising:
[0025] The acquisition module is used to acquire notification messages and model layer information. The notification message indicates that the AI model has finished training, and the model layer information indicates multiple parts of the AI model, each part including at least one layer of the AI model.
[0026] The parameter update module is used to perform multiple parameter updates of the AI model based on the model's hierarchical information during the inference task performed by the AI model. Each time, it obtains the model parameters of a part of multiple parts and updates the model parameters of the corresponding part in the inference device.
[0027] In some embodiments, the parameter update module is configured to: obtain model parameters of a first part among a plurality of parts and update the model parameters of the first part in the inference device; and, during the execution of an inference task based on the updated first part, obtain model parameters of a second part among a plurality of parts and update the model parameters of the second part in the inference device.
[0028] In some embodiments, the duration of performing the inference task based on the updated first part is greater than or equal to the duration of obtaining the model parameters of the second part; the inference apparatus further includes an inference task execution module for: performing an inference task based on the updated second part after performing the inference task based on the updated first part.
[0029] In some embodiments, the inference device further includes a link establishment module, configured to: after obtaining notification messages and model hierarchical information, establish a parameter acquisition link with the training system based on the parameter storage address of the AI model in the training system; and a parameter update module, configured to obtain model parameters of a portion of multiple parts from the training system through the parameter acquisition link.
[0030] In some embodiments, the inference device further includes a layer information determination module, configured to: determine model layer information based on the number of layers in the AI model, the inference time of the AI model, the inference time of each layer in the AI model, and the transmission time of model parameters of each layer in the AI model.
[0031] In some embodiments, the AI model is a first AI model or a second AI model, and the inference device is further configured to: perform an inference task based on the first AI model, generate an inference result, and evaluate the inference result based on the second AI model to generate an evaluation result.
[0032] Fourthly, a computing device is provided, the computing device including a processor and a memory, the processor being configured to execute at least a segment of program code stored in the memory to enable the computing device to implement the parameter update method of the AI model provided by the first aspect or any possible implementation thereof.
[0033] Fifthly, a computer-readable storage medium is provided for storing at least one piece of program code for implementing a parameter update method for an AI model as provided in the first aspect or any possible implementation thereof. The storage medium includes, but is not limited to, volatile memory, such as random access memory, and non-volatile memory, such as flash memory, hard disk drive (HDD), and solid-state drive (SSD).
[0034] Sixthly, a computer program product is provided for implementing the parameter update method for an AI model as provided in the first aspect or any possible implementation thereof. The computer program product can be a software installation package, which can be downloaded and executed on a computing device when the parameter update method for the aforementioned AI model needs to be implemented. Attached Figure Description
[0035] Figure 1 This is a schematic diagram of an implementation environment provided in an embodiment of this application;
[0036] Figure 2 This is a schematic diagram illustrating the operation of an AI service system using an AI model, as provided in an embodiment of this application.
[0037] Figure 3 This is a schematic diagram of the hardware structure of a computing device provided in an embodiment of this application;
[0038] Figure 4 This is a functional architecture diagram of an AI service system provided in an embodiment of this application;
[0039] Figure 5 This is a flowchart of a parameter update method for an AI model provided in an embodiment of this application;
[0040] Figure 6 This is a flowchart illustrating the operation of an AI service system provided in an embodiment of this application.
[0041] Figure 7 This is a flowchart illustrating the operation of an inference system provided in an embodiment of this application;
[0042] Figure 8 This is a schematic diagram of the structure of an inference device provided in an embodiment of this application. Detailed Implementation
[0043] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be further described in detail below with reference to the accompanying drawings. It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application are authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the inference requests, training requests, and artificial intelligence models involved in this application are all obtained under fully authorized conditions.
[0044] To facilitate understanding, the key terms and concepts involved in this application will be explained below.
[0045] Artificial intelligence (AI) models are a class of mathematical algorithm models that use machine learning concepts to solve practical problems. Typically, AI models include a large number of parameters and calculation formulas (or calculation rules).
[0046] Generative large models are used to autonomously generate new content that conforms to semantic logic, such as text, images, audio, and video, based on given inference requests. Examples of generative large models include large language models (LLMs) and multimodal models. Taking a Transformer-based AI model as an example, this model typically includes multiple Transformer layers, also known as Transformer blocks. A Transformer block refers to the basic building block in a Transformer model; multiple Transformer blocks are stacked sequentially to form a complete Transformer model, enabling effective feature extraction and representation learning from the input sequence. Typically, Transformer layers include components such as multi-head self-attention, residual connections, layer normalization, and feedforward networks (FFNs).
[0047] An accelerator, also known as an acceleration chip, acceleration device, acceleration card, or computing card, is a specialized hardware device or computer system designed to accelerate computation in AI scenarios. For example, in the inference scenario of an AI model, an accelerator can be called an inference card, and in the training scenario of an AI model, it can be called a training card. In this application, an accelerator may include, for example, a graphics processing unit (GPU), a neural network processing unit (NPU), an intelligent processing unit (IPU), a tensor processing unit (TPU), or a domain-specific architecture (DSA) chip, and is not limited to these.
[0048] The application scenarios of this application are described below.
[0049] This application pertains to the parameter update scenario of AI models within an integrated model training and inference platform. This platform is a system architecture that coordinates the training and inference of AI models, dynamically synchronizing updated model parameters to the inference process in real-time or near real-time, thus connecting model iteration with inference operations. The AI model used in this application is, for example, a large language model (LLM) based on the Transformer architecture, applicable to fields such as speech processing, text processing, video processing, and audio processing. Specific application scenarios include, but are not limited to, intelligent writing assistants, automatic translation systems, intelligent customer service robots, entertainment, transportation, healthcare, education, finance, law, and information recommendation.
[0050] In related technologies, integrated model training and inference platforms include a training system and an inference system. The training system periodically trains the model based on samples generated by the inference system to update model parameters and broadcasts the updated parameters to the inference system. The inference system then needs to pause its inference operations to deploy the updated model parameters to the underlying hardware computing units and restart its inference operations after deployment. However, due to the large number of model parameters, each deployment of updated model parameters by the inference system takes a considerable amount of time, resulting in prolonged pauses in inference operations and wasted resources.
[0051] Based on this, this application provides an AI model parameter update method for an integrated model training and inference platform, wherein the inference system can deploy the updated model parameters in a timely manner without pausing inference operations, thereby avoiding resource waste.
[0052] The implementation environment of this application is described below.
[0053] Figure 1 This is a schematic diagram of an implementation environment provided in an embodiment of this application. For example... Figure 1 As shown, the implementation environment includes an AI service system 100 (i.e., a model training and inference integrated platform), which includes a scheduling system 101, an inference system 102, and a training system 103. The scheduling system 101, the inference system 102, and the training system 103 are connected via wired or wireless networks.
[0054] The scheduling system 101 is used to schedule the inference system 102 and the training system 103, coordinating the training and inference operations of the AI model. For example, the scheduling system 101 acts as a bridge between the inference system 102 and the training system 103, responsible for transmitting notification messages, model parameters, training samples, etc. Illustratively, the scheduling system 101 can provide users with an interface to interact with the AI service system 100. For example, it receives inference requests initiated by users through this interface and forwards the inference requests to the inference system 102 for execution.
[0055] The inference system 102 is used to execute inference tasks based on the deployed AI model according to the instructions of the scheduling system 101. In some embodiments, the inference system 102 includes an inference scheduling system and an inference execution system. The inference scheduling system, as a unified scheduling center for executing inference tasks, has functions such as initialization processing, deploying inference instances (i.e., AI models) on the inference execution system, issuing inference tasks to the inference execution system, and feeding back various notification messages from the AI service system 100 to the inference execution system. The inference execution system, as an execution component with deployed inference instances, has functions such as executing inference tasks to generate inference results and updating model parameters.
[0056] The training system 103 is used to execute training tasks based on the deployed AI model according to the instructions of the scheduling system 101. In some embodiments, the training system 103 includes a training scheduling system and a training execution system. The training scheduling system, as a unified scheduling center for executing training tasks, has functions such as scheduling different model training execution systems and feeding back various notification messages from the AI service system 100 to different training execution systems. The training execution system, as an execution component with deployed training instances, has functions such as executing training tasks to update model parameters.
[0057] In this application, the AI service system 100 may involve one or more AI models. Illustratively, when the AI service system 100 involves multiple AI models, these multiple AI models include a first AI model and a second AI model.
[0058] The first AI model is used to perform inference tasks to generate inference results. In some embodiments, the first AI model is also called a generator model (G model for short). This model is based on a Large Language Model (LLM) and, through fine-tuning or architectural modification, enables it to have task-oriented generative capabilities. This model can receive tasks input from the system (such as text generation, decision planning, etc.), and utilize the semantic understanding and generative characteristics of LLM, combined with algorithms such as Monte Carlo Tree Search (MCTS), rejection sampling (such as Best of N, BON), or serial revision, to generate inference results with multiple branches. For example, in the inference or testing phase, the first AI model acts as a content producer, expanding the inference results with multiple branches based on the input question, providing rich candidate content for subsequent evaluation.
[0059] The second AI model is used to evaluate the inference results to generate an evaluation result. In some embodiments, the second AI model is also called a verifier model (V model for short). This model is based on LLM and, through fine-tuning or architectural modification, enables it to evaluate the inference results. For example, during the inference or testing phase, the second AI model acts as a content evaluator, evaluating the inference results generated by the first AI model and generating an evaluation result. The evaluation result can be used to determine the final inference result fed back to the user from the inference results of multiple branches.
[0060] In addition, the inference results generated by the first AI model and the evaluation results generated by the second AI model can be used as training samples to participate in the training process of the first AI model and the second AI model during the training-time phase.
[0061] By deploying the first and second AI models in the AI service system 100, a self-reinforcing closed loop of "training-inference-data feedback" can be formed, continuously improving the system's performance in the inference generation and evaluation stages. For example, refer to... Figure 2 , Figure 2 This is a schematic diagram illustrating the operation of an AI service system using an AI model, as provided in an embodiment of this application. Figure 2 As shown, the process of an AI service system running an AI model includes the following stages:
[0062] (1) During the inference phase (Testing-Time) of the AI service system, the inference system inputs a predefined problem or task description into the first AI model (i.e., G model) through system prompts to clarify the target to be processed (such as text generation).
[0063] (2) The reasoning system generates inference results of multiple branches by using the first AI model and combining algorithms such as MCTS, BON or serial revision method.
[0064] (3) The inference system evaluates the inference results output by the first AI model from multiple dimensions, including process and outcome, using the second AI model (i.e., the V model), and obtains the evaluation results. Schematic, the inference system can determine the final inference result output to the user based on the evaluation results.
[0065] (4) The reasoning system stores predefined questions, reasoning results and evaluation results as training samples in the training sample database.
[0066] (5) During the training phase of the AI service system, the training system obtains training samples from the training sample database to train the first AI model and the second AI model respectively.
[0067] (6) The training system updates the model parameters of the first AI model and the second AI model based on the training results of the first AI model and the second AI model.
[0068] (7) The training system sends the updated first AI model and the updated second AI model to the inference system so that the inference system can perform the continued inference phase based on the updated AI model.
[0069] In addition, the AI service system 100 provided in this application typically owns and manages large-scale computing devices, such as accelerator clusters and physical server clusters, which are used to support the functions provided by the AI service system 100.
[0070] In some embodiments, the AI service system 100 can also be a cloud data center, cloud computing platform, cloud server, etc. Taking a cloud computing platform as an example, a cloud computing platform, or simply cloud platform, refers to a service based on hardware and software resources, providing computing, networking, and storage capabilities. Through the network "cloud," massive amounts of data are processed and analyzed remotely before being returned to the user, featuring large scale, distributed computing, virtualization, high availability, scalability, on-demand service, and security. Cloud servers can achieve rapid deployment and release of configurable computing resources with relatively low management costs or low interaction complexity between users and service providers. In some embodiments, the cloud platform uses virtualization technology to virtualize physical computing resources into multiple virtual instances for different tasks. Automated management tools are used to configure, monitor, and manage computing resources. Distributed computing and parallel computing technologies are utilized to improve computing efficiency; this application is not limited to these.
[0071] The aforementioned wireless or wired networks utilize standard communication technologies and / or protocols. These networks are typically Transmission Control Protocol / Internet Protocol (TCP / IP) networks used in data center networks, as well as RDMA networks such as RoCE networks and InfiniBand (IB) networks; no limitation is made thereto. In other embodiments, customized and / or dedicated data communication technologies can be used to replace or supplement the aforementioned data communication technologies.
[0072] Based on the aforementioned implementation environment, this application provides a computing device that can be configured as an independent physical server or any computing node in a cloud data center, etc. Accordingly, the AI model parameter update method provided in this application can be executed by a computing device cluster consisting of at least one computing device. In other embodiments, the AI model parameter update method provided in this application can be executed collaboratively by a first computing device cluster consisting of at least one computing device and a second computing device cluster consisting of at least one computing device, wherein the first computing device cluster is used to implement the functions of the inference system 102, and the second computing device cluster is used to implement the functions of the training system 103.
[0073] Indicatively, for reference Figure 3 , Figure 3 This is a schematic diagram of the hardware structure of a computing device provided in an embodiment of this application. Figure 3 As shown, the computing device 300 includes a memory 301, a processor 302, a communication interface 303, and a bus 304. The memory 301, processor 302, and communication interface 303 are interconnected via the bus 304.
[0074] The memory 301 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or it may be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto. Illustratively, the memory 301 is used to store at least one piece of program code. When the program code stored in the memory 301 is executed by the processor 302, the processor 302 is used to execute the parameter update method of the AI model provided in the following method embodiments.
[0075] Processor 302 may be a network processor (NP), a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), or an integrated circuit used to control the execution of the program in this application. Processor 302 may be a single-core processor or a multi-core processor. The number of processors 302 may be one or more. For example, depending on the actual needs of the AI model parameter update method, computing device 300 may include multiple processors 302, each of which may include one or more processor cores. These multiple processors 302 may include one or more of the following: CPU, GPU, NPU, IPU, TPU, and DSA chip, and are not limited thereto.
[0076] The communication interface 303 uses a transceiver module, such as a transceiver, to enable communication between the computing device 300 and other devices or communication networks. For example, data can be acquired through the communication interface 303.
[0077] The memory 301 and the processor 302 can be set separately or integrated together.
[0078] Bus 304 may include a pathway for transmitting information between various components of computing device 300 (e.g., memory 301, processor 302, communication interface 303).
[0079] The parameter update method for the AI model provided in this application is described below.
[0080] For easier understanding, please refer to the following: Figure 4 The principles of this application will be explained in conjunction with the AI service system described in the aforementioned implementation environment. Figure 4 This is a functional architecture diagram of an AI service system provided in an embodiment of this application. Figure 4 As shown, the AI service system is used to provide inference function 401 and training function 402.
[0081] The inference function 401 is implemented by the inference system in the AI service system. In this application, the inference function 401 includes an acquisition function 4011 and a parameter update function 4012.
[0082] The acquisition function 4011 is used to acquire notification messages and model layering information. The notification message indicates that the training system has trained the AI model, and the model layering information indicates multiple parts of the AI model, each part including at least one layer. That is, the AI model comprises multiple layers, dividing the AI model into multiple parts, each part including at least one layer. Furthermore, the notification message can be sent from the training system to the inference system, or it can be forwarded to the inference system through the scheduling function 403 provided by the AI service system (implemented by the scheduling system); this application does not limit this.
[0083] The parameter update function 4012 is used to perform multiple parameter updates of the AI model during the inference task based on the model's hierarchical information. Each update retrieves model parameters from one of multiple parts of the training system and updates the corresponding model parameters in the inference system. In other words, the inference system does not need to pause its inference operations; instead, it updates the model parameters of each part of the AI model sequentially based on the hierarchical information. The inference system and training system can transmit model parameters via a parameter acquisition link. This link is established between the two systems. For example, after training the AI model, the training system sends a notification message to the inference system, triggering the inference system to establish a parameter acquisition link with the training system based on the parameter storage address (e.g., network IP address) of the AI model. The inference system then retrieves the model parameters of each part of the AI model sequentially from the training system via this link.
[0084] For example, an AI model may consist of 20 layers. The model layering information indicates that the AI model is divided into four parts, each of which contains five layers. When the inference system performs an inference task using the old parameters of the AI model, it obtains new parameters for layers 1-5 from the training system and updates the model parameters for layers 1-5. Then, when the inference system performs an inference task using the new parameters for layers 1-5, it obtains new parameters for layers 6-10 from the training system and updates the model parameters for layers 6-10, and so on, until the inference system obtains all the new parameters of the AI model, thus completing the parameter update of the AI model.
[0085] Furthermore, as described above, when the AI service system involves multiple AI models, these multiple AI models include a first AI model and a second AI model. Based on this, the parameter update function 4012 is used to perform multiple parameter updates for the first AI model during the inference task execution based on the first AI model, according to the model hierarchy information of the first AI model. Each time, model parameters of a portion of the multiple parts of the first AI model are obtained from the training system and the corresponding model parameters in the inference system are updated. Similarly, during the inference task execution based on the second AI model, multiple parameter updates are performed for the second AI model according to the model hierarchy information of the second AI model. Each time, model parameters of a portion of the multiple parts of the second AI model are obtained from the training system and the corresponding model parameters in the inference system are updated. In other words, the inference system can obtain the model parameters of the corresponding model from the training system separately for each deployed AI model, avoiding task conflicts.
[0086] In some embodiments, the inference function 401 further includes a layer information determination function 4013. Illustratively, the layer information determination function 4013 is used to determine model layer information based on the number of layers in the AI model, the inference time of the AI model, the inference time of each layer in the AI model, and the transmission time of model parameters for each layer in the AI model. That is, the model layer information is determined by comprehensively considering multiple dimensions, making the model layering more consistent with the actual operating state, dynamically adapting to different model structures and performance requirements, optimizing the rationality of the layering strategy, and further improving parameter update efficiency. The implementation method of this process is described in subsequent method embodiments and will not be repeated here.
[0087] The training function 402 is implemented by the training system in the AI service system. It is used to train the AI models involved in the AI service system and promptly send a notification message to the inference system when the training is completed.
[0088] In some embodiments, the AI service system also provides a scheduling function 403, implemented by the scheduling system within the AI service system. This facilitates the coordination of work at each stage throughout the entire model training and inference process, avoids information transmission chaos, and ensures that the parameter updates of the inference system are triggered promptly after the training system completes training, achieving seamless integration between training and inference. Illustratively, the flow of information such as the aforementioned notification messages and the parameter storage address of the AI model between the inference system and the training system is all achieved through the scheduling function 403.
[0089] It should be noted that the functional division of the AI service system is not limited to... Figure 4 The content shown can be further customized to include more features in practical applications, such as storage functions for storing training samples. This application does not limit the scope of the application.
[0090] Schematic illustration: The functions provided by the aforementioned AI service system can be installed as a software toolkit component in computing devices and run by the processors and accelerators within those devices. For example, when the hardware architecture of a computing device employs a heterogeneous computing architecture (including architectures using computing units with different instruction sets), users can install heterogeneous computing frameworks on the device. One such framework is the Compute Architecture for NeuroNet (CANN), a heterogeneous computing framework for neural networks. CANN can support users in quickly building AI applications by providing multi-layered programming interfaces. Additionally, users can install deep learning frameworks on computing devices to compile methods for implementing models, construct large-scale computational graphs, and automatically perform gradient calculations within those graphs. The functions provided by the aforementioned AI service system can be interfaced and adapted with deep learning frameworks and heterogeneous computing frameworks.
[0091] The following is for reference. Figure 5 The method embodiment shown takes the parameter update process of any AI model in the AI service system as an example to introduce the parameter update method of the AI model provided in this application.
[0092] Figure 5 This is a flowchart of a parameter update method for an AI model provided in an embodiment of this application. For example... Figure 5 As shown, the method is executed by an AI service system, which includes a training system and an inference system. The training system is used to train an AI model, and the inference system is used to perform inference tasks based on the AI model. The method includes the following steps 501 to 505.
[0093] 501. After the AI model training is completed, the training system sends a notification message to the inference system, indicating that the training system has trained the AI model.
[0094] In this application's embodiments, the AI model is, for example, a Large Language Model (LLM) based on the Transformer architecture. In some embodiments, the AI model is a first AI model used to perform inference tasks to generate inference results; or, the AI model is a second AI model used to evaluate the inference results to generate evaluation results, and this application does not limit the specific model used.
[0095] In some embodiments, the training system sends a notification message to the scheduling system in the AI service system. The scheduling system receives the notification message from the training system and forwards it to the inference system. Through the scheduling system, the work of each stage is coordinated throughout the entire model training and inference process, avoiding message transmission chaos and ensuring that the parameters of the inference system are updated promptly after the training system completes training, thus achieving seamless integration of training and inference.
[0096] 502. The inference system receives notification messages.
[0097] The inference system obtains notification messages from the training system or from the scheduling system; this application does not limit this to either.
[0098] 503. The inference system obtains model layering information, which indicates multiple parts of the AI model, each part including at least one layer of the AI model.
[0099] In this embodiment, the AI model comprises multiple layers. The inference system determines the model layering information based on the number of layers in the AI model, the inference time of the AI model, the inference time of each layer in the AI model, and the transmission time of the model parameters of each layer in the AI model. That is, the model layering information is determined by comprehensively considering multiple dimensions, making the model layering more consistent with the actual operating state, dynamically adapting to different model structures and performance requirements, optimizing the rationality of the layering strategy, and further improving the efficiency of parameter updates.
[0100] Taking the large language model LLM based on the Transformer architecture as an example, the layers of the AI model are usually Transformer layers. The inference system determines the model layer information based on the parameter sizes of the attention layer and feedforward neural network layer of the Transformer layer in the AI model deployed on the inference system. The process of the inference system determining the model layer information is introduced below with reference to formulas (1) to (5).
[0101] For any Transformer layer in the AI model, the inference system determines the parameter size of the attention layer based on the weight matrices of the query (Q) matrix, key (K) matrix, value (V) matrix, and output (O) matrix in the attention layer of the Transformer layer, as shown in the following formula (1):
[0102] sum atten =W K +W V +W Q +W O (1)
[0103] In formula (1), W K W V W Q W O These represent the parameter sizes of the weight matrices for the Q, K, V, and O matrices in the attention layer, respectively.
[0104] The inference system determines the parameter size of the FFN layer based on the weight matrices of the up-sampling layer and the down-sampling layer in the feedforward neural network of the Transformer layer, as shown in the following formula (2):
[0105] sum mlp =W up +W down (2)
[0106] In formula (2), W up W down These represent the parameter sizes of the weight matrices for the up-sampling layer and the down-sampling layer in the FFN layer, respectively.
[0107] Based on the parameter sizes of the attention layer and the FFN layer in the Transformer layer, the transmission time of the model parameters in each layer of the AI model is determined as shown in the following formula (3):
[0108] T TBLayer =(sum mlp +sum atten ( / (B comm ×η comm (3)
[0109] In formula (3), B comm η is the communication bandwidth for transmitting model parameters from the training system to the inference system. comm This refers to the communication efficiency of transmitting model parameters from the training system to the inference system.
[0110] Based on the inference time of the AI model and the number of layers in the AI model, the inference time of each layer in the AI model is determined as shown in the following formula (4):
[0111] T CompLayer =T Compall / L (4)
[0112] In formula (4), T Compall L represents the inference time of the AI model (i.e., the computation time from the input layer to the output layer of the AI model), and L represents the number of layers in the AI model.
[0113] After determining the transmission time of model parameters and the inference time of each layer in the AI model, the number of layers contained in each part of the AI model is determined based on the time difference between transmission and computation. For example, it can be determined by the following formula (5):
[0114] x×T TBLayer ≤y×T compLayer (5)
[0115] In formula (5), x is the number of layers in the AI model, x×T TBLayer The transmission time of model parameters for layer x is given by y, where y is a reference coefficient for measuring computation time, and y×T is the transmission time of model parameters for layer x. compLayer Let x be the inference time of layer y, where x and y are both positive integers. Illustratively, if x and y satisfy the conditions shown in formula (5), it is more advantageous in terms of time cost to obtain the model parameters of layer x each time. In this case, selecting layer x as the layer number of each part in the AI model can avoid the transmission time from exceeding the computation time, thus ensuring the overall efficiency of the system.
[0116] Furthermore, this application does not limit the timing at which the inference system determines model hierarchy information. For example, the inference system may determine model hierarchy information before receiving the notification message, or it may determine model hierarchy information after receiving the notification message, and so on.
[0117] 504. Establish a parameter acquisition link between the inference system and the training system.
[0118] In this embodiment, after learning that the training system has trained the AI model, the inference system can establish a parameter acquisition link with the training system based on the parameter storage address of the AI model in the training system. For example, the training system sends the parameter storage address to the inference system, or the training system sends the parameter storage address to the inference system through the scheduling system. Furthermore, this application does not limit the timing of the inference system acquiring the parameter storage address. The training system can provide the parameter storage address to the inference system at any time before step 504, or the inference system can request the parameter storage address from the training system or the scheduling system after receiving the notification message, and so on.
[0119] In illustration, the parameter storage address is a network IP address. The inference system establishes a parameter acquisition link with the training system based on the parameter storage address of the AI model in the training system. For example, the inference system sends a link establishment request to the training system based on the parameter storage address. The training system verifies the identity of the inference system based on the link establishment request. After successful verification, both parties establish a parameter acquisition link according to the agreed communication protocol (such as HTTP, gRPC, TCP / IP, etc.), providing a channel for subsequent transmission of model parameters.
[0120] 505. During the inference process based on the AI model, the inference system performs multiple parameter updates of the AI model through the parameter acquisition link based on the model's hierarchical information. Each time, it obtains the model parameters of a part of multiple parts from the training system and updates the corresponding part of the model parameters in the inference system.
[0121] In this embodiment, taking any two adjacent parts among the multiple parts indicated by the model hierarchical information as an example (hereinafter referred to as the first part and the second part), the inference system obtains the model parameters of the first part from the training system and updates the model parameters of the first part in the inference system. During the execution of the inference task based on the updated first part, the system obtains the model parameters of the second part from the training system and updates the model parameters of the second part in the inference system. In this way, the process of the inference system executing the inference task and the process of the inference system updating the parameters are executed in parallel, realizing uninterrupted online model updates for inference services and effectively avoiding resource waste. The inference system can obtain the model parameters of the first part from the training system after deleting the model parameters of the first part deployed in the inference system (i.e., the old model parameters), or it can obtain the model parameters of the first part and then delete the model parameters of the first part deployed in the inference system (i.e., the old model parameters). This application does not limit this.
[0122] In some embodiments, the duration for which the inference system performs an inference task based on the updated first part is greater than or equal to the duration for which the inference system obtains the model parameters of the second part from the training system. Illustratively, the inference system obtains the model parameters of the first part from the training system and updates the model parameters of the first part deployed in the inference system; during the execution of the inference task based on the updated first part, it obtains the model parameters of the second part from the training system and updates the model parameters of the second part in the inference system; after performing the inference task based on the updated first part, it performs the inference task based on the updated second part.
[0123] In this way, since the time taken for the inference system to execute the inference task based on the first part is greater than or equal to the time taken for the inference system to acquire the model parameters for the second part, the inference system has already acquired the model parameters for the second part after executing the inference task based on the first part. Therefore, it can continue executing the inference task based on the updated second part, avoiding waiting due to excessive parameter acquisition time. This improves parameter update efficiency while ensuring inference efficiency. In other words, during the period when the inference system is acquiring the next part of the model parameters, the execution time of the currently updated inference task can cover the transmission time of the next part of the model parameters, thus avoiding resource idleness or conflicts. For example, if an AI model includes 20 layers, and the model layering information indicates four parts of the AI model, each part including five layers, the inference time for updated layers 1-5 is 10 seconds, and the parameter transmission time for layers 6-10 is 8 seconds, after the inference system executes the inference task based on updated layers 1-5, the parameters for layers 6-10 have been updated, allowing the inference system to continue executing the inference task based on updated layers 6-10.
[0124] In summary, in the parameter update method for the AI model provided in this application, when the inference system knows that the training system has trained the AI model, it can perform multiple parameter updates of the AI model according to multiple parts of the AI model indicated by the model hierarchical information during the execution of the inference task based on the AI model. In each update, the model parameters of one part of the multiple parts are obtained from the training system and the model parameters of the corresponding part are updated. Thus, the model parameters are gradually updated during the inference process, realizing online model updates for inference business without interruption and effectively avoiding resource waste.
[0125] The following section provides examples illustrating the parameter update method for the AI model provided in this application, using various functions offered by the AI service system.
[0126] Based on the foregoing introduction to AI service systems, it is known that an AI service system includes a scheduling system, an inference system, and a training system. The inference system includes an inference scheduling system and an inference execution system, while the training system includes a training scheduling system and a training execution system. Furthermore, the AI models involved in the AI service system can include a first AI model and a second AI model. Based on this, refer to... Figures 6 to 7 This section introduces the overall operation process of the AI service system. Among other things, Figure 6 This is a flowchart illustrating the operation of an AI service system provided in an embodiment of this application. Figure 7 This is a flowchart illustrating the operation of an inference system provided in an embodiment of this application. Figure 6 As shown, the operation process of the AI service system includes the following steps S1 to S22.
[0127] S1. The scheduling system initiates training and inference. Specifically, the scheduling system sends training requests to the training scheduling system within the training system, and inference requests to the inference scheduling system within the inference system.
[0128] S2. The training scheduling system in the training system starts the training tasks of the first AI model (e.g., G model) and the second AI model (e.g., V model) according to the training request.
[0129] S3. The training execution system of the first AI model and the training execution system of the second AI model obtain training samples from the training sample database and start training to update the model parameters.
[0130] S4. The inference scheduling system in the inference system performs initialization processing according to the inference request, such as allocating affinity hardware and obtaining the initial model file.
[0131] S5, the inference scheduling system deploys inference instances on the inference execution system.
[0132] S6. The inference execution system deploys instances of the first AI model and the second AI model, including loading model parameters.
[0133] S7. The inference scheduling system begins executing inference tasks based on predefined problems or task descriptions.
[0134] S8. The inference execution system uses the first AI model, combined with algorithms such as MCTS, BON, or serial revision method, to generate inference results with multiple branches.
[0135] S9. The inference execution system evaluates the inference results output by the first AI model from multiple dimensions, including process and outcome, using the second AI model, and obtains the evaluation results. S11. Illustratively, the inference system can determine the final inference result output to the user based on the evaluation results.
[0136] S10. The inference execution system uses the second AI model to process multiple branches of the first AI model in a loop, and obtains the evaluation result of the inference result of each branch.
[0137] S11. The inference execution system feeds back the predefined problem, inference results, and evaluation results as samples to the scheduling system.
[0138] S12. The scheduling system stores the received samples into the training sample database.
[0139] S13. The scheduling system obtains the parameter storage addresses of the first AI model and the second AI model from the training scheduling system, such as network IP addresses.
[0140] S14. The training scheduling system returns the parameter storage addresses of the first AI model and the second AI model.
[0141] S15. The scheduling system sends the parameter storage addresses of the first AI model and the second AI model to the inference execution system through the inference scheduling system.
[0142] S16. The inference scheduling system determines the model layering information of the AI model and sends the model layering information to the inference execution system. Specifically, the inference scheduling system determines the model layering information of the first AI model and the model layering information of the second AI model.
[0143] S17. After the first AI model and the second AI model have been trained, the training execution system of the first AI model and the training execution system of the second AI model send a notification message to the scheduling system.
[0144] S18. The scheduling system sends notification messages to the inference execution system through the inference scheduling system.
[0145] S19. In the inference execution system, the deployment instances of the first and second AI models establish parameter acquisition links with the training execution system based on the parameter storage addresses of the first and second AI models, respectively. Illustratively, the inference execution system establishes a first parameter acquisition link with the training execution system based on the parameter storage address of the first AI model, and a second parameter acquisition link with the training execution system based on the parameter storage address of the second AI model. That is, the inference system can establish parameter acquisition links with the training system separately for each deployed AI model and obtain the corresponding model parameters from the training system, avoiding task conflicts.
[0146] S20. Based on the model hierarchical information, the inference execution system deletes the model parameters of the first part of the first AI model and the model parameters of the first part of the second AI model.
[0147] S21. The inference execution system obtains new model parameters of the first part of the first AI model through the first parameter acquisition link between the inference execution system and the training execution system, and obtains new model parameters of the first part of the second AI model through the second parameter acquisition link between the inference execution system and the training execution system.
[0148] S22. The inference execution system updates the model parameters of the first part of the first AI model and the model parameters of the first part of the second AI model based on the obtained model parameters.
[0149] Then, repeat steps S20 to S22 until the model parameters of the first AI model and the second AI model are updated.
[0150] In some embodiments, the functions provided by the inference execution system are implemented by the following components: model management, text generation, model execution, computation unit, and post-processing. Model management is responsible for state management, task scheduling (such as batch processing of tasks based on scheduling strategies), unified memory pool management, KV cache management, and low-rank adaptation (LoRA) weight management. Model management also provides secondary development interfaces such as state monitoring. Text generation is responsible for model configuration, initialization, loading, and autoregressive inference processes, providing a unified autoregressive inference interface for model management. Model execution provides deeply optimized modules and built-in models, possessing model compilation and optimization capabilities to improve model running efficiency. The computation unit, as the actual execution entity (such as an NPU card or module), undertakes specific computation tasks. Post-processing is responsible for processing the model's output results, including token assembly, converting prompts, and returning the final output.
[0151] Combining the functions provided by the inference execution system, the above Figure 6The execution flow of the inference execution system shown can be referenced. Figure 7 It includes the following stages:
[0152] Phase 1: Model Management receives model hierarchical information from the inference scheduling system and distributes the model hierarchical information to the computing units.
[0153] Phase 2: Model Management receives notification messages from the training execution systems of the first AI model and the second AI model.
[0154] Phase 3: Model Management. Based on the parameter storage addresses of the first and second AI models, parameter acquisition links are established with the training execution system respectively. Once the link is successfully established, the computing unit is notified.
[0155] Phase 4: The calculation unit deletes the model parameters of the first part of the first AI model and the model parameters of the first part of the second AI model based on the model hierarchical information.
[0156] Phase 5: The computing unit obtains new model parameters for the first part of the first AI model and updates the corresponding model parameters through the first parameter acquisition link between the computing unit and the training execution system, and obtains new model parameters for the first part of the second AI model and updates the corresponding model parameters through the second parameter acquisition link between the computing unit and the training execution system.
[0157] Phase 6: After the computing unit updates the model parameters of the first part of the first AI model and the first part of the second AI model, it deletes the model parameters of the second part of the first AI model and the second part of the second AI model. This process is repeated until the model parameters are updated.
[0158] As can be seen, through the above-described AI service system operation process, after the inference system learns that the training system has trained the first AI model and the second AI model, during the inference process based on the AI model, according to the multiple parts of the AI model indicated by the model layering information, the inference system sequentially obtains the model parameters of each part from the training system through the parameter acquisition link between the inference system and the training system. Thus, the model parameters are updated gradually and asynchronously during the inference process, realizing uninterrupted online model updates for inference business and effectively avoiding resource waste.
[0159] Based on the parameter update method of the aforementioned AI model, this application also provides an inference device, which can realize some or all of the functions of the inference system in the aforementioned AI service system through software, hardware, or a combination of both. (Reference) Figure 8 , Figure 8 This is a schematic diagram of the structure of a reasoning device provided in an embodiment of this application. Figure 8As shown, the inference device includes an acquisition module 801 and a parameter update module 802.
[0160] The acquisition module 801 is used to acquire notification messages and model layer information. The notification message indicates that the AI model has finished training, and the model layer information indicates multiple parts of the AI model, each part including at least one layer of the AI model.
[0161] The parameter update module 802 is used to perform multiple parameter updates of the AI model based on the model's hierarchical information during the inference task performed by the AI model. Each time, the model parameters of a part of the multiple parts are obtained and the corresponding part of the model parameters in the inference device are updated.
[0162] In some embodiments, the parameter update module 802 is configured to: obtain model parameters of a first part among a plurality of parts and update the model parameters of the first part in the inference device; and, during the execution of an inference task based on the updated first part, obtain model parameters of a second part among a plurality of parts and update the model parameters of the second part in the inference device.
[0163] In some embodiments, the duration of performing the inference task based on the updated first part is greater than or equal to the duration of obtaining the model parameters of the second part; the inference apparatus further includes an inference task execution module 803, configured to: perform an inference task based on the updated second part after performing the inference task based on the updated first part.
[0164] In some embodiments, the inference device further includes a link establishment module 804, configured to: after obtaining notification messages and model hierarchical information, establish a parameter acquisition link with the training system based on the parameter storage address of the AI model in the training system; and a parameter update module, configured to obtain model parameters of a portion of multiple parts from the training system through the parameter acquisition link.
[0165] In some embodiments, the inference device further includes a layer information determination module 805, configured to: determine model layer information based on the number of layers in the AI model, the inference time of the AI model, the inference time of each layer in the AI model, and the transmission time of model parameters of each layer in the AI model.
[0166] In some embodiments, the AI model is a first AI model or a second AI model, and the inference device is further configured to: perform an inference task based on the first AI model, generate an inference result, and evaluate the inference result based on the second AI model to generate an evaluation result.
[0167] When the aforementioned inference device learns that the AI model has finished training, it can perform multiple parameter updates of the AI model according to the multiple parts of the AI model indicated by the model's hierarchical information during the inference task. Each time, it obtains the model parameters of one part of the multiple parts and updates the model parameters of the corresponding part, thereby gradually updating the model parameters during the inference process. This achieves uninterrupted online model updates for inference operations and effectively avoids resource waste.
[0168] Of course, the inference device can also include other functional units to implement other functions involved in the inference system in the above method embodiments. In practical applications, the above functions can be assigned to different functional units as needed, that is, the internal structure of the device can be divided into different functional units to complete all or part of the functions described above. In addition, the inference device provided in the above embodiments and the above method embodiments belong to the same concept, and its specific implementation process is detailed in the method embodiments, which will not be repeated here.
[0169] In this application, the terms "first," "second," etc., are used to distinguish identical or similar items that have substantially the same function and purpose. It should be understood that there is no logical or temporal dependency between "first," "second," and "nth," nor does it limit the quantity or order of execution. It should also be understood that although the following description uses the terms "first," "second," etc., to describe various elements, these elements should not be limited by the terms. These terms are merely used to distinguish one element from another. For example, without departing from the various examples described, a first part can be referred to as a second part, and similarly, a second part can be referred to as a first part. Both the first and second parts can be parts, and in some cases, can be separate and distinct parts.
[0170] In this application, the term "at least one" means one or more, and the term "multiple" means two or more. For example, multiple parts means two or more parts.
[0171] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0172] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, in the form of program structure information. This program structure information includes one or more program instructions. When these program instructions are loaded and executed on a computing device, the processes or functions according to the embodiments of this application are generated, in whole or in part.
[0173] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0174] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A parameter update method for an AI model, characterized in that, The method is applied to an AI service system, which includes a training system and an inference system. The training system is used to train an AI model, and the inference system is used to perform inference tasks based on the AI model. The method includes: The inference system acquires a notification message and model layering information. The notification message indicates that the training system has trained the AI model, and the model layering information indicates multiple parts of the AI model, each part including at least one layer of the AI model. During the inference process of performing inference tasks based on the AI model, the inference system performs multiple parameter updates of the AI model according to the model hierarchical information. Each time, it obtains the model parameters of a portion of the multiple parts from the training system and updates the model parameters of the corresponding part in the inference system.
2. The method according to claim 1, characterized in that, During the inference process, the inference system performs multiple parameter updates of the AI model based on the model's hierarchical information, including: The inference system obtains the model parameters of the first part of the plurality of parts from the training system and updates the model parameters of the first part in the inference system. During the process of performing an inference task based on the updated first part, the inference system obtains the model parameters of the second part from the plurality of parts from the training system and updates the model parameters of the second part in the inference system.
3. The method according to claim 2, characterized in that, The duration for which the inference system performs the inference task based on the updated first part is greater than or equal to the duration for which the inference system obtains the model parameters of the second part from the training system; The method further includes: After performing a reasoning task based on the updated first part, the reasoning system performs a reasoning task based on the updated second part.
4. The method according to any one of claims 1 to 3, characterized in that, After the inference system obtains the notification message and the model hierarchical information, the method further includes: The inference system establishes a parameter acquisition link with the training system based on the parameter storage address of the AI model in the training system; The inference system obtains model parameters of a portion of the plurality of parts from the training system, including: the inference system obtains model parameters of a portion of the plurality of parts from the training system through the parameter acquisition link.
5. The method according to claim 4, characterized in that, The AI service system further includes a scheduling system, which is used to schedule the training system and the inference system. The method further includes: The scheduling system sends the parameter storage address to the inference system; The scheduling system receives the notification message sent by the training system and forwards the notification message to the inference system.
6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: The inference system determines the model layering information based on the number of layers in the AI model, the inference time of the AI model, the inference time of each layer in the AI model, and the transmission time of the model parameters of each layer in the AI model.
7. The method according to any one of claims 1 to 6, characterized in that, The AI model is either a first AI model or a second AI model, and the method further includes: The inference system performs an inference task based on the first AI model, generates an inference result, evaluates the inference result based on the second AI model, generates an evaluation result, and sends the inference result and the evaluation result to the training system. The training system trains the first AI model and the second AI model based on the inference results and the evaluation results.
8. An AI service system, characterized in that, The AI service system includes a training system and an inference system. The training system is used to train an AI model, and the inference system is used to perform inference tasks based on the AI model. The AI service system is used to implement the parameter update method of the AI model as described in any one of claims 1 to 7.
9. A reasoning device, characterized in that, The reasoning device includes: The acquisition module is used to acquire notification messages and model layering information. The notification messages indicate that the AI model has finished training, and the model layering information indicates multiple parts of the AI model, each part including at least one layer of the AI model. The parameter update module is used to perform multiple parameter updates of the AI model based on the model hierarchical information during the inference task performed according to the AI model. Each time, the model parameters of a portion of the multiple parts are obtained and the model parameters of the corresponding part in the inference device are updated.
10. The reasoning device according to claim 9, characterized in that, The parameter update module is used for: Obtain the model parameters of the first part of the plurality of parts and update the model parameters of the first part in the inference device; During the inference task performed based on the updated first part, the model parameters of the second part of the plurality of parts are obtained and the model parameters of the second part in the inference device are updated.
11. The reasoning device according to claim 10, characterized in that, The duration of the inference task performed in the first part after the update is greater than or equal to the duration of obtaining the model parameters in the second part; The inference device further includes an inference task execution module, used for: After performing the reasoning task based on the updated first part, perform the reasoning task based on the updated second part.
12. The reasoning device according to any one of claims 9 to 11, characterized in that, The inference device further includes a link establishment module, used to: after obtaining the notification message and the model hierarchical information, establish a parameter acquisition link with the training system according to the parameter storage address of the AI model in the training system; The parameter update module is used to obtain model parameters of a portion of the multiple parts from the training system through the parameter acquisition link.
13. The reasoning device according to claim 12, characterized in that, The reasoning device further includes a hierarchical information determination module, used for: The model layering information is determined based on the number of layers in the AI model, the inference time of the AI model, the inference time of each layer in the AI model, and the transmission time of the model parameters of each layer in the AI model.
14. The reasoning device according to any one of claims 9 to 13, characterized in that, The AI model is either a first AI model or a second AI model, and the inference device is further used for: The first AI model performs a reasoning task to generate a reasoning result, and the second AI model evaluates the reasoning result to generate an evaluation result.
15. A computing device, characterized in that, The computing device includes a processor and a memory, the processor being configured to execute at least one piece of program code stored in the memory to enable the computing device to implement the parameter update method of the AI model as described in any one of claims 1 to 7.
16. A computer program product, characterized in that, The computer program product is used to implement the parameter update method of the AI model as described in any one of claims 1 to 7.
17. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store at least one piece of program code, which is used to implement the parameter update method of the AI model as described in any one of claims 1 to 7.