Model Deployment Scheme Generation, Model Processing Method, Device and Electronic Device

By generating a deployment scheme that allocates model weights to specific processors based on device capabilities, the method addresses the challenge of deploying large AI models on endpoint devices, improving efficiency and security.

CN119396413BActive Publication Date: 2025-07-15ZHONG KE JIA HE (BEI JING) KE JI YOU XIAN GONG SI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411303641.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-18
Publication Date
2025-07-15
Estimated Expiration
2044-09-18

AI Technical Summary

Technical Problem

In the prior art, large-scale artificial intelligence models are mainly deployed on servers with strong processing capabilities, resulting in difficulties in deploying on other devices, especially in poor network environments or offline environments that cannot provide services in a timely manner and pose information security risks.

Method used

By obtaining the model information of the model to be deployed and the processor parameters of the target device, generating a deployment plan, specifying the processor deployed by the model weights of each processing layer, and reasonably allocating storage resources and computing resources to achieve the rational deployment of the model on the target device.

Benefits of technology

The rational allocation of processor resources on the target device is achieved, the difficulty of model deployment is solved, and the availability and security of the model in poor network environments or offline environments is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119396413B_ABST
    Figure CN119396413B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a method and apparatus for generating a model deployment plan, model processing, and an electronic device, which relate to technical fields such as model deployment and large language models. The specific implementation solution is as follows: Obtain the processor parameters of each processor in the target device, where the target device is used to deploy the model to be deployed; generate a deployment plan based on the model information and the processor parameters, the deployment plan includes first specified information, the model to be deployed includes multiple processing layers, and the first specified information is used to specify the processor on which the model weights of each processing layer are deployed. Based on the deployment plan provided by the embodiments of the present application, it is possible to achieve reasonable allocation of processor resources in the target device, thereby achieving reasonable deployment of the model in the target device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to technical fields such as model deployment and large language models. Specifically, this application relates to a method, apparatus, and electronic device for generating a model deployment plan and processing a model. Background Art

[0002] In recent years, with the rapid development of deep learning and artificial intelligence technologies, the capabilities of artificial intelligence models have become increasingly powerful, and at the same time, the scale of artificial intelligence models has also become increasingly large.

[0003] Currently, large-scale artificial intelligence models are mainly deployed on servers with powerful processing capabilities, but there are difficulties in deploying them on other devices. Therefore, how to reasonably deploy models has become an important technical issue. Summary of the Invention

[0004] This application provides a method, apparatus, and electronic device for generating a model deployment plan and processing a model, aiming to at least solve one of the above technical defects. The technical solutions adopted in this application are as follows:

[0005] In a first aspect, an embodiment of this application provides a method for generating a model deployment plan. The method includes:

[0006] Obtain the model information of the model to be deployed;

[0007] Obtain the processor parameters of each processor in the target device, where the target device is used to deploy the model to be deployed;

[0008] Based on the model information and the processor parameters, generate a deployment plan. The deployment plan includes first specified information. The model to be deployed includes multiple processing layers, and the first specified information is used to specify the processor to which the model weights of each processing layer are deployed.

[0009] In a second aspect, an embodiment of this application provides a method for deploying a model. The method includes:

[0010] Obtain the deployment plan for the model to be deployed. The deployment plan includes first specified information. The model to be deployed includes multiple processing layers. The target device for deploying the model to be deployed includes multiple processors. The first specified information is used to specify the processor to which the model weights of each processing layer are deployed. The model deployment plan is generated by a server based on the model information of the model to be deployed and the processor parameters of the processor;

[0011] Based on the first specified information, deploy the model weights of each processing layer to the specified processor in the target device.

[0012] In a third aspect, an embodiment of this application provides a device for generating a model deployment plan. The device includes:

[0013] A model information acquisition module, configured to acquire model information of a model to be deployed;

[0014] A processor parameter acquisition module, configured to acquire processor parameters of each processor in a target device, where the target device is used to deploy the model to be deployed;

[0015] A deployment plan generation module, configured to generate a deployment plan based on the model information and the processor parameters, where the deployment plan includes first specified information, the model to be deployed includes multiple processing layers, and the first specified information is used to specify the processor to which the model weights of each processing layer are deployed.

[0016] In a fourth aspect, an embodiment of the present application provides a model processing device, which includes:

[0017] A deployment plan acquisition module, configured to acquire a deployment plan for the model to be deployed, where the deployment plan includes first specified information, the model to be deployed includes multiple processing layers, the target device for deploying the model to be deployed includes multiple processors, the first specified information is used to specify the processor to which the model weights of each processing layer are deployed, and the model deployment plan is generated by a server based on the model information of the model to be deployed and the processor parameters of the processor;

[0018] A model deployment module, configured to deploy the model weights of each processing layer to the specified processor in the target device based on the first specified information.

[0019] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the above method are implemented.

[0020] In a sixth aspect, an embodiment of the present application provides an electronic device, which includes:

[0021] One or more processors; and

[0022] A memory associated with the above one or more processors, where the memory is used to store program instructions, and when the program instructions are read and executed by the above one or more processors, the steps of the above method are executed.

[0023] In a seventh aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0024] The beneficial effects brought by the technical solutions provided by the embodiments of the present application are:

[0025] The solution provided by the embodiments of the present application obtains the model information of the model to be deployed, and obtains the processor parameters of each processor in the target device for deploying the model to be deployed. Then, based on the model information and the processor parameters, a deployment plan is generated. The deployment plan includes first specification information for specifying the processors where the model weights of each processing layer are deployed, and second specification information for specifying the processors that execute the model calculations corresponding to each processing layer. Based on the deployment plan provided by the embodiments of the present application, reasonable allocation of processor resources in the target device can be achieved, thereby realizing the reasonable deployment of the model in the target device. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for description in the embodiments of the present application.

[0027] Figure 1 It is a system architecture diagram applicable to the embodiments of the present application;

[0028] Figure 2 It is a flowchart of a method for generating a model deployment plan provided by the embodiments of the present application;

[0029] Figure 3 It is a flowchart of a specific implementation manner of the method for generating a model deployment plan provided by the embodiments of the present application;

[0030] Figure 4 It is a flowchart during the deployment of model weights in the embodiments of the present application;

[0031] Figure 5 It is a flowchart during the model inference calculation in the embodiments of the present application;

[0032] Figure 6 It is a flowchart of a model processing method provided by the embodiments of the present application;

[0033] Figure 7 It is a structural diagram of a model deployment plan generation device provided by the embodiments of the present application;

[0034] Figure 8 It is a structural diagram of a model processing device provided by the embodiments of the present application;

[0035] Figure 9 It is a structural diagram of an electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.

[0037] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application. The singular forms "a", "the" and "said" used in the embodiments of the present application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.

[0038] It should be understood that the term "and / or" used herein is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.

[0039] Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" can be interpreted as "when determined" or "in response to determining" or "when detected (stated condition or event)" or "in response to detecting (stated condition or event)".

[0040] First of all, several terms related to the present application will be introduced and explained:

[0041] The large language model (Large Language model, LLM) is a deep learning model for natural language-related tasks. It can receive text content input and return corresponding outputs. The specific tasks that can be completed include text generation, classification, summarization, rewriting, etc.

[0042] When the LLM processes natural language tasks internally, it can be divided into two different stages: the understanding stage and the generation stage.

[0043] In the understanding stage, the LLM "understands" the meaning of the text input by the user through calculation, manifested as calculating the state tensor (KV-Cache) inside the LLM and the first predicted element (token).

[0044] In the generation stage, the LLM uses the state tensor calculated in the language understanding stage as a parameter to continuously update the KVcache and generate a new next token.

[0045] The overall processing flow of LLM in the understanding stage and the generation stage is the same. The difference between the two stages is that the understanding stage processes multiple or all input data at a time and generates the first token, while the generation stage processes one token at a time.

[0046] The central processing unit (CPU) is a large-scale integrated circuit, which is the computing core and control core of a computer. Its main function is to interpret computer instructions and process data in computer software.

[0047] A graphics processing unit (GPU) is a processor specifically used for image and video processing. GPU was first used in the field of computer games and is currently widely used in the fields of big data analysis and artificial intelligence.

[0048] A neural network processor (NPU) is a specialized processor designed specifically for the computing needs of neural network models. It is designed to perform machine learning, especially deep learning tasks, efficiently and with low power consumption. As an artificial intelligence (AI) accelerator, NPU is a customized circuit for AI acceleration and contains some necessary control units and algorithms to execute deep learning algorithms.

[0049] A digital signal processor (DSP) is a microprocessor dedicated to (usually real-time) digital signal processing. The DSP chip uses a Harvard structure that separates program and data, has a dedicated hardware multiplier, widely uses pipeline operations, and provides special DSP instructions to quickly implement various digital signal processing algorithms. The characteristics of DSP chips include the ability to perform multiple operations in parallel, which makes DSP perform well when processing digital signals that require high-speed mathematical operations. Currently, DSP is also often used to perform deep learning model reasoning.

[0050] At present, the ability of artificial intelligence models to handle complex tasks is getting stronger and stronger. As a result, the number of model parameters is getting larger and larger, and the model's demand for computing power is also increasing, which greatly increases the demand for equipment when deploying models. At present, large-scale models are mainly deployed on servers with powerful processing capabilities, but there are difficulties in deploying them on other devices. Therefore, how to deploy the model reasonably has become an important technical issue.

[0051] Taking the Large Language Model (LLM), which has made breakthroughs in the field of natural language processing in recent years, as an example, the LLM has a huge number of parameters, usually containing tens of billions or hundreds of billions of parameters, which makes the LLM mainly deployed on servers with powerful video memory of Graphics Processing Units (GPUs). In this deployment method, users need to request LLM services through the network, which will lead to the following problems: (1) The LLM service cannot be provided in a timely manner in a poor network environment or an offline environment; (2) There will be certain information security risks in network transmission.

[0052] Deploying the LLM on the edge device can solve the above problems caused by requesting the LLM service through the network. However, the types of edge devices are diverse, and the performance of the processors in the edge devices also varies. The performance parameters of some of the processors cannot fully meet the requirements for deploying the LLM, which makes it difficult to deploy the LLM on the edge device.

[0053] In view of this, the present application provides a new idea. To facilitate the understanding of the present application, the system architecture on which the present application is based will be described first. Figure 1 An exemplary system architecture to which the embodiments of the present application can be applied is shown, as Figure 1 shown in, the system architecture may include: a target device and a server side.

[0054] Among them, the deployment scheme generation device on the server side can adopt the method provided in the embodiments of the present application to generate a deployment scheme based on the model information and processor parameters in the offline stage.

[0055] The target device can interact with the server side through the network. For example, the target device can send the processor parameters to the deployment scheme generation device. The server side can send the model resources and the deployment scheme to the target device when model deployment is required. The model processing device of the target device can implement model deployment based on the deployment scheme.

[0056] The deployment scheme generation device can be set in a server or a server group, and can also be set in an independent or the same cloud server. A cloud server, also known as a cloud computing server or a cloud host, is a host product in the cloud computing service system to solve the problems of high management difficulty and weak service scalability existing in traditional physical hosts and Virtual Private Servers (VPSs).

[0057] It should be understood, Figure 1The numbers of the target devices, deployment scheme generation devices, and model processing devices in [description] are merely illustrative. According to actual requirements, there can be any number of target devices, deployment scheme generation devices, and model processing devices.

[0058] Figure 2 FIG. [figure number] shows a schematic flowchart of a model deployment scheme generation method provided by an embodiment of the present application. This method can be executed by a deployment scheme generation device in the Figure 1 system shown in [figure number]. As shown in Figure 2 FIG. [figure number], this method mainly includes:

[0059] Step 210: Obtain the model information of the model to be deployed;

[0060] Step 220: Obtain the processor parameters of each processor in the target device, where the target device is used to deploy the model to be deployed;

[0061] Step 230: Generate a deployment scheme based on the model information and the processor parameters. The deployment scheme includes first specified information. The model to be deployed includes multiple processing layers, and the first specified information is used to specify the processor on which the model weights of each processing layer are deployed.

[0062] Among them, the model to be deployed may include, but is not limited to, large language models. The model to be deployed may include multiple processing layers, and the model weights of the model to be deployed are composed of the model weights corresponding to each processing layer. Based on the model weights corresponding to each processing layer, model calculations corresponding to each processing layer can be performed.

[0063] The model information of the model to be deployed can reflect the requirements for the processor performance during deployment.

[0064] In an embodiment of the present application, the model to be deployed will be deployed in the target device. The target device may include, but is not limited to, edge devices, terminal devices, etc. As a preferred embodiment, the target device in this case is an edge device, and the model service runs through the edge device.

[0065] The target device includes multiple processors. The processors in the target device can be of various types, including but not limited to CPUs, GPUs, NPUs, DSPs, etc. The processor parameters can reflect the hardware performance of the processors.

[0066] In an embodiment of the present application, the model to be deployed can run on the target device in the form of an executable program. The model deployment involved in the embodiment of the present application mainly refers to loading and storing the model weights into the processors of the target device.

[0067] Based on the model information and processor parameters, it is possible to rationally allocate the processors in the target device according to the requirements of each processing layer of the model to be deployed and the performance of each processor in the target device, so that the model to be deployed can be reasonably deployed on the target device.

[0068] Specifically, when deploying the model weights of each processing layer, the storage space situation of the processor is mainly considered. Therefore, by generating the first specified information, the model weights of each processing layer can be deployed to the processors whose storage space meets the requirements.

[0069] In the embodiment of the present application, by generating a deployment plan, it is possible to effectively allocate the resources in the processors of the target device, which helps to reasonably deploy the model to be deployed on the target device.

[0070] In the embodiment of the present application, the deployment plan can be generated offline and then sent to the target device, so that the target device can deploy the model to be deployed based on the deployment plan.

[0071] The method provided by the embodiment of the present application obtains the model information of the model to be deployed, obtains the processor parameters of each processor in the target device used to deploy the model to be deployed, and then generates a deployment plan based on the model information and processor parameters. The deployment plan includes the first specified information used to specify the processors where the model weights of each processing layer are deployed. Based on the deployment plan provided by the embodiment of the present application, it is possible to rationally allocate the processor resources in the target device, thereby realizing the reasonable deployment of the model on the target device.

[0072] In an alternative manner of the embodiment of the present application, the deployment plan further includes second specified information, and the second specified information is used to specify the processors that execute the model calculations corresponding to each processing layer.

[0073] When executing the model calculations corresponding to each processing layer, the computing power of the processor is mainly considered. Therefore, by generating the second specified information, it is possible to allocate the processors with computing power meeting the requirements to execute the model calculations corresponding to different processing layers.

[0074] It can be understood that for any processing layer of the model to be deployed, the processor used to deploy the model weights corresponding to this processing layer and the processor used to execute the model calculations of this processing layer can be the same or different.

[0075] Based on the deployment plan provided by the embodiment of the present application, it is possible to rationally allocate the storage resources and computing resources of the target device, thereby realizing the reasonable deployment of the model on the target device.

[0076] In the embodiments of the present application, the second specified information can be generated on the server side and sent to the target device, so as to reasonably specify the processor to execute the models of each processing layer. The second specified information can also be converted into execution logic and automatically run on the target device to realize automatically selecting the processor to execute the models of each processing layer.

[0077] In an alternative embodiment of the present application, the second specified information includes the second specified information in the understanding stage. The second specified information in the understanding stage is used to specify the processor that executes the model calculation corresponding to each processing layer in the understanding stage. Based on the model information and the processor parameters, a deployment plan is generated, including:

[0078] For any processing layer, determine the first duration corresponding to the first processor and the second duration corresponding to the second processor. Wherein, the model weight of this processing layer is deployed on the first processor, the computing power of the second processor is higher than that of the first processor, the first duration is the duration for the first processor to execute the model calculation corresponding to this processing layer, and the second duration is the sum of the transmission duration and the calculation duration. The transmission duration is the duration for transmitting the data required for executing the model calculation of this processing layer to the second processor, and the calculation duration is the duration for the second processor to execute the model calculation corresponding to this processing layer;

[0079] Based on the first duration and the second duration, determine the second specified information in the understanding stage.

[0080] In the embodiments of the present application, the model to be deployed can be other types of large models that include but are not limited to text-to-image large models, text-to-video large models, etc., in addition to large language models.

[0081] In the embodiments of the present application, the requirements for the processor performance in the processing of the understanding stage and the generation stage of the model to be deployed are not the same. For the requirements of the model to be deployed in the understanding stage, the second specified information in the understanding stage can be generated to specify the processor that executes the model calculation corresponding to each processing layer in the understanding stage.

[0082] In the understanding stage of the model to be deployed, multiple or all of the data to be processed are input at one time, and all the input data to be processed are processed at one time, which has certain requirements for both the storage space and the computing power of the processor. In actual situations, some processors of the target device may have poor computing power, while some other processors may have insufficient storage space. In response to these situations, by reasonably specifying the processor that executes the model calculation corresponding to each processing layer in the understanding stage, the model calculation efficiency can be effectively improved.

[0083] Specifically, for any processing layer, the processor that deploys the model weights corresponding to this processing layer can be denoted as the first processor, and the time taken by the first processor to process the model calculation corresponding to this processing layer is the first duration. The second processor is a processor in the target device with higher computing power than the first processor. When the second processor processes the model calculation corresponding to this processing layer, it is necessary to first transfer the data required for the model calculation of this processing layer to the second processor, and the time taken for this data transfer is denoted as the transfer duration. The actual second duration corresponding to the second processor is composed of the sum of the transfer duration and the calculation duration for the second processor to execute the model calculation corresponding to this processing layer.

[0084] In the embodiments of the present application, based on the first duration and the second duration, it is possible to determine the processor that can more efficiently complete the processing process corresponding to this processing layer, thereby effectively generating the second specified information in the understanding stage.

[0085] In an alternative manner of the embodiments of the present application, the second specified information includes the second specified information in the generation stage. The second specified information in the generation stage is used to specify that the processor for executing the model calculation corresponding to each processing layer in the generation stage is the first processor, and the first processor is the processor where the model weights corresponding to each processing layer are deployed.

[0086] In the embodiments of the present application, in the generation stage of the model to be deployed, one token is input and processed each time. The demand for processor computing power in the generation stage is relatively low. Therefore, the second specified information in the generation stage can be used to specify the first processor that deploys the model weights corresponding to a certain processing layer to execute the model calculation of this processing layer. That is to say, the second specified information in the generation stage indicates that in the generation stage, it is not necessary to transfer the model weights to a processor with higher computing power, nor to call a processor with higher computing power to execute the model calculation of this processing layer.

[0087] In the embodiments of the present application, since one token is input and processed each time in the generation stage, if the method for determining the processor for executing the model operation of each processing layer in the understanding stage is still followed, it will cause the model weights to be transferred relatively many times, resulting in a large amount of time consumed for the transfer of the model weights and affecting the model processing efficiency. Considering that the demand for processor computing power in the generation stage is relatively low, the processors used for model deployment generally can meet the demand for processor computing power in the generation stage. Therefore, it is possible to specify the first processor that deploys the model weights corresponding to a certain processing layer to execute the model calculation of this processing layer.

[0088] In an alternative manner of the embodiments of the present application, the processor parameters include the memory information of the processor, the model information includes the data volume of the model weights corresponding to each processing layer, and based on the model information and the processor parameters, a deployment plan is generated, including:

[0089] Based on the memory information and the data volume of the model weights corresponding to each processing layer, determine the first specified information.

[0090] In the embodiments of the present application, the processor parameters may include memory information, and the memory information can reflect the ability of the processor to deploy model weights. Therefore, the first specified information can be determined according to the memory information of the processor and the data volume of the model weights corresponding to each processing layer, that is, the processor for deploying the model weights corresponding to each processing layer is determined.

[0091] In the embodiments of the present application, based on the memory information and the data volume of the model weights corresponding to each processing layer, it is possible to determine whether the storage space of the processor can support the deployment of the model weight files of each processing layer, thereby determining the first specified information.

[0092] As an example, the computing power of different processors can be analyzed, and processors with higher computing power are preferentially used to deploy the model weights of each processing layer. After the storage space of the processors with higher computing power is fully utilized, then consider using processors with lower computing power for model weight deployment.

[0093] In the embodiments of the present application, when deploying the model weights of each processing layer to the processor based on the first specified information in the target device, the storage address of the model weights in the memory buffer can be set, and at the same time, the identifier of the processor can be used as the attribute information of the model weights stored in the memory buffer. So that during the model calculation process, it is possible to determine the processor to which the model weights belong according to the attributes of the data in the memory cache area.

[0094] In an alternative manner of the embodiments of the present application, determining the first specified information based on the memory information and the data volume of the model weights corresponding to each processing layer includes:

[0095] Obtain the historical memory usage of each processor;

[0096] Determine the available memory based on the historical memory usage and the memory information;

[0097] Determine the first specified information based on the available memory and the data volume of the model weights corresponding to each processing layer.

[0098] In the embodiments of the present application, the historical memory usage of the processor can reflect the usage of the memory in the processor. Based on the historical memory usage and the memory information, the available memory in the processor that can be used for model weight deployment can be determined.

[0099] Since the target device may be in different application scenarios, based on the historical memory usage of the processor, it is possible to understand the memory occupancy of the processor in the application scenario under normal circumstances, and avoid allocating too much memory resources to deploy the model service, resulting in competition for memory resources between the model service and other services and a decline in processing performance.

[0100] As an example, when there are multiple processors in the target device, if it is determined based on the memory historical usage of a certain processor that the memory resources in this processor are being used at a high load and cannot stably support the normal deployment and operation of the model, this processor can be disabled and other processors in the target device can be used for model deployment.

[0101] In an alternative implementation of the embodiments of the present application, the processor parameters are collected based on a preset test program running on the target device.

[0102] In the embodiments of the present application, a preset test program can be run on the target device to obtain the processor parameters of each processor. The preset test program can include, but is not limited to, a benchmark program, and the collected processor parameters such as the computing power, transmission bandwidth, and storage space of the processor.

[0103] As an example, the operating system of the target device provides relevant application programming interfaces (APIs), and the processor parameters can be obtained through these interfaces.

[0104] As an example, Figure 3 FIG. shows a schematic diagram of a specific implementation manner of the model deployment scheme generation method provided by the embodiments of the present application.

[0105] As Figure 3 shown in, the model deployment scheme generation method in this example includes the following steps:

[0106] Step S310: Load model information and processor parameters;

[0107] Step S320: The offline planner generates a deployment scheme;

[0108] Step S330: Output the deployment scheme.

[0109] In the embodiments of the present application, before deploying the large model, the offline planner can be run in an offline environment first. The offline planner can read the processor parameters through the APIs provided by the operating system of the target device.

[0110] After the offline planner loads the model information and processor parameters, it generates a deployment scheme according to the model information and processor parameters, and formats and outputs the deployment scheme, that is, outputs a deployment scheme that can be directly used by the large language model inference framework.

[0111] As an example, Figure 4 FIG. shows a schematic diagram of the process during model weight deployment in the embodiments of the present application. As Figure 4 shown in, this process includes the following steps:

[0112] Step S410: Load the model;

[0113] Step S420: Load the deployment plan;

[0114] Step S430: Determine the model weight deployment strategy based on the deployment plan;

[0115] Step S440: According to the model weight deployment strategy, load the model weights of each processing layer into the specified processors respectively, and at the same time initialize the corresponding memory buffer areas in the processors and set the attributes of the memory buffer areas.

[0116] In the embodiment of the present application, by loading the model, the model to be deployed runs on the target device in the form of an executable program. The target device can load the deployment plan and determine the model weight deployment strategy from the deployment plan. The model weight deployment strategy is equivalent to the aforementioned first specified information. Based on the first specified information, the model weights of each processing layer can be loaded into the specified processors respectively. At the same time, the corresponding memory buffer areas can be initialized in the processors for storing the model weights, and the attributes of the memory buffer areas can be set as the identifiers of each processor.

[0117] As an example, Figure 5 shows the schematic flowchart of the model inference calculation in the embodiment of the present application. As shown in Figure 5 it, the process includes the following steps:

[0118] Step S510: Load the model weights of each processing layer;

[0119] Step S520: For the current processing layer, determine the processor X for performing the model calculation of the current processing layer;

[0120] Step S530: Whether the model weights of the current processing layer are deployed on the processor X;

[0121] Step S540: Call the processor X to perform the model calculation of the current processing layer;

[0122] Step S550: Transfer the weights of the current processing layer from the processor Y to the processor X, where the model weights of the current processing layer are deployed in the processor Y;

[0123] Step S560: Call the processor X to perform the model calculation of the current processing layer;

[0124] Step S570: Transmit the model calculation result of the current processing layer to the processor for performing the model calculation of the next processing layer;

[0125] Step S580: Finish the model calculation processing of all processing layers;

[0126] Step S590: End.

[0127] In the embodiments of the present application, processor X and processor Y are different processors in the target device.

[0128] The model weights corresponding to each processing layer of the model to be deployed will be deployed to the specified processor. For the current processing layer for which model calculation is about to be performed, processor X that is used to execute the model calculation of the current processing layer can be determined.

[0129] Determine whether the model weights of the current processing layer are deployed to processor X. If the model weights of the current processing layer are deployed to processor X, then processor X can directly execute the model calculation of the current processing layer. If the model weights of the current processing layer are not deployed to processor X, for example, deployed to processor Y, then the weights of the current processing layer can be transferred from processor Y to processor X, and then processor X is called to execute the model calculation of the current processing layer.

[0130] After completing the model calculation of the current processing layer, the model calculation result of the current processing layer can be transmitted to the processor that executes the model calculation of the next processing layer for subsequent model calculation.

[0131] Determine whether the model calculations of all processing layers are completed. If the model calculations of all processing layers are completed, then the process ends; if the model calculations of all processing layers are not completed, then steps S520 to S570 are sequentially executed according to the order of the processing layers.

[0132] Figure 6 shows a schematic flowchart of a model processing method provided by an embodiment of the present application. This method can be Figure 1 executed by the model processing device in the system shown in Figure 6 As shown, this method mainly may include:

[0133] Step S610: Obtain a deployment plan for the model to be deployed. The deployment plan includes first specified information. The model to be deployed includes multiple processing layers. The target device for deploying the model to be deployed includes multiple processors. The first specified information is used to specify the processors to which the model weights of each processing layer are deployed. The model deployment plan is generated by the server based on the model information of the model to be deployed and the processor parameters of the processors;

[0134] Step S620: Based on the first specified information, deploy the model weights of each processing layer to the specified processors in the target device.

[0135] Among them, the model to be deployed may include but is not limited to large language models. The model to be deployed may include multiple processing layers. The model weights of the model to be deployed are composed of the model weights corresponding to each processing layer. Based on the model weights corresponding to each processing layer, model calculations corresponding to each processing layer can be performed.

[0136] In the embodiments of the present application, the model to be deployed will be deployed in a target device. As a preferred embodiment, the target device in this case is an edge device, and the model service is run through the edge device.

[0137] The target device includes multiple processors. The processors in the target device can be of various types, including but not limited to CPUs, GPUs, NPUs, DSPs, etc. The processor parameters can reflect the hardware performance of the processors.

[0138] In the embodiments of the present application, the model to be deployed can run in the target device in the form of an executable program. The model deployment involved in the embodiments of the present application mainly refers to loading and storing the model weights into the processors of the target device.

[0139] In the embodiments of the present application, based on the model information and the processor parameters, the server can rationally allocate the processors in the target device according to the requirements of each processing layer of the model to be deployed and the performance of each processor in the target device, so that the model to be deployed can be reasonably deployed in the target device.

[0140] Specifically, when deploying the model weights of each processing layer, the storage space situation of the processors is mainly considered. Therefore, based on the first specified information in the deployment plan, the model weights of each processing layer can be deployed to the processors whose storage space meets the requirements.

[0141] In the embodiments of the present application, completing the model deployment based on the deployment plan can effectively allocate the processor resources in the target device, which helps to reasonably deploy the model to be deployed on the target device.

[0142] The method provided by the embodiments of the present application obtains a deployment plan for the model to be deployed. The deployment plan includes first specified information and second specified information. The model to be deployed includes multiple processing layers. The target device for deploying the model to be deployed includes multiple processors. The first specified information is used to specify the processors to which the model weights of each processing layer are deployed. The second specified information is used to specify the processors for performing the model calculations corresponding to each processing layer. The model deployment plan is generated by the server based on the model information of the model to be deployed and the processor parameters of the processors. Based on the first specified information, the model weights of each processing layer are deployed to the specified processors in the target device, and the model calculations of the model to be deployed are performed based on the second specified information. Performing model deployment based on the deployment plan provided by the embodiments of the present application can rationally allocate the processor resources in the target device, thereby realizing the reasonable deployment of the model on the target device.

[0143] In the embodiments of the present application, the deployment plan may further include second specified information, which is used to specify the processors for performing the model calculations corresponding to each processing layer.

[0144] When performing the model calculations corresponding to each processing layer, the computing power of the processor is mainly considered. Therefore, by generating the second specified information, it is possible to allocate a processor with computing power that meets the requirements to perform the model calculations corresponding to different processing layers.

[0145] It can be understood that for any processing layer of the model to be deployed, the processor used to deploy the model weights corresponding to this processing layer and the processor used to execute the model calculations for this processing layer can be the same or different.

[0146] Based on the deployment solution provided in the embodiments of the present application, it is possible to reasonably allocate the storage resources and computing resources of the target device, thereby realizing the reasonable deployment of the model on the target device.

[0147] In the embodiments of the present application, the second specified information can be generated on the server side and sent to the target device, so as to reasonably specify the processor to execute the models of each processing layer. The second specified information can also be converted into an execution logic and automatically run on the target device to realize automatically selecting a processor to execute the models of each processing layer.

[0148] In an alternative embodiment of the present application, after deploying the model weights of each processing layer to the specified processor in the target device, the above method further includes:

[0149] For any processing layer, determine the current processor that executes the model calculation of this processing layer in the understanding stage;

[0150] In response to the current processor being the first processor, read the model weights corresponding to this processing layer from the memory buffer of the first processor, and perform the model calculation of this processing layer in the understanding stage based on the model weights corresponding to this processing layer, where the first processor is the processor used to deploy the model weights corresponding to this processing layer;

[0151] In response to the current processor not being the first processor, transfer the model weights corresponding to this processing layer from the memory buffer of the first processor to the memory buffer of the current processor, read the model weights corresponding to this processing layer from the memory buffer of the current processor, and perform the model calculation of this processing layer in the understanding stage based on the model weights corresponding to this processing layer.

[0152] In the embodiments of the present application, the model to be deployed can be other types of large models that include but are not limited to text-to-image large models, text-to-video large models, etc., in addition to large language models, and the processing process includes understanding and generation stages.

[0153] In the understanding stage of the model to be deployed, multiple or all of the data to be processed are input at one time, and all the input data to be processed are processed at one time, which has certain requirements for the storage space and computing power of the processor. In actual situations, some processors of the target device may have poor computing power, while some other processors may have insufficient storage space. In view of these situations, if the processors corresponding to the model calculations of each processing layer in the understanding stage can be reasonably selected, the model calculation efficiency can be effectively improved.

[0154] In the embodiments of the present application, before performing the model calculation of any processing layer, the current processor for performing the model calculation of this processing layer in the understanding stage can be determined. By reasonably selecting the current processor, its computing power can meet the computing power requirements of this processing layer, thereby improving the model calculation efficiency.

[0155] If the current processor is the first processor that deploys the model weights corresponding to this processing layer, it means that using the first processor to perform the model calculation corresponding to this processing layer can have a better processing speed. At this time, the first processor can be directly called to perform the model calculation corresponding to this processing layer.

[0156] If the current processor is the second processor, it means that using the second processor with higher computing power to perform the model calculation corresponding to this processing layer can have a better processing speed. At this time, the model weights corresponding to this processing layer can be first transferred to the second processor, and then the second processor can be called to perform the model calculation of this processing layer in the understanding stage.

[0157] In an alternative embodiment of the present application, transferring the model weights corresponding to this processing layer from the memory buffer of the first processor to the memory buffer of the current processor includes:

[0158] Before calling the current processor to perform the model calculation of this processing layer, control the current processor to start a transfer task, and the transfer task is used to transfer the model weights of this processing layer to the memory buffer of the current processor.

[0159] In the embodiments of the present application, for any processing layer, when the model weights corresponding to this processing layer are not stored in the memory of the current processor, a transfer task for transferring the model weights corresponding to this processing layer from the memory buffer of other processors to the memory buffer of the current processor needs to be executed. At this time, the transfer task can be started before the current processor performs the model calculation of this processing layer. This can greatly reduce the additional waiting for the transfer of model weights when the processing process of the model to be deployed flows to this processing layer, so as to quickly perform the model calculation of this processing layer.

[0160] In the embodiments of the present application, by controlling the current processor to transmit the model weights in advance and reasonably setting the start time of the transmission, the transmission duration of the model weights can be effectively hidden, thereby greatly improving the processing efficiency of the model.

[0161] As an example, it is possible to control the transmission task to be completed before the current processor performs the model calculation of this processing layer. Through reasonable estimation of the transmission duration and reasonable setting of the start time of the model weights, it is ensured that before the model calculation of this processing layer is executed, the transmission task has been completed and the model weights of this processing layer have been stored in the memory buffer of the current processor. This enables the processing flow of the model to be deployed to quickly perform the model calculation of this processing layer without waiting for the model weights to be transmitted additionally when it reaches this processing layer.

[0162] As an example, the model calculations of the third and fourth processing layers of the model to be deployed are both executed by the current processor. The model weights of the third processing layer are deployed in the current processor, and the weights of the fourth processing layer are deployed in other processors. When performing the model calculation of the third processing layer, the transmission task of the model weights is started asynchronously, so that the transmission duration of the model weights can be effectively hidden in the model calculation of the third processing layer, greatly improving the processing efficiency of the model. In this example, if the transmission duration is long, the transmission task of the model weights can also be started at an earlier time point before the model calculation of the third processing layer is performed, so as to ensure that the model weights can be transmitted to the current processor in time to ensure that the current processor can quickly perform the model calculation of this processing layer.

[0163] In an alternative embodiment of the present application, the above method further includes:

[0164] After the model calculation of this processing layer is executed, the model weights of this processing layer are removed from the memory buffer of the current processor.

[0165] In the embodiments of the present application, after the current processor completes the model calculation of this processing layer, the model weights can be promptly removed from the memory buffer of the current processor, thereby releasing the memory space of the current processor.

[0166] As an example, for a certain processor in the target device, the memory buffer of this processor may include a static storage area and a dynamic storage area. Among them, the static storage area is used for static storage of the model weights deposited in the memory buffer of this processor during initial deployment. The dynamic storage area is used for dynamic swapping in and out of the model weights, that is, for storing the model weights transmitted to the memory buffer of this processor during the model operation and removing the corresponding model weights in the dynamic storage area after the corresponding model calculation is executed.

[0167] In the embodiments of the present application, during the execution of the model processing by the target device, the model weights can be transferred between the processor for deploying the model weights and the processor for executing the model calculation, so that the model calculation process of each processing layer can be constructed in the form of a pipeline composed of multiple processors, automatically performing the process of dynamically incoming and outgoing the model weights, and effectively hiding the time-consuming of the model weight transfer, thereby improving the processing efficiency of the model.

[0168] In the embodiments of the present application, by providing a deployment solution and allocating processors for calculation for each processing layer, the processing efficiency of the model can be effectively improved without modifying the structure of the model itself, and it has a wider applicability.

[0169] In an alternative embodiment of the present application, determining the current processor for performing the model calculation of the processing layer in the understanding stage includes:

[0170] Determining a first duration corresponding to a first processor and a second duration corresponding to a second processor, where the computing power of the second processor is higher than that of the first processor, the first duration is the duration for the first processor to execute the model calculation corresponding to the processing layer, and the second duration is the sum of the transmission duration and the calculation duration. The transmission duration is the duration for transmitting the data required for the model calculation of the processing layer to the second processor, and the calculation duration is the duration for the second processor to execute the model calculation corresponding to the processing layer;

[0171] Based on the first duration and the second duration, determining the current processor for performing the model calculation of the processing layer in the understanding stage.

[0172] In the embodiments of the present application, for any processing layer, the first duration for the first processor to process the model calculation corresponding to the processing layer and the second duration for the second processor to process the model calculation corresponding to the processing layer can be determined respectively.

[0173] The second processor is a processor with a higher computing power than the first processor. When the second processor processes the model calculation corresponding to the processing layer, it is necessary to first transmit the data required for the model calculation of the processing layer to the second processor, and record the duration required for this data transmission as the transmission duration. The second duration corresponding to the second processor is actually the sum of the transmission duration and the calculation duration for the second processor to execute the model calculation corresponding to the processing layer.

[0174] In the embodiments of the present application, based on the first duration and the second duration, a processor that can more efficiently complete the processing process corresponding to the processing layer can be determined.

[0175] In an alternative embodiment of the present application, based on the first duration and the second duration, determining the current processor for performing the model calculation of the processing layer in the understanding stage includes:

[0176] In response to the first duration being not greater than the second duration, determine the first processor as the current processor for performing the model calculation of this processing layer in the understanding stage;

[0177] In response to the first duration being greater than the second duration, determine the second processor as the current processor for performing the model calculation of this processing layer.

[0178] In the embodiments of the present application, when the first duration is not greater than the second duration, it means that when the first processor executes the model calculation of this processing layer, the processing duration is shorter and the processing efficiency is higher. At this time, the first processor can be used as the current processor to execute the model calculation of this processing layer.

[0179] When the first duration is greater than the second duration, it means that when the second processor executes the model calculation of this processing layer, the processing duration is shorter and the processing efficiency is higher. At this time, the second processor can be used as the current processor to execute the model calculation of this processing layer.

[0180] In the embodiments of the present application, using the second processor as the current processor to execute the model calculation of this processing layer can be understood as that the model weights of this processing layer are deployed on the first processor with lower computing power (such as a CPU), and the second processor with higher computing power (such as an NPU) is called to obtain the model weights of this processing layer from the memory of the first processor, and then execute the model calculation of this processing layer.

[0181] In the embodiments of the present application, after the current processor executes the model calculation of this processing layer, a calculation result will be obtained, and this calculation result needs to be stored for subsequent model calculations. Specifically, in the memory buffer of the first processor for deploying the model weights of this processing layer, a part of the space for storing the calculation result can be reserved, and this part of the space can be used to store the calculation result. When the first processor executes the model calculation of this processing layer, the first processor can directly store the calculation result in the memory buffer. When the second processor executes the model calculation of this processing layer, the second processor can first send the calculation result to the first processor, and then the first processor stores the calculation result in the memory buffer.

[0182] In an alternative embodiment of the present application, the data required for the model calculation of this processing layer includes the model weights corresponding to this processing layer. Determining the transmission duration includes:

[0183] Obtain the transmission bandwidth of the second processor and the data volume of the model weights corresponding to this processing layer;

[0184] Based on the transmission bandwidth and the data volume, determine the transmission duration for the second processor to obtain the data required for performing the model calculation of this processing layer.

[0185] In an embodiment of the present application, the data required for the model calculation of the processing layer includes the model weights corresponding to the processing layer and the calculation results of the previous processing layer. Among them, the data volume of the model weights corresponding to the processing layer is relatively large and is the main data to be transmitted. Therefore, the transmission of the model weights corresponding to the processing layer can be used as the transmission duration.

[0186] The model information may include the data volume of the model weights corresponding to each processing layer, and the processor parameters may include the transmission bandwidth of the processor. Based on the transmission bandwidth and the data volume, the transmission duration for the second processor to obtain the data required for the model calculation of the processing layer can be determined.

[0187] In an alternative embodiment of the present application, determining the first duration corresponding to the first processor includes:

[0188] Obtain the layer calculation amount of the model calculation corresponding to the processing layer and the computing power of the first processor;

[0189] Based on the layer calculation amount and the computing power of the first processor, determine the first duration for the first processor to execute the model calculation corresponding to the processing layer.

[0190] In an embodiment of the present application, the layer calculation amounts of each processing layer of the model to be deployed can be obtained and used as model information. As an example, the layer calculation amount can be estimated according to the data volume of the model weights corresponding to each processing layer.

[0191] In an embodiment of the present application, the computing power of the first processor can be obtained. Based on the computing power of the first processor and the layer calculation amount, the first duration required for the first processor to process the model calculation corresponding to the processing layer can be determined.

[0192] Similarly, the computing power of the second processor can be obtained. Based on the computing power of the second processor and the layer calculation amount, the calculation duration required for the second processor to process the model calculation corresponding to the processing layer can be determined. Adding it to the transmission duration can obtain the second duration.

[0193] In an alternative embodiment of the present application, after deploying the model weights of each processing layer to the specified processor in the target device, the method further includes:

[0194] For any processing layer, call the first processor to execute the model operation corresponding to the processing layer in the generation stage, where the first processor is the processor used to deploy the model weights corresponding to the processing layer.

[0195] In the embodiment of the present application, in the generation stage of the model to be deployed, each time a token is input and processed. The demand for the computing power of the processor in the generation stage is relatively low. Therefore, in the generation stage, the first processor that deploys the model weights corresponding to a certain processing layer can be directly used to execute the model calculation of this processing layer. That is to say, in the generation stage, it is not necessary to transfer the model weights to a processor with higher computing power to execute the model calculation of this processing layer, and relatively high model calculation efficiency can still be achieved.

[0196] In the embodiment of the present application, since each time a token is input and processed in the generation stage, if the method of specifying the processor for executing the model operation of each processing layer in the understanding stage is still followed, it will cause the model weights to be transferred relatively many times, resulting in a large amount of time consumed for the transfer of the model weights and affecting the model processing efficiency. Considering that the demand for the computing power of the processor in the generation stage is relatively low, the processor used for model deployment generally can meet the demand for the computing power of the processor in the generation stage. Therefore, the first processor that deploys the model weights corresponding to a certain processing layer can be obtained, and the first processor can be called to execute the model calculation of this processing layer.

[0197] In an alternative embodiment of the present application, before obtaining the deployment plan for the model to be deployed, the above method further includes:

[0198] Running a preset test program to collect the processor parameters of each processor of the target device;

[0199] Sending the processor parameters to the server so that the server generates a deployment plan based on the processor parameters and the model information of the model to be deployed.

[0200] In the embodiment of the present application, the preset test program can be run on the target device to obtain the processor parameters of each processor. The preset test program can include but is not limited to the benchmark program, and the collected processor parameters such as the computing power, transmission bandwidth, and storage space of the processor.

[0201] As an example, relevant APIs will be provided in the operating system of the target device, and the processor parameters can be sent to the server through these interfaces.

[0202] Based on the same principle as the method shown in Figure 2 Figure 7 shows a schematic structural diagram of a model deployment plan generation device provided by an embodiment of the present application. As shown in Figure 7

[0203] A model information acquisition module 710, configured to acquire the model information of the model to be deployed;

[0204] A processor parameter acquisition module 720, configured to acquire processor parameters of each processor in a target device, where the target device is used to deploy a model to be deployed.

[0205] A deployment plan generation module 730, configured to generate a deployment plan based on model information and processor parameters, where the deployment plan includes first specified information, the model to be deployed includes multiple processing layers, and the first specified information is used to specify the processor on which the model weights of each processing layer are deployed.

[0206] The device provided in the embodiments of the present application acquires the model information of the model to be deployed, acquires the processor parameters of each processor in the target device used to deploy the model to be deployed, and thus generates a deployment plan based on the model information and the processor parameters. The deployment plan includes first specified information for specifying the processor on which the model weights of each processing layer are deployed. Based on the deployment plan provided in the embodiments of the present application, reasonable allocation of processor resources in the target device can be achieved, and thus reasonable deployment of the model in the target device can be achieved.

[0207] Optionally, the deployment plan further includes second specified information, and the second specified information is used to specify the processor that executes the model calculation corresponding to each processing layer.

[0208] Optionally, the second specified information includes second specified information in the understanding stage, and the second specified information in the understanding stage is used to specify the processor that executes the model calculation corresponding to each processing layer in the understanding stage. The deployment plan generation module is specifically configured to:

[0209] For any processing layer, determine a first duration corresponding to a first processor and a second duration corresponding to a second processor, where the model weights corresponding to this processing layer are deployed on the first processor, the computing power of the second processor is higher than that of the first processor, the first duration is the duration for the first processor to execute the model calculation corresponding to this processing layer, and the second duration is the sum of the transmission duration and the calculation duration. The transmission duration is the duration for transmitting the data required for executing the model calculation of this processing layer to the second processor, and the calculation duration is the duration for the second processor to execute the model calculation corresponding to this processing layer;

[0210] Based on the first duration and the second duration, determine the second specified information in the understanding stage.

[0211] Optionally, the second specified information includes second specified information in the generation stage, and the second specified information in the generation stage is used to specify that the processor that executes the model calculation corresponding to each processing layer in the generation stage is the first processor, and the first processor is the processor on which the model weights corresponding to each processing layer are deployed.

[0212] Optionally, the processor parameters include the memory information of the processor, the model information includes the data volume of the model weights corresponding to each processing layer, and the deployment plan generation module is specifically configured to:

[0213] Determine the first specified information based on the memory information and the data volume of the model weights corresponding to each processing layer.

[0214] Optionally, when determining the first specified information based on the memory information and the data volume of the model weights corresponding to each processing layer, the deployment scheme generation module is specifically configured to:

[0215] Obtain the historical memory usage of each processor;

[0216] Determine the available memory based on the historical memory usage and the memory information;

[0217] Determine the first specified information based on the available memory and the data volume of the model weights corresponding to each processing layer.

[0218] Optionally, the processor parameters are collected based on a preset test program running on the target device.

[0219] Based on the same principle as the method shown in Figure 6 Fig. Figure 8 shows a schematic structural diagram of a model processing device provided in an embodiment of the present application, as Figure 8 shown, the model processing device 80 may include:

[0220] A deployment scheme acquisition module 810, configured to acquire a deployment scheme for a model to be deployed, where the deployment scheme includes first specified information, the model to be deployed includes multiple processing layers, the target device for deploying the model to be deployed includes multiple processors, the first specified information is used to specify the processors on which the model weights of each processing layer are deployed, and the model deployment scheme is generated by the server based on the model information of the model to be deployed and the processor parameters of the processors;

[0221] A model deployment module 820, configured to deploy the model weights of each processing layer to the specified processors in the target device based on the first specified information.

[0222] The device provided in the embodiment of the present application acquires a deployment scheme for a model to be deployed, where the deployment scheme includes first specified information and second specified information, the model to be deployed includes multiple processing layers, the target device for deploying the model to be deployed includes multiple processors, the first specified information is used to specify the processors on which the model weights of each processing layer are deployed, the second specified information is used to specify the processors that execute the model calculations corresponding to each processing layer, and the model deployment scheme is generated by the server based on the model information of the model to be deployed and the processor parameters of the processors. Based on the first specified information, the model weights of each processing layer are deployed to the specified processors in the target device, and model calculations of the model to be deployed are performed based on the second specified information. Model deployment based on the deployment scheme provided in the embodiment of the present application can achieve reasonable allocation of processor resources in the target device, thereby realizing reasonable deployment of the model in the target device.

[0223] Optionally, the above device further includes a model calculation module, and the model calculation module is configured to:

[0224] After deploying the model weights of each processing layer to the specified processor in the target device, for any processing layer, determine the current processor that executes the model calculation of this processing layer in the understanding stage;

[0225] In response to the current processor being the first processor, read the model weights corresponding to this processing layer from the memory buffer of the first processor, and execute the model calculation of this processing layer in the understanding stage based on the model weights corresponding to this processing layer, where the first processor is the processor used to deploy the model weights corresponding to this processing layer;

[0226] In response to the current processor not being the first processor, transfer the model weights corresponding to this processing layer from the memory buffer of the first processor to the memory buffer of the current processor, read the model weights corresponding to this processing layer from the memory buffer of the current processor, and execute the model calculation of this processing layer in the understanding stage based on the model weights corresponding to this processing layer.

[0227] Optionally, when the model calculation module transfers the model weights corresponding to this processing layer from the memory buffer of the first processor to the memory buffer of the current processor, it is specifically configured to:

[0228] Before calling the current processor to execute the model calculation of this processing layer, control the current processor to start a transfer task, and the transfer task is used to transfer the model weights of this processing layer to the memory buffer of the current processor.

[0229] Optionally, the above device further includes:

[0230] A model weight removal module, configured to remove the model weights of this processing layer from the memory buffer of the current processor after the model calculation of this processing layer is completed.

[0231] Optionally, when the model calculation module determines the current processor that executes the model calculation of this processing layer in the understanding stage, it is specifically configured to:

[0232] Determine a first duration corresponding to the first processor and a second duration corresponding to the second processor, where the computing power of the second processor is higher than that of the first processor, the first duration is the duration for the first processor to execute the model calculation corresponding to this processing layer, the second duration is the sum of the transfer duration and the calculation duration, the transfer duration is the duration for transferring the data required for executing the model calculation of this processing layer to the second processor, and the calculation duration is the duration for the second processor to execute the model calculation corresponding to this processing layer;

[0233] Based on the first duration and the second duration, determine the current processor that executes the model calculation of this processing layer in the understanding stage.

[0234] Optionally, when determining the current processor for performing the model calculation of this processing layer in the understanding stage based on the first duration and the second duration, the model calculation module is specifically configured to:

[0235] In response to the first duration being not greater than the second duration, determine the first processor as the current processor for performing the model calculation of this processing layer in the understanding stage;

[0236] In response to the first duration being greater than the second duration, determine the second processor as the current processor for performing the model calculation of this processing layer in the understanding stage.

[0237] Optionally, the data required for the model calculation of this processing layer includes the model weights corresponding to this processing layer. When determining the transmission duration, the model calculation module is specifically configured to:

[0238] Obtain the transmission bandwidth of the second processor and the data volume of the model weights corresponding to this processing layer;

[0239] Based on the transmission bandwidth and the data volume, determine the transmission duration for the second processor to obtain the data required for performing the model calculation of this processing layer.

[0240] Optionally, when determining the first duration corresponding to the first processor, the model calculation module is specifically configured to:

[0241] Obtain the layer calculation amount of the corresponding model calculation of this processing layer and the computing power of the first processor;

[0242] Based on the layer calculation amount and the computing power of the first processor, determine the first duration for the first processor to perform the model calculation corresponding to this processing layer.

[0243] Optionally, the model calculation module is further configured to:

[0244] After deploying the model weights of each processing layer to the specified processor in the target device, for any processing layer, call the first processor to perform the model operation corresponding to this processing layer in the generation stage, where the first processor is the processor used to deploy the model weights corresponding to this processing layer.

[0245] Optionally, the above device further includes:

[0246] A processor parameter acquisition module, configured to run a preset test program and acquire the processor parameters of each processor of the target device before obtaining the deployment plan for the model to be deployed;

[0247] A processor parameter reporting module, configured to send the processor parameters to the server, so that the server generates a deployment plan based on the processor parameters and the model information of the model to be deployed.

[0248] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the apparatus embodiments, since they are basically similar to the method embodiments, they are described relatively simply. For the relevant parts, reference can be made to the descriptions in the method embodiments. The apparatus embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0249] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. And the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions, and corresponding operation entrances are provided for users to select to authorize or refuse.

[0250] The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the method in any one of the foregoing method embodiments are implemented.

[0251] Compared with the prior art, the computer-readable storage medium provided by the embodiment of the present application obtains the model information of the model to be deployed, and obtains the processor parameters of each processor in the target device for deploying the model to be deployed. Then, based on the model information and the processor parameters, a deployment plan is generated. The deployment plan includes first specified information for specifying the processor on which the model weights of each processing layer are deployed and second specified information for specifying the processor for executing the corresponding model calculation of each processing layer. Based on the deployment plan provided by the embodiment of the present application, reasonable allocation of processor resources in the target device can be realized, so as to realize the reasonable deployment of the model in the target device.

[0252] The embodiment of the present application also provides an electronic device, including:

[0253] One or more processors; and

[0254] A memory associated with the above one or more processors, which is used to store program instructions. When the program instructions are read and executed by the above one or more processors, the steps of the method in any one of the foregoing method embodiments are executed.

[0255] Compared with the prior art, the electronic device provided by the embodiment of the present application obtains the model information of the model to be deployed, and obtains the processor parameters of each processor in the target device for deploying the model to be deployed. Then, based on the model information and the processor parameters, a deployment scheme is generated. The deployment scheme includes first specification information for specifying the processors where the model weights of each processing layer are deployed and second specification information for specifying the processors that execute the model calculations corresponding to each processing layer. Based on the deployment scheme provided by the embodiment of the present application, reasonable allocation of the processor resources in the target device can be achieved, so as to realize the reasonable deployment of the model in the target device.

[0256] As an example, Figure 9 FIG. shows a schematic structural diagram of an electronic device applicable to the embodiment of the present application, as Figure 9 shown, the electronic device 2000 includes: a processor 2001 and a memory 2003. Among them, the processor 2001 and the memory 2003 are connected, such as connected through a bus 2002. Optionally, the electronic device 2000 may further include a transceiver 2004. It should be noted that in practical applications, the transceiver 2004 is not limited to one, and the structure of the electronic device 2000 does not constitute a limitation to the embodiment of the present application.

[0257] Among them, the processor 2001 is applied to the embodiment of the present application to implement the method shown in the above method embodiment. The transceiver 2004 may include a receiver and a transmitter, and the transceiver 2004 is applied to the embodiment of the present application to implement the function of communicating between the electronic device of the embodiment of the present application and other devices when executed.

[0258] The processor 2001 may be a CPU (Central Processing Unit, central processor), a general-purpose processor, a DSP (Digital Signal Processor, data signal processor), an ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logic blocks, modules and circuits described in combination with the disclosure of the present application. The processor 2001 may also be a combination for implementing computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0259] The bus 2002 may include a path for transmitting information among the above components. The bus 2002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 2002 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 9 it is only represented by a thick line in the figure, but it does not mean that there is only one bus or one type of bus.

[0260] The memory 2003 may be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or it may also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0261] Optionally, the memory 2003 is used to store the application program code for implementing the solution of this application, and is controlled by the processor 2001 to execute. The processor 2001 is used to execute the application program code stored in the memory 2003 to implement the steps of the method in any one of the foregoing method embodiments.

[0262] This application also provides a computer program product, including a computer program, which implements the steps of the method in any one of the foregoing method embodiments when executed by a processor.

[0263] Compared with the prior art, the computer program product provided by the embodiments of the present application generates a deployment plan by obtaining the model information of the model to be deployed and the processor parameters of each processor in the target device for deploying the model to be deployed. The deployment plan includes first specification information for specifying the processors on which the model weights of each processing layer are deployed and second specification information for specifying the processors for executing the model calculations corresponding to each processing layer. Based on the deployment plan provided by the embodiments of the present application, reasonable allocation of the processor resources in the target device can be achieved, so as to realize the reasonable deployment of the model in the target device.

[0264] It should be understood that although each step in the flowchart of the accompanying drawings is shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps is not strictly restricted by order, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. Their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0265] The above are only some embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for generating a model deployment solution, characterized in that Including: Obtaining model information of a model to be deployed; Obtaining processor parameters of each processor in a target device, where the target device is used to deploy the model to be deployed; Generating a deployment plan based on the model information and the processor parameters, where the deployment plan includes first specified information, the model to be deployed includes multiple processing layers, and the first specified information is used to specify the processor on which the model weights of each of the processing layers are deployed; The deployment plan further includes second specified information, the second specified information is used to specify the processor that executes the model calculation corresponding to each of the processing layers, the second specified information includes second specified information in an understanding stage, and the second specified information in the understanding stage is used to specify the processor that executes the model calculation corresponding to each of the processing layers in the understanding stage; The second specified information includes second specified information in a generation stage, and the second specified information in the generation stage is used to specify that the processor that executes the model calculation corresponding to each of the processing layers in the generation stage is a first processor, and the first processor is the processor on which the model weights of each of the processing layers are deployed.

2. The method according to claim 1, wherein The generating a deployment plan based on the model information and the processor parameters includes: For any one of the processing layers, determining a first duration corresponding to a first processor and a second duration corresponding to a second processor, where the model weights corresponding to this processing layer are deployed on the first processor, the computing power of the second processor is higher than that of the first processor, the first duration is the duration for the first processor to execute the model calculation corresponding to this processing layer, and the second duration is the sum of a transmission duration and a calculation duration, the transmission duration is the duration for transmitting the data required for executing the model calculation of this processing layer to the second processor, and the calculation duration is the duration for the second processor to execute the model calculation corresponding to this processing layer; Determining the second specified information in the understanding stage based on the first duration and the second duration.

3. The method according to claim 1, wherein The processor parameters include the memory information of the processor, the model information includes the data volume of the model weights corresponding to each processing layer, and the generating a deployment plan based on the model information and the processor parameters includes: Determining the first specified information based on the memory information and the data volume of the model weights corresponding to each processing layer.

4. The method according to claim 3, wherein The determining the first specified information based on the memory information and the data volume of the model weights corresponding to each processing layer includes: Obtaining the historical memory usage of each of the processors; Determining the available memory based on the historical memory usage and the memory information; Determining the first specified information based on the available memory and the data volume of the model weights corresponding to each processing layer.

5. The method according to any one of claims 1 to 4, characterized in that, The processor parameters are collected based on a preset test program running on the target device.

6. A model processing method, characterized in that, Including: Obtain a deployment plan for the model to be deployed. The deployment plan includes first specified information. The model to be deployed includes multiple processing layers, and the target device for deploying the model to be deployed includes multiple processors. The first specified information is used to specify the processors to which the model weights of each processing layer are deployed. The model deployment plan is generated by the server based on the model information of the model to be deployed and the processor parameters of the processors. Based on the first specified information, deploy the model weights of each processing layer to the specified processors in the target device; The method further includes: For any one of the processing layers, call the current processor to execute the model calculation of this processing layer in the understanding stage, and call the first processor to execute the model operation corresponding to this processing layer in the generation stage; The first processor is the processor used to deploy the model weights corresponding to this processing layer, and the current processor is the first processor or not the first processor.

7. The method according to claim 6, wherein After deploying the model weights of each processing layer to the specified processors in the target device, the method further includes: For any one of the processing layers, determine the current processor that executes the model calculation of this processing layer in the understanding stage; In response to the current processor being the first processor, read the model weights corresponding to this processing layer from the memory buffer of the first processor, and execute the model calculation of this processing layer in the understanding stage based on the model weights corresponding to this processing layer. The first processor is the processor used to deploy the model weights corresponding to this processing layer; In response to the current processor not being the first processor, transfer the model weights corresponding to this processing layer from the memory buffer of the first processor to the memory buffer of the current processor, read the model weights corresponding to this processing layer from the memory buffer of the current processor, and execute the model calculation of this processing layer in the understanding stage based on the model weights corresponding to this processing layer.

8. The method according to claim 7, characterized in that, The transferring the model weights corresponding to this processing layer from the memory buffer of the first processor to the memory buffer of the current processor includes: Before calling the current processor to execute the model calculation of this processing layer, control the current processor to start a transfer task, and the transfer task is used to transfer the model weights of this processing layer to the memory buffer of the current processor.

9. The method according to claim 8, characterized in that, The method further includes: After the model calculation of this processing layer is completed, move the model weights of this processing layer out of the memory buffer of the current processor.

10. The method according to any one of claims 7-9, characterized in that, The determining the current processor that executes the model calculation of this processing layer in the understanding stage includes: Determine the first duration corresponding to the first processor and the second duration corresponding to the second processor. The computing power of the second processor is higher than that of the first processor. The first duration is the duration for the first processor to execute the model calculation corresponding to this processing layer, and the second duration is the sum of the transfer duration and the calculation duration. The transfer duration is the duration for transferring the data required for executing the model calculation of this processing layer to the second processor, and the calculation duration is the duration for the second processor to execute the model calculation corresponding to this processing layer; Based on the first duration and the second duration, determine the current processor for performing the model calculation of this processing layer in the understanding stage.

11. The method according to claim 10, wherein The determining the current processor for performing the model calculation of this processing layer in the understanding stage based on the first duration and the second duration includes: In response to the first duration being no greater than the second duration, determine the first processor as the current processor for performing the model calculation of this processing layer in the understanding stage; In response to the first duration being greater than the second duration, determine the second processor as the current processor for performing the model calculation of this processing layer in the understanding stage.

12. The method according to claim 10, wherein The data required for the model calculation of this processing layer includes the model weights corresponding to this processing layer. Determining the transmission duration includes: Obtain the transmission bandwidth of the second processor and the data volume of the model weights corresponding to this processing layer; Based on the transmission bandwidth and the data volume, determine the transmission duration for the second processor to obtain the data required for performing the model calculation of this processing layer.

13. The method according to claim 10, wherein The determining the first duration corresponding to the first processor includes: Obtain the layer calculation amount of the corresponding model calculation of this processing layer and the computing power of the first processor; Based on the layer calculation amount and the computing power of the first processor, determine the first duration for the first processor to perform the model calculation corresponding to this processing layer.

14. The method according to any one of claims 7-9, characterized in that, Before obtaining the deployment plan for the to-be-deployed model, the method further includes: Run a preset test program to collect the processor parameters of each processor of the target device; Send the processor parameters to the server so that the server generates the deployment plan based on the processor parameters and the model information of the to-be-deployed model.

15. A model deployment scheme generation device, characterized in that, Includes: A model information acquisition module, configured to acquire the model information of the to-be-deployed model; A processor parameter acquisition module, configured to acquire the processor parameters of each processor in the target device, where the target device is used to deploy the to-be-deployed model; A deployment plan generation module, configured to generate a deployment plan based on the model information and the processor parameters. The deployment plan includes first specified information. The to-be-deployed model includes multiple processing layers. The first specified information is used to specify the processor on which the model weights of each processing layer are deployed; the deployment plan further includes second specified information. The second specified information includes second specified information in the understanding stage, and the second specified information in the understanding stage is used to specify the processor for performing the model calculation corresponding to each processing layer in the understanding stage; The second specified information includes second specified information in the generation stage. The second specified information in the generation stage is used to specify that the processor for performing the model calculation corresponding to each processing layer in the generation stage is the first processor, and the first processor is the processor on which the model weights corresponding to each processing layer are deployed.

16. A model processing device, characterized in that, Includes: A deployment plan acquisition module, configured to acquire a deployment plan for a model to be deployed, where the deployment plan includes first specified information, the model to be deployed includes multiple processing layers, and the target device for deploying the model to be deployed includes multiple processors. The first specified information is used to specify the processors to which the model weights of each processing layer are deployed. The model deployment plan is generated by the server based on the model information of the model to be deployed and the processor parameters of the processors. A model deployment module, configured to deploy the model weights of each processing layer to the specified processors in the target device based on the first specified information. The device further includes a model calculation module, and the model calculation module is configured to: For any one of the processing layers, call the current processor to perform the model calculation of this processing layer in the understanding stage, and call the first processor to perform the model operation corresponding to this processing layer in the generation stage. The first processor is the processor for deploying the model weights corresponding to this processing layer, and the current processor is the first processor or not the first processor.

17. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps of the method according to any one of claims 1-14.

18. An electronic device, characterized in that, including: One or more processors; and A memory associated with the one or more processors, where the memory is used to store program instructions. When the program instructions are read and executed by the one or more processors, they execute the steps of the method according to any one of claims 1-14.

19. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-14.

Citation Information

Patent Citations

  • Service deployment method of hybrid cloud platform, management platform, equipment and medium

    CN116684421A

  • Model processing method and device, model operation method and device, equipment and medium

    CN118259975A