Large model deployment method and device, equipment, medium and product

By quantifying the large-scale model to be quantified in the cloud server and sending the target parameters to the local terminal, the problem of low quantization efficiency and accuracy of the local terminal execution model is solved, the inference performance and efficiency of the large-scale model is improved, and the credibility of the quantitative results of the model is ensured.

CN120218137APending Publication Date: 2025-06-27BEIJING FACE WALL INTELLIGENT TECHNOLOGY CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510281465.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

When deploying a large model on a local terminal, the local terminal needs to execute the model quantization process, resulting in low efficiency and accuracy of the model quantization, which further affects the inference performance and efficiency of the large model.

Method used

By generating a large model to be quantized with the same model structure as the large model to be deployed in the local terminal in the cloud server, and the large model to be quantified by the large model to be quantified by the large model to be quantified by the large model to be quantified by the large model to be quantified by the large model to be quantified by the large model to be deployed in the cloud server. Finally, the resulting target model quantization parameters are sent to the local terminal for model deployment.

Benefits of technology

Because cloud servers have powerful data processing performance, using cloud servers for model quantization can improve the efficiency and accuracy of model quantization, and further improve the inference performance and efficiency of large models to be deployed. At the same time, by aligning the model structure between the cloud side and the end side, the credibility of the model quantization results is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218137A_ABST
    Figure CN120218137A_ABST
Patent Text Reader

Abstract

The invention discloses a large model deployment method and device, equipment, a medium and a product, and relates to the technical field of artificial intelligence, and the method comprises the steps: generating a to-be-quantized large model in response to a structure configuration operation of a user on an initial large model in a cloud server; wherein the to-be-quantized large model and a to-be-deployed large model in the local terminal have the same model structure; performing model quantization on the to-be-quantized large model by adopting at least one large model quantization method, generating at least one group of candidate model quantization parameters, and determining a target model quantization parameter from the candidate model quantization parameters; and sending the target model quantization parameter to a local terminal, so that the local terminal performs model deployment on the to-be-deployed large model according to the target model quantization parameter. According to the method, the cloud server is used for model quantification, so that the model quantification process is not limited by local terminal hardware, the efficiency and precision of model quantification are improved, and the reasoning performance and efficiency of the large model to be deployed are further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, device, equipment, medium and product for deploying a large model. Background Art

[0002] In the era of rapid development of information technology today, significant progress has been made in the field of Natural Language Processing (NLP). Large models play a crucial role in it, and their development and application have profoundly changed the way people interact with computers and process language information.

[0003] When deploying a large model on a local terminal, model quantization of the large model is required. Currently, the model quantization process is usually executed by the local terminal. However, due to the limitations of the local terminal hardware, the efficiency and accuracy of the local terminal in executing model quantization are not high, further affecting the inference performance and efficiency of the large model. Summary of the Invention

[0004] The present invention provides a method, device, equipment, medium and product for deploying a large model to solve the problem that when deploying a large model on a local terminal, the local terminal needs to execute the model quantization process, resulting in low efficiency and accuracy of model quantization, and further affecting the inference performance and efficiency of the large model.

[0005] According to one aspect of the present invention, there is provided a method for deploying a large model, which is executed by a cloud server. The method includes:

[0006] Responding to a user's structure configuration operation on the initial large model in the cloud server, generating a large model to be quantized; wherein, the large model to be quantized has the same model structure as the large model to be deployed in the local terminal;

[0007] Using at least one large model quantization method to respectively perform model quantization on the large model to be quantized, generating at least one set of candidate model quantization parameters, and determining target model quantization parameters from the candidate model quantization parameters;

[0008] Sending the target model quantization parameters to the local terminal, so that the local terminal performs model deployment on the large model to be deployed according to the target model quantization parameters.

[0009] According to another aspect of the present invention, there is provided a method for deploying a large model, which is executed by a local terminal. The method includes:

[0010] Obtain the target model quantization parameters sent by the cloud server; wherein, the target model quantization parameters are obtained by the cloud server quantizing the large model to be quantized in the cloud server, and the large model to be quantized has the same model structure as the large model to be deployed in the local terminal;

[0011] Perform model deployment on the large model to be deployed according to the target model quantization parameters.

[0012] According to another aspect of the present invention, there is provided a deployment device for a large model, configured in a cloud server, and the device includes:

[0013] A large model to be quantized generation module, configured to generate a large model to be quantized in response to a user's structure configuration operation on the initial large model in the cloud server; wherein, the large model to be quantized has the same model structure as the large model to be deployed in the local terminal;

[0014] A model quantization module, configured to respectively perform model quantization on the large model to be quantized by using at least one large model quantization method, generate at least one set of candidate model quantization parameters, and determine the target model quantization parameters from the candidate model quantization parameters;

[0015] A model quantization parameter sending module, configured to send the target model quantization parameters to the local terminal, so that the local terminal performs model deployment on the large model to be deployed according to the target model quantization parameters.

[0016] According to another aspect of the present invention, there is provided a deployment device for a large model, configured in a local terminal, and the device includes:

[0017] A model quantization parameter acquisition module, configured to acquire the target model quantization parameters sent by the cloud server; wherein, the target model quantization parameters are obtained by the cloud server quantizing the large model to be quantized in the cloud server, and the large model to be quantized has the same model structure as the large model to be deployed in the local terminal;

[0018] A model deployment module, configured to perform model deployment on the large model to be deployed according to the target model quantization parameters.

[0019] According to another aspect of the present invention, there is provided an electronic device, and the electronic device includes:

[0020] At least one processor; and

[0021] A memory communicatively connected to the at least one processor; wherein,

[0022] The memory stores a computer program executable by the at least one processor. When executed by the at least one processor, the computer program enables the at least one processor to execute the deployment method of the large model according to any one of the present invention.

[0023] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for implementing the deployment method of the large model according to any one of the present invention when executed by a processor.

[0024] According to another aspect of the present invention, there is provided a computer program product including a computer program which implements the deployment method of the large model according to any one of the present invention when executed by a processor.

[0025] In the present invention, a large model to be quantized having the same model structure as the large model to be deployed in the local terminal is generated in the cloud server, and the cloud server quantizes the large model to be quantized. Finally, the obtained target model quantization parameters are sent to the local terminal for model deployment. The beneficial effects are as follows:

[0026] First, since the cloud server has powerful data processing performance, the cloud server is used for model quantization, and then the local terminal performs model deployment according to the obtained target model quantization parameters. As a result, the model quantization process is not restricted by the local terminal hardware, improving the efficiency and accuracy of model quantization, and further enhancing the inference performance and efficiency of the large model to be deployed.

[0027] Second, since the large model to be quantized in the cloud server has the same model structure as the large model to be deployed in the local terminal, the effect of aligning the model structure between the cloud side and the terminal side is achieved, ensuring the credibility of the model quantization result in the cloud server.

[0028] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0030] Figure 1 It is a flowchart of a deployment method of a large model provided for Embodiment 1 of the present invention;

[0031] Figure 2 Flow chart of a method for determining quantization parameters of a target model provided in the second embodiment of the present invention;

[0032] Figure 3 Flow chart of a method for deploying a large model provided in the third embodiment of the present invention;

[0033] Figure 4 Flow chart of a method for deploying a large model provided in the fourth embodiment of the present invention;

[0034] Figure 5 Schematic structural diagram of a device for deploying a large model provided in the fifth embodiment of the present invention;

[0035] Figure 6 Schematic structural diagram of a device for deploying a large model provided in the sixth embodiment of the present invention;

[0036] Figure 7 Schematic structural diagram of an electronic device for implementing the method for deploying a large model according to the embodiments of the present invention. Detailed implementation manners

[0037] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0038] It should be noted that the terms "candidate", "target", "preferred", "current", "to be recognized", "reference", "first", "second", "third", etc. in the specification and claims of the present invention and the above accompanying drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0039] Embodiment 1

[0040] Figure 1This is a flow chart of a large model deployment method provided in the first embodiment of the present invention. This embodiment is applicable to the case where a local terminal uses the target model quantization parameters obtained by the cloud server to deploy the large model to be deployed in the local terminal. The method can be executed by a large model deployment device, which is configured in the cloud server and can be implemented in the form of hardware and / or software. Figure 1 As shown, the method includes:

[0041] S101. In response to a user's structural configuration operation on an initial large model in a cloud server, a large model to be quantized is generated.

[0042] Among them, cloud servers are also called cloud servers. They are based on cloud computing technology. They integrate resources through cluster applications, distributed computing, grid computing and other technologies, unify hardware, software, network and other resources in a wide area network or local area network, and realize data computing, storage, processing and sharing. It is a virtual server that uses the powerful capabilities of the cloud computing platform to provide users with efficient, flexible and scalable computing services.

[0043] The cloud server pre-deploys an initial large model, which refers to a large model with an initial default structure. After the structural configuration operation of the initial large model, the initial large model is converted into a large model to be quantified, that is, a large model to be quantified is generated. The large model to be quantified has the same model structure as the large model to be deployed in the local terminal. The large model to be deployed refers to a large model that needs to be deployed in the local terminal.

[0044] A large model refers to a deep learning model with a large parameter scale, which usually adopts a Transformer-based architecture, can efficiently process and understand complex language descriptions, and convert them into text information. Usually, the parameter scale of a large model can even reach hundreds of billions. The large model in this embodiment can be a language generation model, a language understanding model, or a multimodal model, etc. This embodiment does not limit the type of the large model.

[0045] The users in this embodiment include cloud computing service providers and terminal enterprise developers, etc., and their structure configuration behaviors are constrained by the RBAC (Role-Based Access Control) permission model.

[0046] Before performing structural configuration operations on the initial large model, the user needs to analyze and understand the operator effects and calculation graphs corresponding to the large model to be deployed, so as to perform structural configuration operations on the initial large model accordingly in the cloud server and ensure that the generated large model to be quantized has the same model structure as the large model to be deployed.

[0047] Among them, regarding the operator effects corresponding to the large model to be deployed: Compared with traditional and mature operator libraries, there are significant differences in the implementation of operator libraries by each chip manufacturer, specifically in aspects such as operator support, optimization strategies, and hardware acceleration. Although many commonly used operators have similar implementations on different platforms, due to the differences in the hardware architectures, underlying optimizations, and performance tuning strategies of each manufacturer, the implementation details of these operators usually vary. To clarify these differences, users need to perform independent compilations for each operator and conduct simulation tests on a cloud environment with higher flexibility (such as Nvidia GPUs, etc.), so as to clarify the specific details of the operator implementation (for example, whether activation functions such as SiLU are pure integer operations, etc.). By confirming the general details of the operator implementation, first, it can help users identify the links where accuracy loss may occur, and second, it can also perform more accurate computational simulations in the cloud environment to ensure that the model can maintain the best performance during local terminal inference.

[0048] Regarding the computation graph corresponding to the large model to be deployed: On the one hand, different manufacturers have different requirements for the channel arrangement of the input. If the arrangement of the input data does not meet the platform requirements, it may lead to incorrect inference results. To solve this problem, users need to separately extract and analyze the operators, adjust the arrangement method of the input data, and find the arrangement form that best meets the hardware requirements. On the other hand, each chip manufacturer usually inserts some specific nodes (such as requant, dequant, etc.) into the computation graph, and these nodes may vary between different platforms. To ensure accurate alignment of the computation graph on the cloud with that on the local terminal, users must comprehensively understand the structure of the entire computation graph and optimize and adjust it according to the implementations of different platforms. This precision alignment of the computation graph can be pre-simulated and adjusted on the cloud to ensure the correctness and efficiency of local terminal inference.

[0049] After the user has analyzed and understood the operator effects and computation graph corresponding to the large model to be deployed, a structure configuration operation is performed on the initial large model in the cloud server. In response to the user's structure configuration operation on the initial large model, a large model to be quantized with the same model structure as the large model to be deployed in the local terminal is correspondingly generated.

[0050] S102: Use at least one large model quantization method to perform model quantization on the large model to be quantized respectively, generate at least one set of candidate model quantization parameters, and determine the target model quantization parameters from the candidate model quantization parameters.

[0051] Among them, due to the limitations of the local terminal hardware, the local terminal cannot efficiently and accurately perform model quantization on the large model to be deployed. Therefore, in this embodiment, the cloud server is used to perform model quantization on the large model to be quantized that has the same model structure as the large model to be deployed, so as to simulate the model quantization process of the large model to be deployed, making the model quantization process not restricted by the local terminal hardware.

[0052] The large model quantization method can be pre-set in the cloud server according to requirements, so that the cloud server can directly call various large model quantization methods to perform model quantization on the large model to be quantized. The large model quantization methods include, but are not limited to, generative pre-trained model quantization, smooth quantization, activation-aware weight quantization, and outlier-aware weight quantization, etc. This embodiment does not limit the specific types of large model quantization methods. It can be understood that due to the powerful data processing ability of the cloud server, some more flexible and advanced large model quantization methods can be used to improve the effect of model quantization.

[0053] Model quantization refers to a technology that optimizes model storage, computing efficiency, and deployment capabilities by reducing the numerical precision of model parameters and converting high-precision floating-point numbers into low-precision formats. Candidate model quantization parameters refer to the model parameters in low-precision format obtained after performing model quantization on the large model to be quantized using the large model quantization method.

[0054] In one implementation, at least one large model quantization method is pre-packaged as a Docker image and deployed to the cloud image repository of the cloud server. The cloud server pulls the Docker image from the cloud image repository through the container engine and instantiates it as a container on the cloud server, and executes at least one large model quantization method defined in the Docker image to perform model quantization on the large model to be quantized respectively, generating at least one set of candidate model quantization parameters.

[0055] In another implementation, at least one large model quantization method is pre-written into a script file. When the cloud server detects through the event listener that the large model to be quantized is generated, it calls at least one large model quantization method in the script file to perform model quantization on the large model to be quantized respectively, generating at least one set of candidate model quantization parameters.

[0056] After generating at least one set of candidate model quantization parameters, further, the cloud server determines the candidate model quantization parameter that is most suitable for the local terminal from each group of candidate model quantization parameters as the target model quantization parameter.

[0057] Among them, the determination methods of the target model quantization parameters include, but are not limited to:

[0058] 1) The cloud server determines the pre-calibrated precision limit conditions, hardware limit conditions, and real-time limit conditions. Further, among the quantization parameters of each group of candidate models, the quantization parameters of the candidate models whose parameter precision meets the precision limit conditions, the adapted hardware meets the hardware limit conditions, and the inference real-time performance meets the real-time limit conditions are used as the target model quantization parameters.

[0059] 2) The cloud server determines the precision loss score, resource consumption score, and inference efficiency score of the quantization parameters of each group of candidate models, and determines the target model quantization parameters from the quantization parameters of each group of candidate models according to the weighted results of the precision loss score, resource consumption score, and inference efficiency score.

[0060] S103. Send the target model quantization parameters to the local terminal, so that the local terminal performs model deployment on the large model to be deployed according to the target model quantization parameters.

[0061] Among them, the local terminal refers to the terminal device that needs to deploy the large model to be deployed, including but not limited to personal computers, smart phones, edge computing devices, intelligent vehicle cockpits, and robots, etc. The type of the local terminal is not limited in this embodiment.

[0062] In one implementation, the local terminal sends a quantization parameter acquisition request to the cloud server. In response to the quantization parameter acquisition request, the cloud server sends the target model quantization parameters to the local terminal. After receiving the target model quantization parameters sent by the cloud server, the local terminal uses the target model quantization parameters to replace the parameters of the large model to be deployed for model deployment of the large model to be deployed.

[0063] The present invention generates a large model to be quantized with the same model structure as the large model to be deployed in the local terminal in the cloud server, and the cloud server performs model quantization on the large model to be quantized, and finally sends the obtained target model quantization parameters to the local terminal for model deployment. The beneficial effects are as follows:

[0064] First, since the cloud server has powerful data processing performance, the cloud server is used for model quantization, and then the local terminal performs model deployment according to the target model quantization parameters obtained by quantization, so that the model quantization process is not restricted by the local terminal hardware, improving the efficiency and precision of model quantization, and further improving the inference performance and efficiency of the large model to be deployed.

[0065] Second, since the large model to be quantized in the cloud server has the same model structure as the large model to be deployed in the local terminal, the effect of aligning the model structure between the cloud side and the terminal side is achieved, ensuring the credibility of the model quantization result in the cloud server.

[0066] Embodiment 2

[0067] Figure 2 This is a flowchart of a method for determining target model quantization parameters provided in the second embodiment of the present invention. This embodiment further optimizes and expands the step of "determining target model quantization parameters from candidate model quantization parameters" in the above embodiment, and can be combined with each of the above optional implementation manners. As Figure 2 shown, the method includes:

[0068] S201. Respectively perform accuracy loss evaluation on at least one group of candidate model quantization parameters to generate accuracy loss scores corresponding to each group of candidate model quantization parameters.

[0069] In one implementation manner, the cloud server runs the large model to be quantized before quantization and the large model after quantization on the same test set, and respectively performs accuracy loss evaluation on each group of candidate model quantization parameters according to the difference in prediction accuracy rates of the large model before and after quantization, to generate accuracy loss scores corresponding to each group of candidate model quantization parameters. It can be understood that the smaller the difference in prediction accuracy rates of the large model after quantization using any group of candidate model quantization parameters, the smaller the accuracy loss, and further the higher the accuracy loss score.

[0070] S202. Respectively perform inference efficiency evaluation on at least one group of candidate model quantization parameters to generate inference efficiency scores corresponding to each group of candidate model quantization parameters.

[0071] In one implementation manner, the cloud server tests the latency and throughput of the large model corresponding to each group of candidate model quantization parameters, and respectively performs inference efficiency evaluation on each group of candidate model quantization parameters according to the test results of latency and throughput, to generate inference efficiency scores corresponding to each group of candidate model quantization parameters.

[0072] Among them, the test of latency can be determined by the time taken for a single inference of the large model from input data to output result. The test of throughput can be determined by the number of samples processed by the large model per unit time (second).

[0073] It can be understood that the smaller the latency, the higher the inference efficiency, and further the higher the inference efficiency score. The larger the throughput, the higher the inference efficiency, and further the higher the inference efficiency score.

[0074] S203. Respectively perform resource consumption evaluation on at least one group of candidate model quantization parameters to generate resource consumption scores corresponding to each group of candidate model quantization parameters.

[0075] In one implementation, the cloud server tests the memory occupancy and energy consumption efficiency of the large model using the quantization parameters of each group of candidate models, and based on the test results of the memory occupancy and energy consumption efficiency, respectively evaluates the resource consumption of each group of candidate model quantization parameters to generate resource consumption scores corresponding to each group of candidate model quantization parameters.

[0076] Among them, the test of memory occupancy can be determined by the peak value of the video memory / RAM during the operation of the large model. The test of energy consumption efficiency can be determined by the power consumption under the unit calculation amount of the large model.

[0077] It can be understood that the less the memory occupancy, the less the resource consumption, and further the higher the resource consumption score. The higher the energy consumption efficiency, the less the resource consumption, and further the higher the resource consumption score.

[0078] S204. Determine the target model quantization parameters from the candidate model quantization parameters according to the accuracy loss scores, inference efficiency scores, and resource consumption scores corresponding to each group of candidate model quantization parameters.

[0079] In one implementation, the cloud server obtains the preset accuracy loss score threshold, inference efficiency score threshold, and resource consumption score threshold, further compares the accuracy loss score with the accuracy loss score threshold, the inference efficiency score with the inference efficiency score threshold, and the resource consumption score with the resource consumption score threshold respectively, and then uses the candidate model quantization parameters whose accuracy loss score meets the accuracy loss score threshold, inference efficiency score meets the inference efficiency score threshold, and resource consumption score meets the resource consumption score threshold as the target model quantization parameters.

[0080] In another implementation, the cloud server performs a weighted sum of the accuracy loss scores, inference efficiency scores, and resource consumption scores corresponding to each group of candidate model quantization parameters, and determines the target model quantization parameters from the candidate model quantization parameters according to the weighted sum result.

[0081] By respectively performing accuracy loss evaluation on at least one group of candidate model quantization parameters to generate accuracy loss scores corresponding to each group of candidate model quantization parameters; respectively performing inference efficiency evaluation on at least one group of candidate model quantization parameters to generate inference efficiency scores corresponding to each group of candidate model quantization parameters; respectively performing resource consumption evaluation on at least one group of candidate model quantization parameters to generate resource consumption scores corresponding to each group of candidate model quantization parameters; and determining the target model quantization parameters from the candidate model quantization parameters according to the accuracy loss scores, inference efficiency scores, and resource consumption scores corresponding to each group of candidate model quantization parameters, the beneficial effects are as follows:

[0082] First, achieve the optimal balance of the multi-dimensional performance of the target model quantization parameters and avoid the limitations of optimizing a single indicator. Second, improve the adaptability of the model to the hardware environment. Third, reduce the trial-and-error cost and development cycle. Fourth, the quantization economic benefits are significant.

[0083] Optionally, determine the target model quantization parameters from the candidate model quantization parameters according to the accuracy loss scores, inference efficiency scores, and resource consumption scores respectively corresponding to each group of candidate model quantization parameters, including:

[0084] S2041. Perform weighted calculation on the accuracy loss scores respectively corresponding to each group of candidate model quantization parameters to determine the first weighted scores respectively corresponding to each group of candidate model quantization parameters.

[0085] S2042. Perform weighted calculation on the inference efficiency scores respectively corresponding to each group of candidate model quantization parameters to determine the second weighted scores respectively corresponding to each group of candidate model quantization parameters.

[0086] S2043. Perform weighted calculation on the resource consumption scores respectively corresponding to each group of candidate model quantization parameters to determine the third weighted scores respectively corresponding to each group of candidate model quantization parameters.

[0087] Exemplarily, assume that the accuracy loss score of any group of candidate model quantization parameters is X1, the inference efficiency score is X2, and the resource consumption score is X3. Assume that the preset accuracy loss score weight is a1, the inference efficiency score weight is a2, and the resource consumption score weight is a3. Then the first weighted score of this group of candidate model quantization parameters is a1×X1; the second weighted score is a2×X2; the third weighted score is a3×X3.

[0088] S2044. Determine the weighted total scores respectively corresponding to each group of candidate model quantization parameters according to the first weighted scores, second weighted scores, and third weighted scores respectively corresponding to each group of candidate model quantization parameters, and determine the target model quantization parameters from the candidate model quantization parameters according to the weighted total scores.

[0089] In one implementation, sum the first weighted scores, second weighted scores, and third weighted scores respectively corresponding to each group of candidate model quantization parameters to determine the weighted total scores respectively corresponding to each group of candidate model quantization parameters. Further, determine the target model quantization parameters from the candidate model quantization parameters according to the sorting results of the weighted total scores of each group of candidate model quantization parameters. For example, use the candidate model quantization parameters with the highest weighted total score as the target model quantization parameters.

[0090] By performing weighted calculations on the accuracy loss scores corresponding to the quantization parameters of each group of candidate models respectively, the first weighted scores corresponding to the quantization parameters of each group of candidate models are determined; by performing weighted calculations on the inference efficiency scores corresponding to the quantization parameters of each group of candidate models respectively, the second weighted scores corresponding to the quantization parameters of each group of candidate models are determined; by performing weighted calculations on the resource consumption scores corresponding to the quantization parameters of each group of candidate models respectively, the third weighted scores corresponding to the quantization parameters of each group of candidate models are determined; according to the first weighted scores, second weighted scores and third weighted scores corresponding to the quantization parameters of each group of candidate models respectively, the weighted total scores corresponding to the quantization parameters of each group of candidate models are determined, and the target model quantization parameters are determined from the candidate model quantization parameters. On the one hand, it can avoid the negative impact caused by the deviation of a single indicator, such as the accuracy collapse caused by only pursuing low latency, or the deployment failure caused by excessive model compression, etc.; on the other hand, it can also enhance the dynamic business adaptation ability. For example, the weight allocation can be dynamically adjusted according to different business requirements, or the weights can be adjusted to adapt to different deployment platforms, etc.

[0091] Embodiment 3

[0092] Figure 3 The flowchart of a deployment method for a large model provided in Embodiment 3 of the present invention further optimizes and expands the above embodiments and can be combined with the above various optional implementation manners. As Figure 3 shown, the method includes:

[0093] S301. In response to a user's structure configuration operation on the initial large model in the cloud server, a large model to be quantized is generated.

[0094] Among them, the large model to be quantized has the same model structure as the large model to be deployed in the local terminal.

[0095] S302. Use at least one large model quantization method to perform model quantization on the large model to be quantized respectively, generate at least one group of candidate model quantization parameters, and determine the target model quantization parameters from the candidate model quantization parameters.

[0096] S303. Store the target model quantization parameters in the cloud server using the target storage path.

[0097] Among them, the target storage path is a storage path allocated by the cloud server for storing the target model quantization parameters.

[0098] S304. Establish a target association relationship between the target storage path and the target terminal identifier of the local terminal, and add the target association relationship to the association relationship set.

[0099] Among them, the association relationship set includes the association relationship between the terminal identifiers of each local terminal and the storage paths storing the model quantization parameters corresponding to each local terminal. The form of the association relationship set can be a KV key-value pair set.

[0100] In one implementation, the cloud server establishes a KV key-value pair between the target storage path and the target terminal identifier, where the target terminal identifier is used as the Key value and the target storage path is used as the Value value. Further, the KV key-value pair is added to the KV key-value pair set.

[0101] By storing the target model quantization parameters in the cloud server using the target storage path, establishing the target association relationship between the target storage path and the target terminal identifier of the local terminal, and adding the target association relationship to the association relationship set, it is possible to finely control the storage path and permissions, and also optimize the parameter retrieval and loading efficiency.

[0102] S305. Perform permission verification on the local terminal according to the to-be-identified terminal identifier sent by the local terminal, and in the case of successful verification, perform matching in the association relationship set according to the to-be-identified terminal identifier.

[0103] Among them, the to-be-identified terminal identifier refers to the un-verified terminal identifier sent by the local terminal.

[0104] In one implementation, the cloud server receives the parameter acquisition request sent by the local terminal, and parses the parameter acquisition request to obtain the to-be-identified terminal identifier carried therein. Further, the cloud server performs permission verification on the local terminal according to the to-be-identified terminal identifier, and in the case of successful verification, performs matching of the to-be-identified terminal identifier in the association relationship set.

[0105] Optionally, the to-be-identified terminal identifier includes at least one of terminal hardware characteristics, software environment characteristics, and behavior pattern characteristics.

[0106] Among them, the terminal hardware characteristics refer to the unique attributes and technical specifications of the local terminal at the physical level, including but not limited to hardware serial numbers, chip identifiers, storage media, CPU architectures, RAM specifications, etc. The software environment characteristics refer to the set of all software-related configurations, dependencies, and interaction conditions that affect software operation and functions in a specific computing environment of the local terminal, including but not limited to operating system types and versions, system services and drivers, programming language environments, environment variables, and configuration files, etc. The behavior pattern characteristics refer to the regular operation habits, interaction methods, and resource calling rules shown by the local terminal during use, including but not limited to application startup frequencies and time period distributions, background process residence rules, data transmission characteristics, touch operation hot zone distributions, power consumption patterns, etc.

[0107] Optionally, perform permission verification on the local terminal according to the terminal identifier to be recognized sent by the local terminal, including:

[0108] Generate a dynamic fingerprint to be recognized corresponding to the local terminal according to the terminal hardware characteristics, software environment characteristics, and behavior pattern characteristics; perform a similarity match between the dynamic fingerprint to be recognized and the registered dynamic fingerprint of the local terminal. If the similarity between the dynamic fingerprint to be recognized and the registered dynamic fingerprint is greater than the threshold, it is determined that the local terminal passes the verification.

[0109] In one implementation, the cloud server uses a dynamic fingerprint generation algorithm to calculate the dynamic fingerprint according to the terminal hardware characteristics, software environment characteristics, and behavior pattern characteristics, and generates a dynamic fingerprint to be recognized corresponding to the local terminal. Further, the cloud server performs a similarity match between the dynamic fingerprint to be recognized and the registered dynamic fingerprint of the local terminal. If the similarity between the dynamic fingerprint to be recognized and the registered dynamic fingerprint is less than or equal to the threshold, it is determined that the local terminal fails the verification; if the similarity between the dynamic fingerprint to be recognized and the registered dynamic fingerprint is greater than the threshold, it is determined that the local terminal passes the verification.

[0110] By generating a dynamic fingerprint to be recognized corresponding to the local terminal according to the terminal hardware characteristics, software environment characteristics, and behavior pattern characteristics; performing a similarity match between the dynamic fingerprint to be recognized and the registered dynamic fingerprint of the local terminal, and if the similarity between the dynamic fingerprint to be recognized and the registered dynamic fingerprint is greater than the threshold, it is determined that the local terminal passes the verification. Since the dynamic fingerprint is updated in real time according to the local terminal status (such as software upgrade, hardware wear), it is more difficult to be fixed and cracked compared to static identifiers (such as device serial numbers), enhancing the accuracy and security of terminal identity recognition.

[0111] S306. Determine the target association relationship from the association relationship set according to the matching result, and determine the target storage path according to the target association relationship.

[0112] S307. Obtain the target model quantization parameter according to the target storage path, and send the target model quantization parameter to the local terminal.

[0113] Exemplarily, assume that the terminal identifier to be recognized is "xxyyzz", and assume that the target storage path associated with the target terminal identifier "xxyyzz" in the association relationship set is "aa: / / bb / cc / dd", then obtain the target model quantization parameter according to the target storage path "aa: / / bb / cc / dd", and send the target model quantization parameter to the local terminal.

[0114] By performing permission verification on the local terminal according to the terminal identifier to be recognized sent by the local terminal, and when the verification passes, matching according to the terminal identifier to be recognized in the association relationship set; determining the target association relationship from the association relationship set according to the matching result, and determining the target storage path according to the target association relationship; obtaining the target model quantization parameters according to the target storage path, and sending the target model quantization parameters to the local terminal, the beneficial effects are as follows:

[0115] Firstly, it ensures that only authorized devices can obtain the target model quantization parameters, effectively preventing unauthorized terminals from stealing or tampering with sensitive quantization configurations and reducing the risk of model leakage.

[0116] Secondly, the association relationship set between the target storage path and the terminal identifier is stored in the form of key-value pairs, reducing the query latency compared to traditional traversal algorithms and being suitable for the real-time parameter distribution requirements in high-concurrency scenarios.

[0117] Thirdly, it realizes the distributed storage and on-demand loading of parameter files, avoiding the cloud storage redundancy caused by preloading all parameters and reducing the storage space occupation.

[0118] Optionally, the method further includes:

[0119] A. Using the first model conversion method to perform model conversion on the large model to be quantized, generating the first converted large model.

[0120] Among them, the first model conversion method is the same as the second model conversion method adopted by the local terminal, and the second model conversion method is used to perform model conversion on the large model to be deployed.

[0121] Model conversion is the process of converting a large model from one form to another, usually involving operations such as framework adaptation, format compatibility, and hardware optimization.

[0122] Different manufacturers often adopt different settings and strategies during the model conversion process. For example, whether to perform suppress processing on the input and output, etc. Different conversion methods will have different impacts on the inference performance and accuracy of the model. Therefore, it is necessary to carefully align these differences. To ensure that the converted model performs the same as the simulation on the cloud side at the edge side, users must deeply understand the conversion processes and parameter settings of each manufacturer to ensure that no unnecessary accuracy loss or performance bottleneck is introduced during the conversion process. In addition, according to the characteristics of different platforms, adjust the conversion strategy so that the finally obtained model can fully utilize the hardware advantages at the edge side, thereby achieving the optimal inference effect.

[0123] B. Using at least one large model quantization method to perform model quantization on the first converted large model respectively, generating at least one set of first model quantization parameters, and determining the preferred model quantization parameters from the first model quantization parameters.

[0124] Among them, the generation method of the first model quantization parameters and the determination method of the preferred model quantization parameters can refer to the methods described in other embodiments, which will not be elaborated here.

[0125] C. Send the preferred model quantization parameters to the local terminal, so that the local terminal performs model deployment on the large model to be deployed according to the preferred model quantization parameters.

[0126] By using the first model conversion method to perform model conversion on the large model to be quantized, a first converted large model is generated; among them, the first model conversion method is the same as the second model conversion method adopted by the local terminal, and the second model conversion method is used to perform model conversion on the large model to be deployed. The beneficial effects are as follows:

[0127] In the first aspect, through a unified conversion method, the parameter optimization in the quantization stage can be directly reused in the deployment stage, avoiding repeated calculations.

[0128] In the second aspect, the unified method ensures that the numerical truncation rules in the quantization stage are strictly consistent with the computing precision of the local terminal, eliminating the cumulative error caused by differences in conversion tools in the traditional solution.

[0129] Embodiment 4

[0130] Figure 4 The flowchart of a method for deploying a large model provided in Embodiment 4 of the present invention. This embodiment is applicable to the situation where the local terminal uses the target model quantization parameters obtained from the cloud server to perform model deployment on the large model to be deployed in the local terminal. This method can be executed by a large model deployment device, and the large model deployment device is configured in the local terminal and can be implemented in the form of hardware and / or software. As Figure 4 shown, the method includes:

[0131] S401. Obtain the target model quantization parameters sent by the cloud server.

[0132] Among them, the target model quantization parameters are obtained by the cloud server performing model quantization on the large model to be quantized in the cloud server, and the large model to be quantized has the same model structure as the large model to be deployed in the local terminal.

[0133] In one implementation, after the user analyzes and understands the operator effect and computational graph corresponding to the large model to be deployed, a structure configuration operation is performed on the initial large model in the cloud server. The cloud server responds to the user's structure configuration operation on the initial large model and correspondingly generates a large model to be quantized that has the same model structure as the large model to be deployed in the local terminal.

[0134] The cloud server uses at least one large model quantization method to perform model quantization on the large model to be quantized respectively, generates at least one set of candidate model quantization parameters, and determines the target model quantization parameters from the candidate model quantization parameters.

[0135] The local terminal sends a quantization parameter acquisition request to the cloud server, and in response to the quantization parameter acquisition request, the cloud server sends the target model quantization parameters to the local terminal.

[0136] S402. Perform model deployment on the large model to be deployed according to the target model quantization parameters.

[0137] In one implementation, after receiving the target model quantization parameters sent by the cloud server, the local terminal uses the target model quantization parameters to perform parameter replacement on the large model to be deployed, for performing model deployment on the large model to be deployed.

[0138] By obtaining the target model quantization parameters sent by the cloud server; wherein, the target model quantization parameters are obtained by the cloud server performing model quantization on the large model to be quantized in the cloud server, and the large model to be quantized has the same model structure as the large model to be deployed in the local terminal, and performing model deployment on the large model to be deployed according to the target model quantization parameters, the beneficial effects are as follows:

[0139] First, since the cloud server has powerful data processing performance, model quantization is performed using the cloud server, and then the local terminal performs model deployment according to the obtained target model quantization parameters, so that the model quantization process is not restricted by the local terminal hardware, improving the efficiency and accuracy of model quantization, and further improving the inference performance and efficiency of the large model to be deployed.

[0140] Second, since the large model to be quantized in the cloud server has the same model structure as the large model to be deployed in the local terminal, the effect of aligning the model structure between the cloud side and the terminal side is achieved, ensuring the credibility of the model quantization result in the cloud server.

[0141] Optionally, performing model deployment on the large model to be deployed according to the target model quantization parameters includes:

[0142] Determine the target chip manufacturer to which the processor of the local terminal belongs, and determine the target model programming interface from the candidate model programming interfaces according to the manufacturer identification information of the target chip manufacturer; call the target model programming interface to replace the current model parameters of the large model to be deployed with the target model quantization parameters, generate an optimized large model, and deploy the optimized large model in the local terminal.

[0143] Among them, the target chip manufacturer refers to an enterprise that designs or manufactures local terminal processors, and its products (such as CPUs, GPUs, NPUs) are the core computing units of electronic devices. The manufacturer identification information is a data set used to uniquely identify the identity of the target chip manufacturer, usually existing in the form of codes, labels, certification documents, etc. The target model programming interface is a standardized interaction protocol for users to call, manage, and deploy the large model to be deployed, and is a set of predefined functions, classes, or protocols that provide programmatic access to functions such as model inference, training, and monitoring. The current model parameters refer to the model parameters in the large model to be deployed that have not been model-quantized.

[0144] In one implementation, the local terminal obtains the unique identification information of the chip manufacturer by reading the processor hardware identification register or calling a system interface (such as / proc / cpuinfo), and determines the target chip manufacturer to which the processor of the local terminal belongs according to the unique identification information. Further, according to the manufacturer identification information of the target chip manufacturer, the corresponding target model programming interface is selected from the candidate model programming interfaces, and the current model parameters of the large model to be deployed are replaced with the target model quantization parameters through the quantization parameter conversion function (such as convert_weights_to_quantized()) of the target model programming interface. The local terminal links the quantized large model to be deployed with the runtime library provided by the target chip manufacturer and compiles it to generate a lightweight model file that can run on the local terminal.

[0145] By determining the target chip manufacturer to which the processor of the local terminal belongs, and determining the target model programming interface from the candidate model programming interfaces according to the manufacturer identification information of the target chip manufacturer; calling the target model programming interface to replace the current model parameters of the large model to be deployed with the target model quantization parameters to generate an optimized large model, and deploying the optimized large model on the local terminal can, on the one hand, improve the deployment efficiency and performance of the large model on the local terminal side, and on the other hand, enhance cross-platform compatibility and development flexibility.

[0146] Optionally, deploying the optimized large model on the local terminal includes:

[0147] A1. Analyze the current computation graph corresponding to the optimized large model, and serialize the current computation graph to generate a current sequence value.

[0148] Among them, the computation graph is a graphical representation method used to describe the computation process of the large model. It represents the data flow and computation dependency relationships between the layers in the large model through nodes and edges. Each node usually represents a computation operation (such as convolution, matrix multiplication, etc.), and the edges represent the transfer of data between these operations.

[0149] In one embodiment, the local terminal uses a lightweight graph parsing framework to parse the computational graph of the optimized large model to determine the current computational graph corresponding to the optimized large model. Further, the current computational graph is converted into a standardized intermediate representation, and a unified encoding rule for nodes, edges, and weights is defined. Then, the current computational graph is serialized and encoded according to the encoding rule to generate a current sequence value.

[0150] B1. Calculate the hash value of the current sequence value to generate a current hash value, and determine whether the current inference logic of the optimized large model is the same as the benchmark inference logic according to the comparison result between the benchmark hash value and the current hash value.

[0151] Among them, the benchmark hash is calculated based on the benchmark sequence value corresponding to the benchmark computational graph, and the benchmark computational graph reflects the benchmark inference logic of the large model, that is, the inference logic that meets the expectations.

[0152] In one embodiment, a preset hash algorithm is used to calculate the hash value of the current sequence value to generate a current hash value, and the current hash value is compared with the benchmark hash value to determine whether the benchmark hash value and the current hash value are the same, and further determine whether the current inference logic is the same as the benchmark inference logic.

[0153] C1. If the current inference logic is the same as the benchmark inference logic, deploy the optimized large model on the local terminal.

[0154] In one embodiment, if the benchmark hash value and the current hash value are different, it means that the current inference logic is different from the benchmark inference logic, and then the technical personnel are notified to adjust the model code of the optimized large model so that the current inference logic is the same as the benchmark inference logic. If the benchmark hash value and the current hash value are the same, it means that the current inference logic is the same as the benchmark inference logic, and then the optimized large model is deployed on the local terminal.

[0155] By parsing the current computational graph corresponding to the optimized large model, serializing the current computational graph to generate a current sequence value, calculating the hash value of the current sequence value to generate a current hash value, and determining whether the current inference logic of the optimized large model is the same as the benchmark inference logic according to the comparison result between the benchmark hash value and the current hash value. If the current inference logic is the same as the benchmark inference logic, deploy the optimized large model on the local terminal. On the one hand, it can ensure that the inference logic of the optimized large model meets the expected inference logic. On the other hand, it can confirm through hash calculation that the computational graph has not been maliciously tampered with, avoid the leakage of privacy data, and improve the security of end-side deployment.

[0156] Optionally, deploying the optimized large model on the local terminal includes:

[0157] A2. Determine the first model vocabulary corresponding to the optimized large model, and determine the second model vocabulary corresponding to the large model to be quantized.

[0158] Among them, the model vocabulary, also known as Token ID, is the core tool for the model to process natural language, used to convert text into discrete symbols that the model can understand. Its essence is a preset symbol set, which determines the parsing granularity and semantic understanding ability of the model for the input text.

[0159] In one implementation, the local terminal determines the first model vocabulary corresponding to the optimized large model and sends a vocabulary acquisition request to the cloud server to obtain the second model vocabulary corresponding to the large model to be quantized from the cloud server.

[0160] B2. If the first model vocabulary is the same as the second model vocabulary, deploy the optimized large model on the local terminal.

[0161] In one implementation, compare the first model vocabulary with the second model vocabulary. If the first model vocabulary is different from the second model vocabulary, notify the technician to adjust the model code of the optimized large model so that the first model vocabulary is the same as the second model vocabulary. If the first model vocabulary is the same as the second model vocabulary, deploy the optimized large model on the local terminal.

[0162] By determining the first model vocabulary corresponding to the optimized large model and the second model vocabulary corresponding to the large model to be quantized, if the first model vocabulary is the same as the second model vocabulary, deploy the optimized large model on the local terminal. On the one hand, it ensures the alignment of the model vocabulary on the cloud side and the end side, and guarantees the accuracy of the model inference logic; on the other hand, it enhances the cross-platform deployment compatibility.

[0163] Embodiment Five

[0164] Figure 5 The structure diagram of a large model deployment device provided by Embodiment Five of the present invention. This device is configured on the cloud server and is applicable to the situation where the local terminal uses the target model quantization parameters obtained from the cloud server to perform model deployment on the large model to be deployed in the local terminal, as Figure 5 shown. This device includes:

[0165] A large model to be quantized generation module 51, configured to generate a large model to be quantized in response to a user's structure configuration operation on the initial large model in the cloud server; wherein, the large model to be quantized has the same model structure as the large model to be deployed in the local terminal;

[0166] The model quantization module 52 is used to perform model quantization on the to-be-quantized large model by using at least one large model quantization method, generate at least one set of candidate model quantization parameters, and determine target model quantization parameters from the candidate model quantization parameters;

[0167] The model quantization parameter sending module 53 is used to send the target model quantization parameters to the local terminal, so that the local terminal performs model deployment on the to-be-deployed large model according to the target model quantization parameters.

[0168] Optionally, the model quantization module 52 is specifically used for:

[0169] Perform accuracy loss evaluation on each of the at least one set of candidate model quantization parameters, and generate an accuracy loss score corresponding to each set of the candidate model quantization parameters;

[0170] Perform inference efficiency evaluation on each of the at least one set of candidate model quantization parameters, and generate an inference efficiency score corresponding to each set of the candidate model quantization parameters;

[0171] Perform resource consumption evaluation on each of the at least one set of candidate model quantization parameters, and generate a resource consumption score corresponding to each set of the candidate model quantization parameters;

[0172] Determine target model quantization parameters from the candidate model quantization parameters according to the accuracy loss scores, the inference efficiency scores, and the resource consumption scores corresponding to each set of the candidate model quantization parameters.

[0173] Optionally, the model quantization module 52 is specifically further used for:

[0174] Perform weighted calculation on the accuracy loss scores corresponding to each set of the candidate model quantization parameters to determine a first weighted score corresponding to each set of the candidate model quantization parameters;

[0175] Perform weighted calculation on the inference efficiency scores corresponding to each set of the candidate model quantization parameters to determine a second weighted score corresponding to each set of the candidate model quantization parameters;

[0176] Perform weighted calculation on the resource consumption scores corresponding to each set of the candidate model quantization parameters to determine a third weighted score corresponding to each set of the candidate model quantization parameters;

[0177] Determine a weighted total score corresponding to each set of the candidate model quantization parameters according to the first weighted score, the second weighted score, and the third weighted score corresponding to each set of the candidate model quantization parameters, and determine target model quantization parameters from the candidate model quantization parameters according to the weighted total score.

[0178] Optionally, the device further includes an association relationship establishment module, which is specifically configured to:

[0179] Store the target model quantization parameters in the cloud server using the target storage path;

[0180] Establish a target association relationship between the target storage path and the target terminal identifier of the local terminal, and add the target association relationship to the association relationship set;

[0181] Optionally, the model quantization parameter sending module 53 is specifically configured to:

[0182] Verify the permissions of the local terminal according to the terminal identifier to be recognized sent by the local terminal, and perform matching in the association relationship set according to the terminal identifier to be recognized when the verification is passed;

[0183] Determine the target association relationship from the association relationship set according to the matching result, and determine the target storage path according to the target association relationship;

[0184] Obtain the target model quantization parameters according to the target storage path, and send the target model quantization parameters to the local terminal.

[0185] Optionally, the terminal identifier to be recognized includes at least one of a terminal hardware feature, a software environment feature, and a behavior pattern feature;

[0186] The model quantization parameter sending module 53 is further specifically configured to:

[0187] Generate a dynamic fingerprint to be recognized corresponding to the local terminal according to the terminal hardware feature, the software environment feature, and the behavior pattern feature;

[0188] Perform a similarity match between the dynamic fingerprint to be recognized and the registered dynamic fingerprint of the local terminal. If the similarity between the dynamic fingerprint to be recognized and the registered dynamic fingerprint is greater than a threshold, it is determined that the local terminal passes the verification.

[0189] Optionally, the device further includes a model conversion module, which is specifically configured to:

[0190] Perform model conversion on the large model to be quantized using a first model conversion method to generate a first converted large model; wherein, the first model conversion method is the same as a second model conversion method used by the local terminal, and the second model conversion method is used to perform model conversion on the large model to be deployed;

[0191] Use at least one large model quantization method to perform model quantization on the first converted large model respectively, generate at least one set of first model quantization parameters, and determine the preferred model quantization parameters from the first model quantization parameters;

[0192] Send the preferred model quantization parameters to the local terminal, so that the local terminal performs model deployment on the large model to be deployed according to the preferred model quantization parameters.

[0193] The large model deployment device provided in the fifth embodiment of the present invention can execute the large model deployment method provided in the first to third embodiments of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0194] Embodiment Six

[0195] Figure 6 FIG. is a schematic structural diagram of a large model deployment device provided in the sixth embodiment of the present invention. The device is configured in a local terminal and is applicable to the situation where the local terminal uses the target model quantization parameters obtained from a cloud server to perform model deployment on the large model to be deployed in the local terminal, as Figure 6 shown. The device includes:

[0196] A model quantization parameter acquisition module 61, configured to acquire target model quantization parameters sent by a cloud server; wherein, the target model quantization parameters are obtained by the cloud server performing model quantization on the large model to be quantized in the cloud server, and the large model to be quantized has the same model structure as the large model to be deployed in the local terminal;

[0197] A model deployment module 62, configured to perform model deployment on the large model to be deployed according to the target model quantization parameters.

[0198] Optionally, the model deployment module 62 is specifically configured to:

[0199] Determine the target chip manufacturer to which the processor of the local terminal belongs, and determine the target model programming interface from the candidate model programming interfaces according to the manufacturer identification information of the target chip manufacturer;

[0200] Call the target model programming interface to replace the current model parameters of the large model to be deployed with the target model quantization parameters, generate an optimized large model, and deploy the optimized large model in the local terminal.

[0201] Optionally, the model deployment module 62 is specifically further configured to:

[0202] Parse the current computation graph corresponding to the optimized large model, and serialize the current computation graph to generate a current sequence value;

[0203] Perform a hash calculation on the current sequence value to generate a current hash value, and determine whether the current inference logic of the optimized large model is the same as the benchmark inference logic according to the comparison result between the benchmark hash value and the current hash value;

[0204] If the current inference logic is the same as the benchmark inference logic, deploy the optimized large model in the local terminal.

[0205] Optionally, the model deployment module 62 is further specifically configured to:

[0206] Determine the first model vocabulary corresponding to the optimized large model, and determine the second model vocabulary corresponding to the large model to be quantized;

[0207] If the first model vocabulary is the same as the second model vocabulary, deploy the optimized large model in the local terminal.

[0208] The large model deployment device provided in Embodiment VI of the present invention can execute the large model deployment method provided in Embodiment IV of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0209] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0210] Embodiment VII

[0211] Figure 7 FIG. shows a schematic structural diagram of an electronic device 70 that can be used to implement an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0212] As Figure 7As shown, the electronic device 70 includes at least one processor 71 and a memory communicatively connected to the at least one processor 71, such as a read-only memory (ROM) 72, a random access memory (RAM) 73, etc. The memory stores a computer program executable by the at least one processor. The processor 71 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 72 or the computer program loaded from the storage unit 78 into the random access memory (RAM) 73. In the RAM 73, various programs and data required for the operation of the electronic device 70 can also be stored. The processor 71, the ROM 72, and the RAM 73 are connected to each other via a bus 74. An input / output (I / O) interface 75 is also connected to the bus 74.

[0213] Multiple components in the electronic device 70 are connected to the I / O interface 75, including: an input unit 76, such as a keyboard, a mouse, etc.; an output unit 77, such as various types of displays, speakers, etc.; a storage unit 78, such as a disk, an optical disc, etc.; and a communication unit 79, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 79 allows the electronic device 70 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0214] The processor 71 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 71 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 71 executes the various methods and processes described above, such as the method for deploying a large model.

[0215] In some embodiments, the method for deploying a large model can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 78. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 70 via the ROM 72 and / or the communication unit 79. When the computer program is loaded into the RAM 73 and executed by the processor 71, one or more steps of the method for deploying a large model described above can be executed. Alternatively, in other embodiments, the processor 71 can be configured to execute the method for deploying a large model by any other appropriate means (e.g., by means of firmware).

[0216] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0217] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0218] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0219] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0220] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0221] The computing system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0222] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.

[0223] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A large model deployment method, characterized in that: Executed by a cloud server, the method includes: In response to the user's structural configuration operation on the initial large model in the cloud server, a large model to be quantized is generated; wherein the large model to be quantized has the same model structure as the large model to be deployed in the local terminal; Using at least one large model quantization method to perform model quantization on the large models to be quantized, generate at least one set of candidate model quantization parameters, and determine target model quantization parameters from the candidate model quantization parameters; The target model quantization parameters are sent to the local terminal, so that the local terminal performs model deployment on the large model to be deployed according to the target model quantization parameters.

2. The method according to claim 1, characterized in that The step of determining the target model quantization parameter from the candidate model quantization parameters comprises: Performing precision loss evaluation on the at least one group of candidate model quantization parameters respectively, and generating precision loss scores corresponding to each group of candidate model quantization parameters respectively; Performing inference efficiency evaluation on the at least one group of candidate model quantization parameters respectively, and generating inference efficiency scores corresponding to each group of candidate model quantization parameters respectively; Performing resource consumption evaluation on the at least one group of candidate model quantization parameters respectively, and generating resource consumption scores corresponding to each group of candidate model quantization parameters respectively; According to the precision loss score, the inference efficiency score and the resource consumption score respectively corresponding to each group of the candidate model quantization parameters, the target model quantization parameter is determined from the candidate model quantization parameters.

3. The method according to claim 2, characterized in that The step of determining the target model quantization parameter from the candidate model quantization parameters according to the precision loss score, the inference efficiency score, and the resource consumption score respectively corresponding to each group of the candidate model quantization parameters includes: Performing weighted calculation on the precision loss scores respectively corresponding to the candidate model quantization parameters of each group, to determine first weighted scores respectively corresponding to the candidate model quantization parameters of each group; Performing weighted calculation on the inference efficiency scores respectively corresponding to the candidate model quantization parameters of each group, to determine second weighted scores respectively corresponding to the candidate model quantization parameters of each group; Performing weighted calculation on the resource consumption scores respectively corresponding to the candidate model quantization parameters of each group, to determine third weighted scores respectively corresponding to the candidate model quantization parameters of each group; According to the first weighted score, the second weighted score and the third weighted score respectively corresponding to each group of the candidate model quantization parameters, the weighted total score corresponding to each group of the candidate model quantization parameters is determined, and the target model quantization parameter is determined from the candidate model quantization parameters according to the weighted total score.

4. The method according to claim 1, after determining the target model quantization parameter from the candidate model quantization parameters, further comprising: Using a target storage path to store the target model quantization parameters in the cloud server; Establishing a target association relationship between the target storage path and the target terminal identifier of the local terminal, and adding the target association relationship to the association relationship set; The sending the target model quantization parameter to the local terminal includes: Performing an authority check on the local terminal according to the terminal identifier to be identified sent by the local terminal, and matching the terminal identifier to be identified in the association relationship set if the check passes; Determine the target association relationship from the association relationship set according to the matching result, and determine the target storage path according to the target association relationship; The target model quantization parameter is acquired according to the target storage path, and the target model quantization parameter is sent to the local terminal.

5. The method according to claim 4, characterized in that The terminal identification to be identified includes at least one of terminal hardware features, software environment features and behavior pattern features; The performing authority verification on the local terminal according to the terminal identifier to be identified sent by the local terminal includes: Generate a dynamic fingerprint to be identified corresponding to the local terminal according to the terminal hardware characteristics, the software environment characteristics and the behavior pattern characteristics; The dynamic fingerprint to be identified is matched with the registered dynamic fingerprint of the local terminal for similarity, and if the similarity between the dynamic fingerprint to be identified and the registered dynamic fingerprint is greater than a threshold, it is determined that the local terminal has passed the verification.

6. The method according to claim 1, further comprising: Using a first model conversion method to perform model conversion on the large model to be quantized to generate a first converted large model; wherein the first model conversion method is the same as a second model conversion method used by the local terminal, and the second model conversion method is used to perform model conversion on the large model to be deployed; Using at least one large model quantization method to perform model quantization on the first conversion large models respectively, generating at least one set of first model quantization parameters, and determining preferred model quantization parameters from the first model quantization parameters; The preferred model quantization parameters are sent to the local terminal, so that the local terminal performs model deployment on the large model to be deployed according to the preferred model quantization parameters.

7. A large model deployment method, characterized in that: Executed by a local terminal, the method includes: Obtaining a target model quantization parameter sent by a cloud server; wherein the target model quantization parameter is obtained by the cloud server performing model quantization on a large model to be quantized in the cloud server, and the large model to be quantized has the same model structure as the large model to be deployed in the local terminal; Model deployment is performed on the large model to be deployed according to the target model quantization parameters.

8. The method according to claim 7, characterized in that The step of performing model deployment on the large model to be deployed according to the target model quantization parameters includes: Determine the target chip manufacturer to which the processor of the local terminal belongs, and determine the target model programming interface from the candidate model programming interfaces according to the manufacturer identification information of the target chip manufacturer; The target model programming interface is called to replace the current model parameters of the large model to be deployed with the target model quantization parameters, to generate an optimized large model, and to deploy the optimized large model in the local terminal.

9. The method according to claim 8, wherein deploying the optimized large model in the local terminal comprises: Parsing the current calculation graph corresponding to the optimized large model, and serializing the current calculation graph to generate a current sequence value; Performing hash calculation on the current sequence value to generate a current hash value, and determining whether the current reasoning logic of the optimized large model is the same as the benchmark reasoning logic based on a comparison result between the benchmark hash value and the current hash value; If the current reasoning logic is the same as the benchmark reasoning logic, the optimized large model is deployed in the local terminal.

10. The method according to claim 8, wherein deploying the optimized large model in the local terminal comprises: Determine a first model vocabulary corresponding to the optimized large model, and determine a second model vocabulary corresponding to the large model to be quantized; If the first model vocabulary is the same as the second model vocabulary, the optimized large model is deployed in the local terminal.

11. A large model deployment device, characterized in that: Configured on a cloud server, the device includes: A large model generation module to be quantified, used to generate a large model to be quantified in response to a user's structural configuration operation on the initial large model in the cloud server; wherein the large model to be quantified has the same model structure as the large model to be deployed in the local terminal; A model quantization module, configured to use at least one large model quantization method to perform model quantization on the large model to be quantized, generate at least one set of candidate model quantization parameters, and determine target model quantization parameters from the candidate model quantization parameters; The model quantization parameter sending module is used to send the target model quantization parameter to the local terminal, so that the local terminal performs model deployment on the large model to be deployed according to the target model quantization parameter.

12. A large model deployment device, characterized in that: Configured in a local terminal, the device includes: A model quantization parameter acquisition module, used to acquire a target model quantization parameter sent by a cloud server; wherein the target model quantization parameter is obtained by the cloud server performing model quantization on a large model to be quantized in the cloud server, and the large model to be quantized has the same model structure as the large model to be deployed in the local terminal; The model deployment module is used to deploy the large model to be deployed according to the target model quantization parameters.

13. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the large model deployment method described in any one of claims 1-6 and / or 7-10.

14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to execute the large model deployment method described in any one of claims 1-6 and / or 7-10.

15. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method for deploying a large model according to any one of claims 1-6 and / or 7-10.

Citation Information

Patent Citations

  • Quantization strategy determination method of neural network and image identification method and device

    CN110348562A

  • Model quantification method and device, electronic equipment and storage medium

    CN114970883A

  • Data processing method, system and device

    CN115766173A

  • Generative large model reasoning result verification method and verification system

    CN118964913A

  • Model optimization method, service processing method, server and terminal equipment

    CN119272822A