Model deployment method, electronic device and readable storage medium
By pruning and optimizing the parameter of the artificial intelligence model, the operating speed and efficiency problems caused by the huge model size are solved, and the model volume is reduced and performance improvement is achieved.
Patent Information
- Application Number
- CN202311872500.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2043-12-29
AI Technical Summary
As the data processing volume and complexity increase, the size of the artificial intelligence model becomes huge, resulting in a huge amount of model parameters, increasing the burden on electronic devices, and affecting the running speed and efficiency of the model.
By pruning the deployed model, removing removable parameters, obtaining a simplified model, reducing the model volume, and optimizing the simplified model through pre-training and fine-tuning samples to ensure the maintenance of model performance.
It realizes the reduction of the model volume, improves the running speed of the model, reduces power consumption and delay, and ensures the prediction capability and performance of the model. Especially in the case of large-volume models, the effect is significant.
Smart Images

Figure CN118444931B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of electronic devices, and in particular to a model deployment method, an electronic device, and a readable storage medium. Background Art
[0002] With the development of computer technology, artificial intelligence has emerged, and artificial intelligence models have also emerged and are constantly optimized and upgraded. Artificial intelligence models have the ability to learn sample data based on data, and predict events that have not occurred based on the data change patterns contained in the sample data. Artificial intelligence models can even perform complex and large-scale operations on the input data to obtain prediction results that are close to the data's change process in reality. However, as the amount of data processed or the complexity of the data processing process increases, the size of the artificial intelligence model also becomes huge, making the amount of model parameters involved in the operation of the artificial intelligence model huge, thereby increasing the burden on electronic devices that call the artificial intelligence model. Summary of the invention
[0003] This application provides a model deployment method, an electronic device, and a readable storage medium to reduce the size of the model and increase the model running speed. The technical solution is as follows:
[0004] In a first aspect, a model deployment method is provided, comprising: for at least one set of model parameters, removing each set of model parameters in the model to be deployed in turn to obtain at least one model to be tested; wherein each set of model parameters is used to perform model calculations on all elements in the matrix of input data, and the matrix of input data is completely transformed; the matrix of input data is generated when the input data is calculated inside the model to be deployed; the model to be deployed includes multiple sets of model parameters, and the multiple sets of model parameters include at least one set of model parameters; for each model to be tested, inputting test data into the model to be tested to obtain a test loss value; determining removable parameters in at least one set of model parameters based on the test loss values of all models to be tested; removing the removable parameters in the model to be deployed to obtain a simplified model of the model to be deployed; and deploying the model to be deployed based on the simplified model and the model to be deployed to obtain a deployed model.
[0005] By inputting test data into the test model after removing the model parameters of the model to be deployed, and then calculating the loss function of the test model on the output result of the test data to obtain the test loss value, the test loss value is compared with the loss value of the model to be tested, so as to test the prediction effect of the test model, determine whether the model parameters of the model to be deployed can be removed, and remove the removable parameters to obtain a simplified model, implement the pruning operation of the model to be deployed, and then simplify the model parameters of the model to be deployed. When applied to large-volume models, the data volume of the large-volume model in deployment is reduced.
[0006] In one embodiment, the model to be deployed is a large language model; for at least one set of model parameters, each set of model parameters is removed in turn in the model to be deployed to obtain at least one model to be tested, and the model deployment method also includes: using pre-training samples to pre-train a model framework of the large language model, so that the model framework predicts the N1+N2th word in each pre-training sample based on the first N1 words in each pre-training sample to obtain a pre-training model; for each pre-training sample, N1 is a variable greater than 0 and less than the total number of words in each training sample - N2, and N2 is an integer greater than 0; using fine-tuning samples to train the pre-training model, so that the pre-training model answers questions in the fine-tuning sample based on the text in the fine-tuning sample to obtain the model to be deployed.
[0007] At the same time, the embodiment of the present application simplifies the model to be deployed obtained by model training. In the process of simplification, it can accurately determine whether a specific set of model parameters are removable parameters. When the embodiment of the present application is applied to a large language model with more than 7 billion model parameters, while ensuring LLM performance, the optimization of power consumption and latency can achieve a positive return of about 14%+.
[0008] In one embodiment, removable parameters in at least one group of model parameters are determined based on the test loss values of all models to be tested, including: determining the importance scores of the removable parameters corresponding to the model to be tested based on the test loss value of each model to be tested and the reference loss value obtained by inputting test data into the model to be deployed; wherein the test data is data that is consistent with the deployment target of the model to be deployed; the importance score represents the difference between the model to be tested and the model to be deployed; the importance scores are sorted to obtain a sorting result; and the removable parameters are determined based on the sorting result.
[0009] The above-mentioned test data may be test data specifically set for the terminal where the model is deployed or the application using the model, or it may be data specifically extracted from the training data of the model to be tested, or it may be data obtained from the same type of terminal or application.
[0010] According to the test loss value of the model to be tested and the reference loss value of the original version of the model to be deployed, the model to be tested whose model prediction ability is less different from the original version of the model to be deployed is determined, and the model parameters of the determined model to be tested are used as removable parameters, thereby ensuring as much as possible that after each removable parameter is removed from the model to be deployed, the performance of the simplified model is close enough to the model to be deployed, reducing the decline in the prediction ability of the simplified model caused by removing the model parameters. The model deployed by the method provided in the embodiment of the present application still has the same level of processing power as the original model. For example, the large model deployed by the method provided in the embodiment of the present application still has the powerful data processing function of the large model and retains the characteristics of the large model.
[0011] In one embodiment, the importance scores of removable parameters corresponding to the model to be tested are determined based on the test loss value of each model to be tested and the reference loss value obtained by inputting test data into the model to be deployed, including: determining the functional relationship between the loss value and the input data based on the loss function and the model to be deployed; determining the test loss value and the importance score based on each group of model parameters in at least one group of model parameters, the partial derivatives of the functional relationship with respect to each group of model parameters, and the partial derivative matrix of a set scalar function with respect to each model parameter in each group of model parameters.
[0012] By determining the importance score of each set of model parameters through functional relationships, the importance of each set of model parameters can be judged more accurately when there is a large amount of test data.
[0013] In one embodiment, the model to be deployed also includes a vocabulary; deploying the model to be deployed according to the simplified model and the model to be deployed also includes: simplifying the vocabulary to obtain a simplified vocabulary; deploying the model to be deployed according to the simplified model, the simplified vocabulary and the model to be deployed.
[0014] In the embodiment of the present application, by simplifying the vocabulary, the storage space occupied by the vocabulary can be reduced and the event of the deployed model calling the vocabulary to process input data can be shortened.
[0015] In one embodiment, simplifying a vocabulary to obtain a simplified vocabulary includes: using the vocabulary to perform word segmentation operations on pre-trained samples and fine-tuning samples to obtain word segmentation of the pre-trained samples; determining the frequency of occurrence of each word in the vocabulary based on the word segmentation of the pre-trained samples and the fine-tuning samples; and deleting long-tail words in the vocabulary based on the frequency of occurrence to obtain a simplified vocabulary.
[0016] In the present embodiment, the pre-training samples used to simplify the vocabulary are samples for pre-training the model framework of the large language model. Through the pre-training samples, the model framework of the large language model can basically learn relatively complete language processing capabilities. Therefore, in the pre-training samples, the frequency of appearance of vocabulary words in the vocabulary can basically represent the frequency of each word appearing in each input data during the use of the model after deployment. Therefore, in the embodiment of the present application, determining whether to delete each word in the vocabulary based on the pre-training samples can not only reduce the data volume of the vocabulary, but also ensure the normal use of the model after deployment.
[0017] In one implementation, simplifying the word library to obtain a simplified word library includes: deleting non-Chinese and non-English words in the word library to obtain a simplified word library.
[0018] In most cases, people do not use non-Chinese and non-English words in their daily lives, such as Japanese and Korean. In most scenarios, people do not even use input methods that contain such symbols. Therefore, excluding non-Chinese and non-English words from the vocabulary can greatly simplify the vocabulary and significantly reduce its size, while also ensuring that the vocabulary meets the normal call requirements after the model is deployed.
[0019] In one embodiment, a simplified model is pre-trained using pre-training samples, so that the simplified model predicts the N1+N2th word in each pre-training sample based on the first N1 words in each pre-training sample, thereby obtaining a simplified model after post-training; for each pre-training sample, N1 is a variable greater than 0 and less than the total number of words in each training sample - N2, and N2 is an integer greater than 0; the pre-trained simplified model is trained using fine-tuning samples, so that the pre-trained simplified model answers questions in the fine-tuning samples based on the text in the fine-tuning samples, thereby obtaining a fine-tuned simplified model; the fine-tuned simplified model is deployed to obtain a deployed model.
[0020] Since the model to be deployed has been pruned, the simplified model obtained after pruning has some performance loss. Therefore, the simplified model is trained in the same way as the model to be deployed to improve the prediction ability of the simplified model.
[0021] In one embodiment, the model deployment method further includes: receiving a model request from a terminal; generating response data to the model request according to the deployed model; and sending the response data to the terminal.
[0022] In this embodiment, the deployed model is obtained according to the simplified model, and response data to the model request is generated according to the deployed model, so that the amount of data transmitted from the server to the terminal is small, the burden on the terminal is reduced, and the burden on data transmission is reduced.
[0023] In a second aspect, an electronic device is provided, the electronic device comprising: a processor and a memory;
[0024] The memory is used to store a program for the electronic device to execute the method provided in any embodiment of the present application, and to store data involved in implementing the method provided in any embodiment of the present application;
[0025] The processor is configured to execute the program stored in the memory.
[0026] Optionally, there are one or more processors and one or more memories.
[0027] Optionally, the memory may be integrated with the processor, or the memory may be provided separately from the processor.
[0028] The processing device in the above-mentioned second aspect can be a chip. The processor can be implemented by hardware or by software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc.; when implemented by software, the processor can be a general-purpose processor, which is implemented by reading the software code stored in the memory. The memory can be integrated in the processor or can be located outside the processor and exist independently.
[0029] In the specific implementation process, the memory can be a non-transitory memory, such as a read-only memory (ROM), which can be integrated with the processor on the same chip or can be set on different chips. This application does not limit the type of memory and the setting method of the memory and the processor.
[0030] In a third aspect, a computer-readable storage medium is provided, in which instructions are stored. When the instructions stored in the storage medium are executed on a computer, the computer can execute the method of the first aspect.
[0031] In a fourth aspect, the present application provides a computer program product. When the computer program product runs on a computer, it enables the computer to execute the schedule processing method of any possible implementation manner in the first aspect.
[0032] In a fifth aspect, an embodiment of the present application further provides a processor, comprising: an input circuit, an output circuit, and a processing circuit. The processing circuit is used to receive a signal through the input circuit and transmit a signal through the output circuit, so that the processor executes the method in any one of the embodiments of the first aspect above.
[0033] In the specific implementation process, the above-mentioned processor can be a chip, the input circuit can be an input pin, the output circuit can be an output pin, and the processing circuit can be a transistor, a gate circuit, a trigger, and various logic circuits. The input signal received by the input circuit can be, for example, but not limited to, received and input by a receiver, and the signal output by the output circuit can be, for example, but not limited to, output to a transmitter and transmitted by the transmitter, and the input circuit and the output circuit can be the same circuit, which is used as an input circuit and an output circuit at different times. This application does not limit the specific implementation of the processor and various circuits.
[0034] The technical effects obtained by the above-mentioned second, third, fourth and fifth aspects are similar to the technical effects obtained by the corresponding technical means in the above-mentioned first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 This is a schematic diagram of an application scenario of an embodiment of the present application;
[0036] Figure 2 is another application scenario schematic diagram of an embodiment of the present application;
[0037] Figure 3 This is another application scenario schematic diagram of an embodiment of the present application;
[0038] Figure 4 A schematic diagram of the model deployment method flow provided in an embodiment of the present application;
[0039] Figure 5 This is a flow chart of a model deployment method according to an example of this application;
[0040] Figure 6 This is a schematic diagram of the structure of a model deployment device in an example of this application;
[0041] Fig. 7A , Figure 7B , Figure 7C This is a schematic diagram of a sub-process of a model deployment method in an example of this application;
[0042] Figure 8 A schematic diagram of the software structure of an electronic device involved in an embodiment of the present application;
[0043] Fig. 9 A schematic diagram of the structure of an electronic device shown in an embodiment of the present application. DETAILED DESCRIPTION
[0044] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below in conjunction with the accompanying drawings.
[0045] It should be understood that the "multiple" mentioned in this application refers to two or more. In the description of this application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in order to facilitate the clear description of the technical solution of this application, the words "first" and "second" are used to distinguish between the same items or similar items with basically the same functions and effects. Those skilled in the art can understand that the words "first" and "second" do not limit the quantity and execution order, and the words "first" and "second" do not limit them to be different.
[0046] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0047] The artificial intelligence model (also referred to as the model in this application) can process the input data of the input model, and obtain the output data after complex calculations inside the model. With the development of computer technology, people can simulate the changes that may occur in data in specific scenarios in reality through artificial intelligence models, so as to predict according to the input data and obtain the corresponding output data as the prediction result. Artificial intelligence models with prediction functions include deep learning models and machine learning models. Among them, the deep learning model is a machine learning model based on artificial neural networks. The deep learning model may have multiple hidden layers, which can automatically learn complex features and patterns from the data and predict the data results that are currently unknown but may occur in the future. Deep learning models are usually trained with a large amount of data and computing resources to obtain better performance and accuracy. A machine learning model is a mathematical model. A machine learning model can also use data to learn the natural development and change laws of data and predict the data results that may be generated in the future. Machine learning models are usually trained using supervised learning or unsupervised learning algorithms, which can automatically learn development patterns and change laws from data.
[0048] In the model building stage, a large amount of sample data is usually needed to train the model so that the model can learn the rules in the sample data and change the model parameters inside the model according to the change rules from the input sample data to the output sample data in the sample data. Therefore, when receiving input data with unknown output results, the model can process the input data according to the learned model parameters and the change rules from the input sample data to the output sample data, predict the possible development results of the input data in reality, and obtain the output data.
[0049] Machine learning models and deep learning models can be used in various network applications, such as stock price prediction, air ticket price change prediction, passenger flow change prediction, product sales prediction; they can also repair and identify images, detect and screen diseases, predict various human indicators, or analyze and process natural languages. In reality, the law of data change is often very complex. Therefore, the artificial intelligence model used to predict based on input data and generate output data also has a complex internal model structure due to the complexity of the law of data change, and is also accompanied by complex model parameters. There are even cases where the number of internal parameters of artificial intelligence models is huge, which results in a large model.
[0050] Generally speaking, a large model refers to a neural network model that contains ultra-large-scale parameters (usually more than one billion). Large models are usually huge in scale and may contain billions of parameters (for example, more than 7 billion). The size of a large model can reach hundreds of gigabytes (GB) or even larger. This huge model scale provides it with powerful expression and learning capabilities, but it also consumes a lot of resources during the operation of the model. Even if some models are not large models in the general sense, there are still a large number of parameters (i.e., model parameters) inside the model, so that the client needs to obtain more data from the server that provides the model in order to use the model, or a large number of model parameters need to be used for calculation during the operation of the model, which affects the speed of generating output data.
[0051] Figure 1 Schematic diagram of an application scenario of the method for model deployment provided in an embodiment of the present application. Figure 1As shown, the user uses the dialogue model to perform text processing operations, and the user inputs dialogue initiation information through the terminal device 101. The dialogue initiation information may be information for initiating a topic when the user chats with the intelligent chat module, which may include keywords that the user wants to chat, or sentences about the chat topic, etc. The user enters the dialogue initiation information in the terminal device 101, and the dialogue initiation information may be a specific chat sentence, or an instruction to the model. In response to the dialogue initiation information input by the user, the terminal device 101 can use the dialogue processing model to process the dialogue initiation information. According to the characteristics of natural language, the dialogue processing model generates a response sentence to realize the chat between the terminal device 101 and the user. Before realizing this function, the terminal device 101 first obtains the dialogue processing model from the target server 102 (or other electronic devices that provide the model), and downloads the parameters inside the model and the function file used by the model to process data. In the case where the target server 102 updates the structure of the dialogue processing model, the terminal device 101 also needs to obtain the updated model file from the target server 102.
[0052] exist Figure 1 In the application scenario shown, the terminal device 101 may also be another server other than the target server 102. The target server 102 uses a large amount of sample data to train an untrained language processing model to obtain a deployable language processing model, and sends the model file of the trained language processing model to the model deployment server 103. Then the model file of the language processing model needs to be transmitted between the target server 102 and the model deployment server 103. In the case that some terminal devices 101 need to use the language processing model, they can send the processed conversation initiation information to the model deployment server 103 and receive the output data returned by the model deployment server 103.
[0053] Figure 2 Another application scenario schematic diagram of the method for model deployment provided in an embodiment of the present application. The target server 102 pre-trains an untrained model to obtain an image processing model. The user's terminal device 201 obtains the model file of the image processing model from the target server 102. After the terminal device 201 obtains multiple images from the image acquisition device 202, the image processing model can be called to perform target detection processing operations on the multiple images to obtain information related to the target objects in the multiple images, such as the number of target objects detected in the multiple images, the locations where the target objects appear, the color and / or shape and / or size of the target objects, and other attribute information, and the specific images where the target objects appear. After the user's terminal device 201 obtains the information related to the target object, it can summarize or filter the detection results of the multiple images, thereby tracking the target from the multiple images, reducing the time for manual image viewing, and improving the accuracy of image recognition.
[0054] Figure 2 In the application scenario shown, the image processing model can be a model obtained by training with a large number of sample images with annotation information. When training the model, the targets in a large number of sample images can be annotated by manual annotation, and the annotated sample images can be input into the image processing model to be trained to obtain the trained image processing model. After the trained image processing model is deployed on the terminal device 201, it can receive at least one image received or stored by the terminal device 201, and identify the target in the at least one image according to the target to be identified.
[0055] In another possible implementation scenario, the target server 102 may provide multiple different image processing models, each of which is used to identify a target, for example, vehicles, human bodies, trees, buildings, etc. may be identified through multiple different image processing models. Therefore, in order to implement the image recognition function, the terminal device may obtain model files of multiple image processing models corresponding to each function from the target server 102.
[0056] Figure 3 Schematic diagram of another application scenario of the method for model deployment in the embodiment of the present application. The target server 102 stores the model to be deployed. At the same time, the target server 102 is configured with multiple terminal devices 301. The terminal device 301 may be an intelligent terminal cabinet, an intelligent home appliance, etc., or may be a personal computer-type terminal. The terminal device 301 can obtain the model to be deployed from the target server 102, use the model to be deployed to process the user data that the terminal device 301 can obtain, and optimize the model to be deployed according to the feedback information in the user data, obtain the optimized model parameters, and then send the optimized model parameters to the target server 102. The target server 102 updates the model to be deployed according to the optimized model parameters. Each terminal device 301 can obtain the updated model to be deployed from the target server 102 according to the update notification of the model to be deployed, or obtain some parameters updated by the model to be deployed.
[0057] exist Figures 1 to 3 In the scenario shown, the terminal device can carry the deployed model file when it leaves the factory. The terminal device can also send a call instruction through the network to call the model deployed on the server side, send the input data to the server side, and the server side model obtains the output data based on the input data, and then returns the output data to the client. For example, Figure 1In the scenario shown, the user can use the network to perform text processing operations, and the user can input the dialogue initiation information through the terminal device 101. The dialogue initiation information may be the information for initiating a topic when the user chats with the intelligent chat module, which may include keywords that the user wants to chat, or sentences about the chat topic, etc. After receiving the dialogue initiation information, the terminal device 101 can process the dialogue initiation information, extract key information from the dialogue initiation information, or hide sensitive information from the dialogue initiation information, and then send the processed dialogue initiation information to the target server 102 where the dialogue processing model is deployed. After receiving the dialogue initiation information, the target server 102 calls the dialogue processing model to process the dialogue initiation information. The dialogue initiation information will be used to generate input information of the dialogue processing model, and the dialogue processing model performs natural language processing operations according to the input information to obtain output information. The output information is used as the basis for the response information to the dialogue initiation information, and the target server 102 sends the response information to the dialogue initiation information to the terminal device 101, thereby performing intelligent chat behavior with the user.
[0058] exist Figure 1 , Figure 2 and Figure 3 Based on the scenario shown, the specific implementation of the model deployment method provided in the embodiment of the present application is introduced below. Figure 4 The provided method can also be applied to electronic devices with other different layer structures. Figure 4 It is a flowchart of a model deployment method according to an exemplary embodiment. The method may include at least a part or all of the contents related to the following steps S41 to S45.
[0059] Step S41: for at least one set of model parameters, each set of model parameters is removed in turn in the model to be deployed to obtain at least one model to be tested; each set of model parameters is used to perform model calculation on all elements in the matrix of input data, and the matrix of input data is converted into one dimension; the matrix of input data is generated when the input data is calculated inside the model to be deployed; the model to be deployed includes multiple sets of model parameters, and the multiple sets of model parameters include at least one set of model parameters.
[0060] In an embodiment of the present application, the model to be deployed can process the input data, convert the input data into at least one-dimensional vectors, and then input the at least one-dimensional vectors into a subsequent data processing layer (or data processing module), so that the subsequent data processing layer performs vector operations on the at least one-dimensional vectors, and after multiple layers of vector operations, obtains a processed vector for calculating the output data, performs data conversion on the processed vectors, and obtains the output data. The model to be deployed may be a quantity prediction model, for example, it can predict the number of a certain type of fish in a river based on various factors. The model to be deployed may also be an image processing model, which performs binary classification of images based on images or image frames. In addition, the model to be deployed may also be an audio processing model, a facial feature processing model, and the like.
[0061] The at least one set of model parameters in step S41, which may also be referred to as model weights, may be model parameters that can be calculated with all data in at least one-dimensional vectors obtained by converting the input data. Converting the matrix of the input data by one dimension may refer to performing at least one round of operations on all elements in the matrix of the input data, each round converting all elements in the matrix of the input data into elements in another matrix.
[0062] Therefore, each set of model parameters is actually one or a set of vectors. In each layer of the model to be deployed, there may be at least one set of model parameters. For example, the model to be deployed includes an input layer, a hidden layer, and an output layer. In the input layer, the input character information is vectorized and converted into an input data row vector containing 100 elements. In the hidden layer, first, the input data row vector is matrix multiplied, and the input data column vector is converted into a first intermediate data matrix of 100×100 using a hidden layer column vector containing 100 elements. Then, each of the 100 elements of the hidden layer column vector can be used as a set of model parameters. Subsequently, for the intermediate data matrix, it is converted into a second intermediate data matrix of 100×3 through a hidden layer matrix of 100×3. Then, each column vector in the hidden layer matrix is a set of model parameters. Then, in the output layer, the 100×3 second intermediate data matrix is converted into a 3×3 output data matrix through a 3×20 output data conversion matrix, and finally the data is output according to the output data matrix. Then, the 20 column vectors in the second intermediate data matrix are 20 groups of model parameters.
[0063] It can be seen from the above content that when only vector multiplication (or matrix multiplication) is performed on the matrix of input data inside the model to be deployed, the number of parameters contained in each group of model parameters in at least one group of model parameters is the same as the number of elements in the column vector of the intermediate data matrix when the input data is calculated inside the model to be deployed. And each group of model parameters is a matrix column vector that performs matrix multiplication operation with the intermediate data matrix when the input data is calculated inside the model to be deployed. The aforementioned input data row vector, output data matrix, and intermediate data matrix can all be called the matrix of input data in step S41.
[0064] For example, when a matrix of input data is multiplied inside the model to be deployed, the input data is converted into an M×1 row vector by the model input layer after being input into the model to be deployed, where M is an integer greater than 1. Then, the matrix of the input data is a 1×M row vector containing M elements, and each group of model parameters in at least one group of model parameters includes M parameters. For another example, the input data is converted into an M1×1 row vector after being converted or undergoing at least one round of calculation in the model to be deployed, and then converted into an M2×1 column vector after at least one subsequent round of calculation. Then, each group of model parameters in at least one group of model parameters includes M1 or M2 parameters. For another example, the input data is converted into an M3×M4 intermediate data matrix after being converted or undergoing at least one round of calculation in the model to be deployed, so that each group of model parameters required for each round of calculation with the matrix of the input data is an M4×1 row vector. Wherein M, M1, M2, M3, and M4 are all integers greater than 1.
[0065] For another example, when mixed operations such as addition, subtraction, multiplication and division are performed on vectors within the model to be deployed, the dimension and / or number of elements of the matrix of the input data may continue to change as the calculation progresses, so that the number of parameters included in at least one set of model parameters changes with the change of the matrix of the input data.
[0066] The model to be deployed may include a multi-layer model structure, and each layer of the model structure may further include multiple sub-blocks or sub-layers, and each sub-block or sub-layer may include multiple groups of model parameters. The at least one group of model parameters in step S41 may be a part of the multiple groups of model parameters included in the sub-block or sub-layer.
[0067] In another possible implementation, if a certain layer within the model to be deployed performs a matrix addition operation on an intermediate data matrix, then the matrix added to the intermediate data matrix is a set of model parameters.
[0068] In a possible implementation, at least one set of model parameters may include all parameters in the model to be deployed, and each set of model parameters in the model to be deployed that can be used to operate the intermediate data matrix corresponding to the input data is removed tentatively to obtain multiple models to be tested. After each set of model parameters is removed, the parameters of the model to be deployed are incomplete parameters and become a model to be tested.
[0069] In a possible implementation, a group of model parameters can be removed by setting the group of model parameters in the model to be deployed to 0 or an invalid value, so that when the matrix of input data is calculated in the model to be deployed, the removed model parameters do not generate intermediate data corresponding to the input data. Alternatively, the group of model parameters can be deleted to reduce the dimension of the model parameter vector group where the reorganized model parameters are located.
[0070] In one possible implementation, at least one set of model parameters may be a column vector of a weight matrix in the model.
[0071] Step S42: For each model to be tested, input the test data into the model to be tested to obtain a test loss value.
[0072] The test data is equivalent to the input data of the model to be tested, and is used to test the prediction performance of the model to be tested. For the model to be tested, after the test data is input as input data, it is forward calculated inside the model, and after being processed by each layer inside the model to be tested, it is transmitted to the output layer of the model to be tested, and finally the output data is generated through the output layer of the model to be tested.
[0073] In the aforementioned forward calculation process, the test data is passed as input data to the model to be tested, and the output data is calculated according to the model architecture and model parameters of the model to be tested, which serves as the basis for evaluating the performance of the test model. When the model to be tested is a neural network model, forward calculation refers to a series of operations from the input layer to the output layer of the input data, where each operation is composed of the model's weights and activation functions. This process is also called forward propagation.
[0074] In the model to be deployed, when the model parameters include multiple groups, after each group of model parameters is removed, a corresponding model to be tested is obtained, and after the same test data is input into the model to be tested, the corresponding output data is obtained. The test data input for each different model to be tested is the same, and the test data can include multiple pieces, so as to test the prediction performance of the model to be tested from multiple different angles.
[0075] Step S42 may specifically include: for each model to be tested, inputting test data into the model to be tested to obtain output data of the model to be tested, and calculating a loss value for a single model to be tested based on the output data of the model to be tested and labeled data (or reference data), as the test loss value of the single model to be tested.
[0076] In a possible implementation, during the execution of steps S41 and S42, after each model to be tested is obtained, the test data may be used for testing to obtain a corresponding test loss value, and then the next model to be tested may be determined and the corresponding test loss value may be obtained, until each set of model parameters in the model to be deployed is deleted once, and the test loss values of all models to be tested are obtained. Alternatively, each set of model parameters of the model to be deployed may be deleted to obtain all the models to be tested, and then the test data may be input one by one for each model to be tested to obtain a corresponding test loss value. Alternatively, in the model to be deployed, the deleted model parameters may be randomly extracted, the model parameters may be deleted to obtain the model to be tested, the performance of the model to be tested may be tested, and if the set requirements are met, the randomly extracted model parameters may be used as removable parameters until the determined removable parameters meet the preset requirements for reducing the size of the model.
[0077] Step S43: Determine removable parameters in at least one group of model parameters according to the test loss values of all the models to be tested.
[0078] When determining removable parameters based on the test loss values, the test loss values can be sorted, and the model parameters corresponding to the smaller test loss values can be selected as removable parameters. Alternatively, an acceptable test loss value range can be determined, and the model parameters corresponding to the test loss values within the test loss value range can be determined as removable parameters.
[0079] In step S43, the test data is input into the model to be deployed with all model parameters, and forward calculation is performed inside the model to be deployed to obtain the output data of the model to be deployed. Then, based on the output data of the model to be deployed and the labeled data of the test data, the loss value of the model to be deployed is calculated as a reference loss value.
[0080] When determining removable parameters based on the test loss value, the reference loss value can be compared with the test loss value corresponding to each model to be tested to obtain the difference between each test loss value and the reference loss value. The model parameters that have no obvious impact on the loss value of the model to be tested and the original model to be deployed after removal are determined as removable parameters.
[0081] Step S44: remove removable parameters in the model to be deployed to obtain a simplified model of the model to be deployed.
[0082] In an embodiment of the present application, the removable parameters in the model to be deployed are removed, that is, all model parameters determined as removable parameters are removed, and all removable parameters are set to 0, or all removable parameters are set to invalid parameters, or some removable parameters are set to 0 and some are set to invalid parameters, so that the removable parameters do not participate in the calculation of the matrix of input data in the simplified model. In a possible implementation, the removable parameters of the model to be deployed do not include the parameters pre-deployed in the model to be deployed, but the parameters obtained by learning in the model training stage of the model to be deployed. The removable parameters are one of at least one set of model parameters in step S41.
[0083] After removing the removable parameters, the number of parameters of the simplified model is smaller than the number of parameters of the model to be deployed, and the volume of the simplified model is smaller than the volume of the model to be deployed, so that the speed at which the simplified model processes the input data is correspondingly improved.
[0084] Step S45: deploy the model to be deployed according to the simplified model and the model to be deployed to obtain a deployed model.
[0085] In this embodiment, the model to be deployed is deployed according to the simplified model and the model to be deployed, and the simplified model can be deployed as a simplified version of the model to be deployed, and the simplified model is finally used as the deployed model. Alternatively, the simplified model and the model to be deployed are deployed together on the server side, and when the client requests the simplified model, the model file of the simplified model is provided to the client, and when the client requests the model to be deployed, the model file of the model to be deployed is provided to the client.
[0086] Alternatively, based on the simplified model and the model to be deployed, the model to be deployed may be deployed by further adjusting the simplified model according to the model prediction effect of the model to be deployed to obtain an adjusted simplified model, and deploying the adjusted simplified model as the deployed model.
[0087] Alternatively, the model to be deployed is deployed based on the simplified model and the model to be deployed. The model to be deployed can be appropriately simplified based on the model prediction intersection of the model to be deployed and the simplified model to obtain another simplified model, and the other simplified model of the model to be deployed is deployed as the deployed model.
[0088] Alternatively, according to the simplified model and the model to be deployed, the model to be deployed may be deployed by setting the simplified model on the client or server. When the simplified model is set on the server, the client may obtain the model file of the simplified model from the server.
[0089] The present application inputs test data to the test model after removing the model parameters of the model to be deployed, and then the test model calculates the loss function on the output result of the test data to obtain a test loss value, and compares the test loss value with the loss value of the model to be tested, so as to test the prediction effect of the test model, determine whether the model parameters of the model to be deployed can be removed, and remove the removable parameters to obtain a simplified model, implement the pruning operation of the model to be deployed, and then simplify the model parameters of the model to be deployed and reduce the data volume of the model deployment.
[0090] In a possible implementation, the above Figure 4 Steps S41-S45 in the above can be executed by a dedicated model deployment end or a dedicated model training end. After completing the training of the model to be deployed, the same side executes the above steps S41-S45. In the case where the execution subject of steps S41-S45 is not a dedicated model training end, the dedicated model training end (model training server) sends the model to be deployed to the dedicated model deployment end after completing the training of the model to be deployed.
[0091] If the deployed model is newly upgraded and new model parameters are obtained, the upgraded new model can be reduced in size using the method of any embodiment of the present application when the model volume exceeds a preset threshold.
[0092] Since artificial intelligence technology has been developing rapidly in recent years, there are many different types and branches in the field of models. Before model training, the method provided in the embodiment of the present application can be used to reduce the size of the untrained model, but in the untrained stage, it is not possible to accurately judge the impact of different model parameters on the model prediction ability. Therefore, in one embodiment, the model to be deployed is a large language model (LLM); before step S41, the model deployment method also includes: using pre-trained samples, pre-training the model framework (bottom plate model) of the large language model, so that the model framework predicts the N1+N2th word in each pre-training sample according to the first N1 words in each pre-training sample, and obtains a pre-trained model; for each pre-training sample, N1 is a variable greater than 0 and less than the total number of words in each training sample - N2, and N2 is an integer greater than 0; using fine-tuning samples, the pre-training model is trained so that the pre-training model answers the questions in the fine-tuning sample according to the text in the fine-tuning sample, and obtains the model to be deployed.
[0093] A model framework (base model) usually refers to a basic model or a general model that can serve as the underlying framework or foundation for other models. Base models usually have some basic features and functions and can be used in different application scenarios and tasks. For example, the neural network model in deep learning can be regarded as a base model, which can be used for a variety of tasks such as image recognition and natural language processing.
[0094] In this embodiment, the pre-training samples may be samples used to pre-train the model framework, and the aforementioned characters may be Chinese language characters, English vocabulary, English abbreviations, characters that can be used for daily expressions (such as commonly used α, β, Ω), emoticons, etc. The aforementioned number of characters and total number of characters may be the number of characters in the Chinese language characters, the total number of characters in the Chinese language characters, the total number of English vocabulary, the total number of English abbreviations, or the total number of Chinese language characters, English vocabulary, and common characters.
[0095] In the case where the model to be deployed is a large language model, the large language model may include two stages: base model pre-training and supervised fine tuning (SFT) instruction fine tuning. The base model in this embodiment may be a large language model that has not undergone any training, and the parameters therein may include model parameters and model hyperparameters, wherein the model parameters are optimized as the training process progresses, and the model hyperparameters may be fixed parameters that do not change as the model is trained.
[0096] By performing unsupervised pre-training on the base model, a general base model (or open source model), i.e., a pre-trained model, can be obtained. This model can be applied to a variety of natural language processing scenarios.
[0097] The model to be deployed can be used to process language and text type input data and generate language and text type output results. Accordingly, the pre-training samples used to pre-train the model framework can be training samples of unlabeled data. The model framework can perform unsupervised training (or unsupervised learning) through the pre-training samples, and learn from the pre-training samples the organization of the language, language structure, vocabulary combination characteristics and other characteristics that constitute the pre-training samples, capture the general characteristics and structure of the input data, without paying attention to the labels of specific tasks, and without paying attention to the labeled answer information. Therefore, when receiving other similar language type input data, it can provide feedback on the input data according to the learned language organization, language structure and other characteristics.
[0098] In this embodiment, the pre-training sample can be a natural language paragraph, a chat sentence, an article, etc., and can also be text, an image or other forms of data on the Internet. For each (or each) pre-training sample, the model framework can predict the N1+1th (or N1+2, N1+3...)th word in the pre-training sample for the first N1 words in the input pre-training sample. N1 is a variable from 1 to the total number of words in the training sample - 1, that is, when N2 is 1, the model framework can predict the second word in the pre-training sample for the first word in the input pre-training sample, and predict the third word in the pre-training sample based on the first two words in the pre-training sample, and so on.
[0099] In this embodiment, the words in the pre-training sample may refer to characters, function symbols, punctuation marks, Chinese characters, English words, words in other languages, or single characters, numbers, etc. For natural language paragraphs, chat sentences, articles, etc. used as pre-training samples, they can be organized or split into single sentences with complete meanings, and the single sentences with complete meanings can be used as one (or one) pre-training sample. Alternatively, multiple paragraphs with clear logical order relationships can be used as one pre-training sample, so that the base model can learn the organizational relationship between paragraphs based on the pre-training samples.
[0100] In this embodiment, the fine-tuning sample can be a supervised fine-tuning sample containing annotation information, which is used to fine-tune the pre-trained model to adapt the pre-trained model to a specific task or field. Each (or each) fine-tuning sample can include reading content, question statements and annotated answers (i.e., annotation information). The fine-tuning sample can be generated based on the pre-training sample, or can be obtained by random sampling after mixing open source language samples with SFT fine-tuning data. The number of fine-tuning samples can be determined based on the computing power of the model training device, and can be increased as the computing power increases.
[0101] In the case where the model to be deployed is a model to be deployed for a target application, the fine-tuning samples are adapted to the purpose of use of the target application. For example, if the model to be deployed is expected to be applied to a translation model, then the fine-tuning samples are samples that tend to train the translation capabilities of the pre-trained model. The fine-tuning samples may include translation instructions, content to be translated, and annotated translation results. For another example, if the model to be deployed is expected to be applied to a dialogue model, then the fine-tuning samples are samples that tend to train the dialogue capabilities of the pre-trained model. The fine-tuning samples may include dialogue sentences, prompt words, and annotated dialogue results.
[0102] When the model to be deployed is a general model for the terminal, the fine-tuning sample is adapted to the operating environment of the terminal device. For example, the specific terminal device has a corresponding hardware environment and software environment, and the type of the fine-tuning sample belongs to the instructions that may be generated in the target user group of the terminal device, and the type of input data that may be generated in the hardware environment and software environment unique to the terminal device.
[0103] The pre-trained model is trained using fine-tuning samples, that is, the process of supervised learning on fine-tuning samples of one or more specific tasks. The tasks of the fine-tuning samples are usually tasks related to natural language processing (NLP) or other related processing methods, such as text classification, named entity recognition, machine translation, etc. In the embodiment of the present application, for each fine-tuning sample, a specific promt encoding instruction prompt is given to the pre-trained model, and then a target answer is given, so that the pre-trained model can learn the ability to answer questions through instructions. Prompt in the field of artificial intelligence generally refers to an input text or instruction used to guide the model to generate a specific output. In natural language processing and machine learning, prompts are often used to guide models to generate text related to a specific topic, context, or goal. For example, a user can provide a piece of text as input, and then use the model to generate output related to the text.
[0104] In natural language processing, prompts are often used for text generation tasks, such as dialogue systems, machine translation, and text summarization. By providing appropriate prompts, the model can generate high-quality output related to the input text. In machine learning, prompts can be used to guide the model to perform specific tasks, such as classification, regression, or clustering. By providing appropriate prompts, the model can better understand the input data and generate accurate output results.
[0105] The prompt code in the embodiment of the present application represents the prompt word code, which can be a programming method based on natural language processing, which guides the large language model to generate specific text output by providing prompt words to the large language model.
[0106] For example, in a fine-tuning sample, it includes user input data (user prompt), promt instruction prompt (system prompt) and target answer (result). The content of the system prompt may include multiple character information separated by symbols to indicate the direction of the target answer. As an example, the content of the system prompt may be: "The following is the content sent by the user to the voice assistant, which may contain several events. Each event contains <'title', 'start time', 'end time', 'recurrence', 'location', 'url link', 'participant'> fields. The fields in each event are output as a line in json format, and no additional response is made.\n\n".
[0107] The user prompt is the text content that needs to be understood by the pre-trained model. As an example, the content of the user prompt may be: "At 3 pm on Wednesday, I will go to Xinhua Bookstore to apply for a card and buy some books. I need a short period of time to choose. On Sunday, I will go to Shanghai General Motors to repair the main engine and transmission and be responsible for car washing and interior cleaning and maintenance. Well, this matter is very important. On Saturday, I will go to clean the house, including the refrigerator, TV, and computer. I may be busy then."
[0108] Result is the marked reference answer, which may include multiple reference answers for the same question. As an example, the content of result may be: "[{'Title':'Apply for a card and buy some books','Start time':'Wednesday 03:00pm','Location':'Xinhua Bookstore'}, {'Title':'Cleaning the house','Start time':'Saturday','Location':'Home'}, {'Title':'Car washing and interior cleaning and maintenance','Start time':'Sunday','Location':'Shanghai GM repairs the main engine and transmission'}]".
[0109] The system prompt, user prompt and result are all data of fine-tuning samples obtained by prompt encoding. Through the system prompt and user prompt, the pre-trained model can learn the content of the result, so that the pre-trained model can master the ability to extract information.
[0110] By fine-tuning the pre-trained model using fine-tuning samples, the weaker aspects of the pre-trained samples can be targeted to make up for them and expand the scope of application of the model. In the process of training the pre-trained model using fine-tuning samples, the pre-trained model can refer to the methods in the fine-tuning samples to learn how to answer question sentences, how to read text accurately, and how to give output data close to the labeled answer.
[0111] In another possible implementation, the above process of obtaining the model to be deployed can be performed on a dedicated model training end. When the model training end completes the training of the externally deployed model, the dedicated test end performs the above process on the model to be deployed. Figure 4 Steps S41-S45 shown in FIG. 1 , and Figure 4 The dedicated model training end may be a server end dedicated to providing a general model to be deployed.
[0112] When the embodiments of the present application are applicable to a large language model, the large language model is usually a model structure with model parameters of billions or tens of billions or more, and the large language model is composed of many layers of transformer decoder blocks with the same structure. Taking a large language model with more than seven billion model parameters as an example, the main structure configuration of this large language includes 32 layers of blocks, the hidden layer dimension of each layer is 4096, the number of attention heads is 32, the input vector dimension is the dimension of 120,000 words, and each word is represented by a vector of 4096 floating-point numbers; the total number of model parameters is more than seven billion. Such a large language model with a large number of parameters occupies about 15G of read-only memory (ROM) when using floating point (FP) 16 precision for storage. When using integer data (INT) 4 quantization for storage, it also needs to reach 3.42 gigabyte (G) model + 980 megabyte (MB) vocabulary = 4.38G. When the large language model is loaded on the mobile phone side for operation, the storage resource occupation of 4.38G of ROM is at least equivalent to the need to call more than 4.38G of random access memory (RAM) resources on the mobile phone side. Through actual measurement, when the model is called by the same version of software, when the model has a model parameter amount of the order of 7B, the end-to-end delay includes several parts of the following formula:
[0113] Total latency = model loading latency + total number of input tokens × prompt encoding latency + total number of generated tokens × single token inference latency. Token represents a character.
[0114] The above total delay, after end-to-end completion, ranges from a few seconds to tens of seconds, depending on the task. This is a considerable burden on the user experience and the power consumption of the mobile phone. Therefore, in an embodiment of the present application, the time is compressed by reducing the size of the deployed model.
[0115] By pre-training the model framework and then fine-tuning the pre-trained model in a targeted manner, a large language model with input data processing capabilities can be obtained, thereby indirectly ensuring the language processing function of the simplified model.
[0116] In other possible implementations, the model to be deployed may be an image processing model, a classification model, an audio processing model, a video processing model, or other types of models, or other types of large models. At the same time, the embodiment of the present application simplifies the model to be deployed obtained by model training, and in the process of simplification, it can accurately determine whether a specific set of model parameters are removable parameters. When the embodiment of the present application is applied to a large language model with more than 7 billion model parameters, while ensuring LLM performance, the optimization of power consumption and latency can achieve a positive return of about 14%+.
[0117] Within a model, each set of model parameters plays a certain role in accurately obtaining output data. When the model includes multiple sets of model parameters, different model parameters have different effects on model performance (that is, whether the model can accurately obtain output data). Therefore, if the model parameters that have a smaller impact on model performance are determined as removable parameters, it will not only reduce the size of the model but also retain the original performance of the model.
[0118] Therefore, in one embodiment of the present application, reference Fig. 7A , according to the test loss values of all the models to be tested, determining the removable parameters in at least one group of model parameters, also includes: step S71: determining the importance score of the removable parameters corresponding to the model to be tested according to the test loss value of each model to be tested and the reference loss value obtained by inputting the test data into the model to be deployed; wherein the importance score represents the difference between the model to be tested and the model to be deployed; step S72: sorting the importance scores to obtain the sorting result; step S73: determining the removable parameters according to the sorting result.
[0119] In this embodiment, the importance score can be calculated by a certain calculation formula for the reference loss value, and the importance of the model parameter corresponding to the reference loss value to the model performance can be evaluated based on the obtained value.
[0120] Since the model to be deployed is a trained model, the model to be deployed is used as a reference. After the test data is input into the model to be deployed, the output data of the model to be deployed is obtained, and the output data of the model to be deployed corresponds to a loss value. The loss value corresponding to the output data of the model to be deployed is used as the reference loss value, and the reference loss value is compared with the test loss value of each test model, so as to select the test model with the smallest difference with the original version of the model to be deployed after the model parameters are removed.
[0121] In another embodiment, other methods can be used to determine the difference between the test loss value and the reference loss value to represent the importance score. The smaller the difference, the lower the corresponding importance. According to the difference information, the models to be tested are sorted, and the models to be tested with the smaller difference between the test loss value and the reference loss value are selected. The model parameters corresponding to the selected models to be tested are used as removable parameters. The test data may include multiple, for example, the test data may include 10-50 data randomly selected from open source language sample data and SFT fine-tuning samples. The test loss values corresponding to the output data obtained from different test data may be different, and the reference loss values corresponding to different test data may also be different. Thus, a certain calculation method can be used to determine the difference between the test loss value and the reference loss value. For example, the L1 / L2 norm value, random forest algorithm or Taylor expansion can be used to determine the difference between the test loss value and the reference loss value corresponding to different test data, and the importance of each group of vectors in the LLM is evaluated. Alternatively, the difference between the test loss value and the reference loss value can be evaluated using the square difference between the test loss value and the reference loss value, the absolute value of the difference between the test loss value and the reference loss value, etc. The L1 norm value may be the sum of the absolute values of the elements in the vector, and the L2 norm value may be the square root of the sum of the squares of the elements in the vector.
[0122] Determining the removable parameters according to the sorting results may be to use the model parameters corresponding to the models to be tested that are sorted in a previously set proportional order as the removable parameters according to the sorting results.
[0123] In this embodiment, based on the test loss value of the model to be tested and the reference loss value of the original version of the model to be deployed, a model to be tested having a smaller difference in model prediction ability with the original version of the model to be deployed is determined, and the model parameters of the determined model to be tested are used as removable parameters, thereby ensuring as much as possible that after removing each removable parameter from the model to be deployed, the performance of the simplified model is sufficiently close to that of the model to be deployed, thereby reducing the decline in the prediction ability of the simplified model caused by removing the model parameters.
[0124] In another possible implementation, during the model optimization stage, model parameters that play a smaller role in model performance may be removed, and other model parameters may be added to replace the removed parameters. When other model parameters play a greater role in model performance than the removed parameters, other model parameters may be retained to improve the overall performance of the model before it is reduced in size, thereby indirectly improving the performance of the model after it is reduced in size.
[0125] Since there may be multiple test data used to test the model performance of the model to be tested, each test data corresponds to a different loss value to be tested, and if the performance of the model after the size reduction needs to be guaranteed as much as possible, it is necessary to determine the removable parameters as accurately as possible. The model parameters that have the least impact on the model performance are used as removable parameters, so it is necessary to accurately calculate and compare the multiple loss values to be tested corresponding to each group of model parameters and the reference loss value of the model to be deployed corresponding to each test data. Therefore, in one embodiment, the importance score of the removable parameter corresponding to the model to be tested is determined according to the test loss value of each model to be tested and the reference loss value obtained by inputting the test data into the model to be deployed, including: determining the functional relationship between the loss value and the input data according to the loss function and the model to be deployed; determining the test loss value and the importance score according to each group of model parameters in at least one group of model parameters, the partial derivatives of the functional relationship with respect to each group of model parameters, and the partial derivative matrix of the set scalar function with respect to each model parameter in each group of model parameters.
[0126] In an embodiment of the present application, the output data of the model to be deployed can be summarized as multiple functions about the input data. Inside the model to be deployed, a series of calculations are performed on the matrix of the input data using functions and model parameters. The functions inside the model to be deployed are summarized and expressed, and a function mapping from the input data to the output data can be obtained. At the same time, the relationship between the reference loss value of the model to be deployed and the output data can be expressed using a loss function. Therefore, the functional relationship between the reference loss value of the model to be deployed and the input data of the model to be deployed can be jointly determined by the function inside the model to be deployed and the loss function.
[0127] For example, when the model to be deployed is a fully connected neural network, the model to be deployed may include a neural network with a hidden layer. Assume that the input feature corresponding to the test data is x, the activation function of the hidden layer is h, and the activation function of the output layer is y. In the fully connected neural network, the weight matrix from the input layer to the hidden layer is Win, the weight from the hidden layer to the output layer is Wout, the bias coefficient of the input layer is bin, and the bias coefficient of the output layer is bout. Then, the forward calculation process from the input data to the output data includes the formula from the input layer to the hidden layer and the formula from the hidden layer to the output layer.
[0128] The formula from the input layer to the hidden layer may include: zin = Win × x + bin; h = activation (zin); activation represents the activation function, such as the sigmoid function, ReLu function, etc., and zin represents the matrix of input data. The formula from the hidden layer to the output layer may include: zout = Wout × h + bout; y = activation (zout), where zout represents the vector of output data. Through zin, h, zout, and y, the functional relationship from the input data to the output data of the model to be deployed can be determined.
[0129] In practical applications, the structure of the model to be deployed may include more complex formulas, and different model structures may require different forward calculation formulas.
[0130] At the same time, since the removal of model parameters of the model to be deployed will not affect the model layer structure or sub-layer functions within the model to be deployed, the model to be tested can also be summarized as a function from input data to output data. Therefore, the functional relationship from input data to output data of the model to be tested is the same as the functional relationship from input data to output data of the model to be deployed. The difference between the two lies in the parameters in the functional relationship. In the functional relationship of the model to be tested, a set of model parameters is missing.
[0131] In one embodiment, the importance score can be calculated by Taylor expansion to calculate the difference between the loss value to be tested and the reference loss value. Use the model to be tested and the model to be deployed to perform forward calculations on the same data to obtain loss-ori, and then remove each group of model parameters Wi to calculate the corresponding loss-i; then sort from large to small by loss-ori and the change Δloss-i of loss-i, and then according to the preset pruning rate P, each group of model parameters includes i model parameters, and the product of the total number of model parameters and the pruning rate P is p. The first i×p parameters are pruned to obtain the final pruned model LLM-pruned (the pruned model is the simplified model); assuming that a model with deployment includes 4096×4096 parameters and the pruning rate P is 0.2, then the model to be deployed becomes 4096×4096×0.8 parameters after pruning. Through Taylor expansion, the importance score can be calculated using the following formula:
[0132]
[0133] in,
[0134] Among them, W i is the model parameter removed from the model to be tested, D is the test data, is the loss function of the model to be deployed, is the loss function of the model to be tested, and the superscript T represents the transposed matrix. O(||W i || 3 ) is an infinitesimal quantity, and H is the Hessian matrix preset for the current model parameter Wi. The Hessian Matrix, also known as the Hessian Matrix, the Hessian Matrix, the Hessian Matrix, etc., is a square matrix composed of the second-order partial derivatives of a multivariate function, which describes the local curvature of the function. The Hessian matrix in the embodiment of the present application is a matrix containing the second-order partial derivatives of a scalar function, with a matrix size of n×n, where the element Hij represents the second-order partial derivatives of the preset scalar function f with respect to the i-th and j-th variables.
[0135] When the embodiment of the present application is applied to a large language model, the large language model needs to search for words from a certain vocabulary repository (i.e., a vocabulary library) to help the large language model generate output data of input data and perform operations related to natural language processing. In order to support the performance of the large language model, the vocabulary in the vocabulary library almost includes a large number of words in human language, and a large part of these words are basically not used in the model application process.
[0136] Therefore, in the embodiment of the present application, when the model to be deployed also includes a vocabulary, the following is adopted: Figure 7B The process shown simplifies the word library. Then, in the model deployment process, the model to be deployed is deployed according to the simplified model and the model to be deployed, and also includes: step S73: simplifying the word library to obtain a simplified word library; step S74: deploying the model to be deployed according to the simplified model, the simplified word library and the model to be deployed.
[0137] In the embodiment of the present application, when the model to be deployed is a large language model, the model to be deployed needs to use a vocabulary tool to generate output data or understand the input data when processing input data. The number of words in the vocabulary of the large language model is also very large. In the embodiment of the present application, in addition to simplifying the model parameters of the model to be deployed, the vocabulary of the model to be deployed is also simplified, and the vocabulary is trimmed to further reduce the size of the model after deployment. When simplifying the vocabulary, the vocabulary in the vocabulary can be simplified based on the public word frequency data.
[0138] In other possible implementations, when the model to be deployed is an image processing model, an audio and video processing model, a data computing model or other types of models, the model to be deployed may also have a corresponding vocabulary, and the vocabulary of other types of models to be deployed may also be simplified according to the method provided in the embodiments of the present application.
[0139] In a possible implementation, when users use large language models in daily life, they are unlikely to use some extremely professional words, such as highly professional compound name words. Therefore, highly professional words can be classified and provided to users when requested by users. By default, highly professional words are omitted from the model file provided to users to further reduce the model size and improve the model running speed.
[0140] In another possible implementation, the model architecture file of the model to be deployed, the model parameter file of the model to be deployed, and the vocabulary may be provided by different servers. Then, a model test end dedicated to simplifying the model can obtain the model to be deployed from the server end providing the model to be deployed, perform the simplification operation on the model to be deployed, and at the same time, obtain the vocabulary from the server end providing the vocabulary, and perform the simplification operation on the vocabulary.
[0141] Alternatively, if the model architecture file, model parameter file and vocabulary of the model to be deployed are provided by the same server, then a model testing end dedicated to simplifying the model can be used to obtain the model to be deployed and the vocabulary from the server that provides the model to be deployed and the vocabulary, and perform simplification operations on the model to be deployed and the vocabulary, respectively.
[0142] Alternatively, the simplification operation of the model to be deployed and the simplification operation of the vocabulary are both performed by different servers. When deploying the model, the server that deploys the model obtains the deployable model from another server that provides a simplified version of the deployable model, and obtains the simplified vocabulary from another server that provides a simplified vocabulary.
[0143] There may be many ways to simplify the vocabulary, but in order to ensure the performance of the model after the size is reduced, it is necessary to selectively delete or trim the vocabulary in the vocabulary to avoid not being able to obtain sufficient vocabulary support during the operation of the model. In one embodiment, simplifying the vocabulary to obtain a simplified vocabulary includes: using the vocabulary to perform a segmentation operation on a pre-training sample to obtain the segmentation of the pre-training sample; determining the frequency of occurrence of each word in the vocabulary based on the segmentation of the pre-training sample; and deleting long-tail words in the vocabulary based on the frequency of occurrence to obtain a simplified vocabulary.
[0144] In the embodiment of the present application, the vocabulary includes multiple words (i.e., words in the embodiment of the present application). The original vocabulary in the vocabulary can be used to perform byte pair encoding (BPE) segmentation on the pre-trained samples and fine-tuning samples, and then perform vocabulary screening by combining word frequency information with preset rules. BPE is a commonly used unsupervised text segmentation algorithm, which is usually used to deal with vocabulary segmentation problems in natural language processing tasks. The main idea of BPE is to build a vocabulary by iteratively merging the most frequently adjacent character pairs (byte pairs). A simple example is that, according to the target data set, adjacent characters that appear frequently are combined together to form a word.
[0145] After determining the frequency of occurrence (word frequency) of all words in the vocabulary, the vocabulary can be simplified from at least two aspects: deleting all non-Chinese and English character strings and symbols; and deleting long-tail words with a word frequency < 5 (or other set values, such as 1-10) by word frequency sorting.
[0146] In actual applications, through the above two aspects of screening, the vocabulary of the large language model is reduced from 12W to 9.4W, and the linear layer of the final character prediction is also reduced from 4096×12w to 4096×9.4w. Since vocabulary pruning only narrows the scope of the vocabulary in the vocabulary, and only targets uncommon words and long-tail words, it will not change the values of more than 7 billion parameters of the large language model, so it will not have a visible impact on the performance of the large language model. The pre-training samples that determine the vocabulary frequency are samples for pre-training the model framework of the large language model. Through the pre-training samples, the model framework of the large language model can basically learn a relatively complete language processing capability, so that in the pre-training samples, the frequency of appearance of vocabulary in the vocabulary can basically represent the frequency of each vocabulary in each input data during the use process after the model is deployed. Therefore, in the embodiment of the present application, determining whether to delete each vocabulary in the vocabulary based on the pre-training samples can not only reduce the data volume of the vocabulary, but also ensure the normal use of the model after deployment.
[0147] In the embodiment of the present application, the long-tail words are words distributed at relatively flat positions in the vocabulary frequency curve. The flat position and non-flat position of the curve can be distinguished by presetting a frequency threshold.
[0148] When the method provided in the embodiment of the present application is applicable to a large language model, in daily life, the vocabulary that users of each country may involve in the process of using the large language is in a relatively concentrated range. When the model is applied to a Chinese user group, in one embodiment of the present application, the word library is simplified to obtain a simplified word library, including: deleting non-Chinese and non-English words in the word library to obtain a simplified word library.
[0149] In most cases, Chinese users do not use non-Chinese and non-English words in their daily lives, such as Japanese and Korean. In most scenarios, people do not even use input methods that contain such symbols. Therefore, excluding non-Chinese and non-English words from the vocabulary can greatly simplify the vocabulary and significantly reduce its size, while also ensuring that the vocabulary meets the normal call requirements after the model is deployed.
[0150] Although the embodiments of the present application try to ensure that when the model is reduced in size, the removable model parameters are the parameters that have a smaller impact on the model performance, there will always be a certain performance loss when the model is reduced in size.
[0151] Therefore, in one embodiment of the present application, still refer to Figure 7C , deploying the model to be deployed according to the simplified model and the model to be deployed, also includes: step S75: using the pre-training samples to pre-train the simplified model, so that the simplified model predicts the N1+N2th word in each pre-training sample according to the first N1 words in each pre-training sample, and obtains the simplified model after post-training; for each pre-training sample, N1 is a variable greater than 0 and less than the total number of words in each training sample - N2, and N2 is an integer greater than 0; step S76: using the fine-tuning samples to train the pre-trained simplified model, so that the pre-trained simplified model answers the questions in the fine-tuning samples according to the text in the fine-tuning samples, and obtains the fine-tuned simplified model; step S77: deploying the fine-tuned simplified model to obtain the deployed model.
[0152] Since the model to be deployed has been pruned, the obtained simplified model will have a certain loss in performance relative to the model to be deployed. Therefore, in this embodiment, the simplified model is post-trained again with pre-training data to restore the performance of the simplified model to a certain extent, thereby obtaining a simplified model after post-training. Post-training here refers to performing some pre-training again as recovery training for the simplified model, which is equivalent to letting the simplified model learn some content again after completing the pruning of the pre-training model, resulting in some loss in the performance of the pre-training model, so as to achieve the purpose of restoring the model performance. During the post-training process, a portion of the pre-training data can be resampled and then pre-trained.
[0153] In a possible implementation, the operation of training the model and the operation of simplifying the model may be performed by two different servers. Then, after generating the model to be deployed, the server that simplifies the model obtains the model to be deployed from the server that trains the model. After the server that simplifies the model generates the simplified model, it sends the simplified model to the server that trains the model, and the server that trains the model performs post-training and fine-tuning operations on the simplified model.
[0154] In a possible implementation, the operation of training the model and the operation of simplifying the model may be performed by two different servers. Then, after generating the model to be deployed, the server that simplifies the model obtains the model to be deployed from the server that trains the model. After the server that simplifies the model generates the simplified model, the simplified model can be post-trained and fine-tuned directly on the server that simplifies the model.
[0155] In step S710, the (simplified) vocabulary is deployed together with the fine-tuned model.
[0156] When restorative training is performed on the simplified model, the pre-trained samples and fine-tuned samples used can be different from those used when training the base model. After restorative training is performed on the simplified model, the SFT instruction fine-tuning training with full parameters is performed on the simplified model obtained by restorative training, and the model capabilities required by business applications, such as information extraction and dialogue capabilities, are injected, and finally a model that can be directly deployed is obtained.
[0157] In this embodiment, the pre-training samples include multiple ones, and the original data of the pre-training samples can be various chat sentences, articles, paragraphs, and the pre-training samples can be a sample formed by splitting and combing natural language materials such as chat sentences, articles and paragraphs. Since the model to be deployed is pruned, the simplified model obtained after pruning has some performance loss. Therefore, the simplified model is trained in the way of training the model to be deployed to improve the prediction ability of the simplified model.
[0158] Generally, the terminal may carry the model file when leaving the factory, or request to download the model file from the server when installing the application, or request the model file from the server while the application is running, or request a new version of the model file from the server after the server provides the new version of the model file.
[0159] Therefore, in one embodiment, the model deployment method further includes: receiving a model request from a terminal; generating response data of the model request according to the deployed model; and sending the response data to the terminal.
[0160] The model request may be a request to download the model file of the deployed model to the terminal so as to run the model file on the terminal. Alternatively, it may be a request to use the model deployed on the server side to obtain the model output data. In the case where the terminal carries the model file when it leaves the factory, the model request may be sent by the storage module for the terminal. According to the model request, the simplified large model is deployed in the chip system to be used to assemble the terminal.
[0161] In this embodiment, the deployed model is obtained according to the simplified model, and response data to the model request is generated according to the deployed model, so that the amount of data transmitted from the server to the terminal is less, the burden on the terminal is reduced, and the burden on data transmission is reduced.
[0162] In a specific example, Figure 5 As shown, in the stage without any training, the model in the example of the present application is a base model 51. After the base model 51 is trained, the model to be deployed 52 is obtained. Among them, the training process of the base model 51 includes a pre-training process of the base model 51 using pre-training samples, and a process of fine-tuning the base model 51 using fine-tuning samples. For the model to be deployed, the importance of the model parameters is evaluated using test data to obtain the importance score of each group of model parameters. The test data may include 10-30 data randomly sampled from a mixture of SFT samples and open source data. For each group of model parameters that can be removed, the model parameters are removed to obtain a test model, the test data is input into the test model to obtain output data, and the importance score corresponding to each group of model parameters is calculated according to the output data of the test model. According to the importance score corresponding to the model parameters, at least one group of model parameters (i.e., model weights) is deleted to obtain a simplified model. The simplified model is post-trained using the pre-trained samples 53 of the training base model 51 to obtain a simplified model after post-training, and then the SFT samples 54 are used to fine-tune the simplified model after post-training to obtain a simplified model after fine-tuning 56. The pre-trained samples 53 may use public data, which may include 200,000 to 160,000 samples. For the vocabulary corresponding to the model, the vocabulary in the vocabulary is trimmed by performing BPE segmentation on the data of the SFT samples and the open source samples, and using the segmentation to determine the frequency of the vocabulary in the vocabulary to obtain a trimmed (i.e. simplified) vocabulary. The trimmed vocabulary and the fine-tuned simplified model 56 are used as the final deployable model 55. Among them, the open source samples may include samples disclosed through the network platform.
[0163] Corresponding to the model deployment method provided in the embodiment of the present application, the embodiment of the present application also provides a model training device for implementing the model deployment method provided in any embodiment of the present application. In one embodiment, the structure of the model training device is as follows: Figure 5As shown, it includes: a model module to be tested 61, a test loss value module 62, a removable parameter module 63, a simplified model module 64 and a deployment module 65.
[0164] Among them, the model to be tested module 61 is used to remove each group of model parameters in the model to be deployed in turn for at least one group of model parameters to obtain at least one model to be tested; wherein each group of model parameters is used to perform model calculations on all elements in the matrix of input data, and the matrix of input data is completely converted; the matrix of input data is generated when the input data is calculated inside the model to be deployed; the model to be deployed includes multiple groups of model parameters, and the multiple groups of model parameters include at least one group of model parameters.
[0165] The test loss value module 62 is used to input the test data into each model to be tested to obtain the test loss value.
[0166] The removable parameter module 63 is used to determine removable parameters in at least one group of model parameters according to the test loss values of all the models to be tested.
[0167] The simplified model module 64 is used to remove removable parameters in the model to be deployed to obtain a simplified model of the model to be deployed.
[0168] The deployment module 65 is used to deploy the model to be deployed according to the simplified model and the model to be deployed to obtain the deployed model.
[0169] Due to the huge number of parameters of large language models, which usually reach more than billions of parameters, if you want to deploy them on terminal devices, they require a very large ROM / RAM occupancy, and the inference latency and power consumption are also huge costs. Therefore, this patent proposes a model compression solution, which mainly prunes the weights of the base model, performs post-training recovery training, and then fine-tunes the SFT instructions on the business, and finally performs vocabulary pruning. Finally, a smaller model with performance that is not inferior to the original model can be output. When it is deployed on terminal devices, it can achieve ROM, RAM, latency, and power consumption benefits proportional to the weight pruning rate.
[0170] It should be noted that: when recording information, the model deployment device provided in the above embodiment only uses the division of the above-mentioned functional modules as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0171] For example, you can remove the parameter module execution Fig. 7A The deployment module also performs the steps Figure 7B , Figure 7CThe model deployment device also has a functional module for training the baseboard model to obtain the model to be deployed.
[0172] In the embodiment of the present application, the user can use an electronic device as a user terminal to call the deployed model. Figure 8 is a schematic diagram of a software architecture of an electronic device according to an exemplary embodiment. Figure 8 The electronic device shown runs the Android system.
[0173] Figure 8 is a block diagram of a software system of an electronic device to which the method provided in the embodiment of the present application is applied. Figure 8 , the layered architecture divides the software into several layers, each with clear roles and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system includes multiple layers, from top to bottom, the application layer, the application framework layer, the Android runtime (Android runtime) and the system layer, and the kernel layer. When the instruction of the model call is issued, the instruction is reported from the bottom layer to the application framework layer or the application layer layer by layer.
[0174] The application layer can include a series of application packages. Figure 8 As shown, the application package may include desktop applications, cameras, gallery, calls, maps, navigation, Bluetooth, music, videos and other applications. These applications can be set in the process of program running, according to the network data received by the application or according to the user's instructions, call the model set in the terminal to obtain the output data of the model.
[0175] The application framework layer provides application programming interface (API) and programming framework for the applications in the application layer. The application framework layer includes some predefined functions. Figure 4 As shown, the application framework layer may include a window manager, a content provider, an interface layer, a phone manager, a resource manager, a notification manager, etc. When a user issues a call model instruction through the touch screen of an electronic device, information about the user's call model operation may be transmitted to the application layer through the application framework layer.
[0176] The system layer includes the Android Runtime. The Android Runtime includes the core library and the virtual machine. The Android Runtime is responsible for the scheduling and management of the Android system.
[0177] The kernel layer is the layer between hardware and software. The kernel layer includes at least display driver, camera driver, audio driver, and sensor driver. When a user operates the touch screen or buttons of an electronic device, the kernel layer can generate corresponding information, such as model call information, based on the information of the user's operation on the touch screen or buttons.
[0178] Under the above four-layer architecture, the receiving device is also provided with a hardware layer, which may include the aforementioned electronic device hardware components, such as display screens, buttons, etc.
[0179] At the same time, the model deployment device of the embodiment of the present application can be set at Figure 8 The electronic device shown may also include other functional modules to implement any steps and functions in the method for processing click operation information provided in any embodiment of the present application.
[0180] The functional units and modules in the above embodiments may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit, and the above integrated units may be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the protection scope of the embodiments of the present application.
[0181] The device for determining location information provided in the above embodiments and the method embodiment for determining location information belong to the same concept. The specific working process of the units and modules in the above embodiments and the technical effects brought about can be found in the method embodiment part and will not be repeated here.
[0182] As an example of the present application, the electronic device can access the base station and also has the ability to access the wireless local area network, for example, the electronic device is a mobile phone, a tablet, a smart watch, a portable notebook, etc. Please refer to Fig. 9 , Fig. 91 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. Among them, the sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0183] It is to be understood that the structure illustrated in the embodiment of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown in the figure, or combine some components, or split some components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0184] The processor 110 may include one or more processing units, for example: the processor 110 may include an application processor (application processor, AP), a modem processor, a graphics processor (graphics processing unit, GPU), an image signal processor (image signal processor, ISP), a controller, a memory, a video codec, a digital signal processor (digital signal processor, DSP), and / or a neural-network processing unit (neural-network processing unit, NPU), etc. Among them, different processing units can be independent devices or integrated in one or more processors. In an embodiment of the present application, the processor is used to generate instructions recognizable by the application based on the signal received by the sensor, and the application can call the model deployed on the electronic device 100 according to the recognizable instructions.
[0185] The controller may be the nerve center and command center of the electronic device 100. The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions. For example, the controller may generate a control signal for shutting down the electronic device 100 and a control signal for restarting the electronic device 100 within a short time after shutting down the electronic device 100 according to the instruction operation code.
[0186] The processor 110 may also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory may store data of a model deployed on the electronic device 100, including a model framework and model parameters, and a resource library used by the model, such as a vocabulary library used by a large language model.
[0187] In some embodiments, the processor 110 may include one or more interfaces, such as an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface.
[0188] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.
[0189] Among them, the wireless communication module 160 can provide wireless communication solutions including wireless local area networks (WLAN) (WiFi network), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication technology (NFC), infrared technology (IR), etc. for application in the electronic device 100.
[0190] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, so that electronic device 100 can communicate with the network and other devices through wireless communication technology and call models deployed on other devices.
[0191] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. In some embodiments, the electronic device 100 may include 1 or N display screens 194, where N is an integer greater than 1. The output data of the model of the electronic device 100 may be displayed through the display screen 194.
[0192] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement a data storage function, such as persistent storage of different versions of models.
[0193] The internal memory 121 can be used to store computer executable program codes, which include instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area may store data (such as audio data, a phone book, etc.) created by the electronic device 100 during use. In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (universal flash storage, UFS), etc. The internal memory may also be used to store model information. When the model deployed in the embodiment of the present application is called, the internal memory may store a matrix of input data generated when the model is called.
[0194] The touch sensor 180K is also called a "touch panel". The touch sensor 180K can be set on the display screen 194, and the touch sensor 180K and the display screen 194 form a touch screen, also called a "touch screen". The touch sensor 180K is used to detect touch operations acting on or near it. The touch sensor 180K can pass the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 194. In other embodiments, the touch sensor 180K can also be set on the surface of the electronic device 100, which is different from the position of the display screen 194. The touch sensor 180K can convert the analog signal of the user input instruction into a digital signal of the calling model when receiving the user's input instruction.
[0195] The button 190 includes a power button, a volume button, etc. The button 190 can be a mechanical button or a touch button. The electronic device 100 can receive the button input and generate a key signal input related to the user settings and function control of the electronic device 100. Through the button 190, a model call instruction can also be generated. For example, a special gesture for pressing the button 190 can be set. When the user presses the button 190 with a special gesture, the model deployed on the electronic device 100 is called.
[0196] Indicator 192 may be an indicator light, which may be used to indicate charging status, power changes, messages, missed calls, notifications, etc.
[0197] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When loading and executing computer instructions on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from a website site, computer, server or data center by wired (such as: coaxial cable, optical fiber, data subscriber line (Digital Subscriber Line, DSL)) or wireless (such as: infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server, data center, etc. that contains one or more available media integrated. The available media may be magnetic media (such as floppy disks, hard disks, and magnetic tapes), optical media (such as digital versatile discs (DVD)), or semiconductor media (such as solid state disks (SSD)).
[0198] The above are optional embodiments provided for the present application and are not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the technical scope disclosed in the present application shall be included in the protection scope of the present application.
Claims
1. A model deployment method, It is characterized in that include: For at least one set of model parameters, each set of the model parameters is removed in turn in the model to be deployed to obtain at least one model to be tested; wherein each set of the model parameters is used to perform model calculation on all elements in the matrix of input data, and the matrix of the input data is completely converted; the matrix of the input data is generated when the input data is calculated inside the model to be deployed; the model to be deployed includes multiple sets of model parameters, and the multiple sets of model parameters include the at least one set of model parameters; each set of model parameters is a matrix vector that performs matrix multiplication operation with an intermediate data matrix when the input data is calculated inside the model to be deployed; For each of the models to be tested, input the test data into the model to be tested to obtain a test loss value; Determine, according to the test loss values of all the models to be tested, a removable parameter in the at least one set of model parameters; the removable parameter includes one of the at least one set of model parameters; Removing removable parameters in the model to be deployed to obtain a simplified model of the model to be deployed; Deploy the model to be deployed according to the simplified model and the model to be deployed to obtain a deployed model; The step of determining the removable parameters in the at least one set of model parameters according to the test loss values of all the models to be tested comprises: Determine the importance score of the removable parameters corresponding to the model to be tested according to the test loss value of each of the models to be tested and the reference loss value obtained by inputting the test data into the model to be deployed; wherein the test data is adapted to the deployment target of the model to be deployed; and the importance score represents the difference between the model to be tested and the model to be deployed; Sorting the importance scores to obtain a sorting result; The removable parameters are determined according to the sorting result.
2. The method according to claim 1, It is characterized in that The model to be deployed is a large language model; for at least one set of model parameters, each set of model parameters is removed in turn from the model to be deployed, before obtaining at least one model to be tested, the method further includes: Using the pre-training samples, pre-training the model framework of the large language model, so that the model framework predicts the N1+N2th word in each pre-training sample according to the first N1 words in each pre-training sample to obtain a pre-training model; for each pre-training sample, N1 is a variable greater than 0 and less than the total number of words in each pre-training sample minus N2, and N2 is an integer greater than 0; The pre-trained model is trained using the fine-tuning sample, so that the pre-trained model can answer questions in the fine-tuning sample according to the text in the fine-tuning sample, thereby obtaining the model to be deployed.
3. The method according to claim 1, It is characterized in that The step of determining the importance score of the removable parameters corresponding to the model to be tested according to the test loss value of each model to be tested and the reference loss value obtained by inputting the test data into the model to be deployed comprises: Determine a functional relationship between a loss value and input data according to the loss function of the model to be deployed and the model to be deployed; The test loss value and the importance score are determined according to each group of model parameters in the at least one group of model parameters, the partial derivatives of the functional relationship with respect to each group of model parameters, and the partial derivative matrix of a set scalar function with respect to each model parameter in each group of model parameters.
4. The method according to claim 2, It is characterized in that The model to be deployed also includes a vocabulary; and deploying the model to be deployed according to the simplified model and the model to be deployed further includes: Simplifying the vocabulary to obtain a simplified vocabulary; The model to be deployed is deployed according to the simplified model, the simplified vocabulary and the model to be deployed.
5. The method according to claim 4, It is characterized in that The simplifying of the vocabulary to obtain a simplified vocabulary includes: Using the vocabulary, performing a word segmentation operation on the pre-training sample and the fine-tuning sample to obtain word segmentations of the pre-training sample; Determine the frequency of occurrence of each word in the vocabulary according to the word segmentation of the pre-training sample and the fine-tuning sample; According to the occurrence frequency, long-tail words in the vocabulary are deleted to obtain a simplified vocabulary.
6. The method according to claim 4, It is characterized in that The simplifying of the vocabulary to obtain a simplified vocabulary includes: The non-Chinese and non-English words in the word library are deleted to obtain the simplified word library.
7. The method according to any one of claims 1 to 6, It is characterized in that The step of deploying the model to be deployed according to the simplified model and the model to be deployed includes: Using the pre-training samples, the simplified model is pre-trained, so that the simplified model predicts the N1+N2th word in each pre-training sample according to the first N1 words in each pre-training sample, and obtains the simplified model after post-training; for each pre-training sample, N1 is a variable greater than 0 and less than the total number of words in each pre-training sample minus N2, and N2 is an integer greater than 0; Using the fine-tuning sample, training the pre-trained simplified model, so that the pre-trained simplified model can answer questions in the fine-tuning sample according to the text in the fine-tuning sample, thereby obtaining a fine-tuned simplified model; The fine-tuned simplified model is deployed to obtain the deployed model.
8. The method according to any one of claims 1 to 7, It is characterized in that The method further comprises: Receive a model request from a terminal; Generate response data for the model request according to the deployed model; The response data is sent to the terminal.
9. An electronic device, It is characterized in that The electronic device comprises: a processor and a memory; The memory is used to store a program for the electronic device to execute the method according to any one of claims 1 to 8, and to store data involved in implementing the method according to any one of claims 1 to 8; The processor is configured to execute the program stored in the memory.
10. A chip system, It is characterized in that The chip system is applied to an electronic device, and the chip system includes one or more processors, and the processor is used to call computer instructions so that the electronic device executes the method as described in any one of claims 1 to 8.
11. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed on an electronic device, the electronic device executes the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Network model compression method and device, storage medium and computer equipment
CN110490323A
Model deployment method, device, electronic equipment and storage medium
CN110991643A