Question and answer method and product of question and answer large language model based on data-driven regularization
By pruning the Q&A large language model on lightweight devices, we migrate important information to the rest, solving the problem of insufficient hardware configuration, achieving efficient Q&A performance and low resource consumption, and improving user experience.
Patent Information
- Application Number
- CN202510828333.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-20
AI Technical Summary
When deploying a large Q&A language model on lightweight devices, there are problems such as insufficient hardware configuration of the device, high memory usage, large computing resources consumption, and too long response time for problem reasoning, which affects the user experience.
By obtaining the device parameter value of the target device, determining the channel that needs to be pruned, using the loss function with regular loss terms to update the pre-trained model parameters, and performing channel pruning, migrating important information to the remaining part, and obtaining the target Q&A large language model.
While reducing the scale of the model, maintain model performance, reduce hardware configuration and computing resource overhead, improve problem reasoning and response efficiency, and improve user experience.
Smart Images

Figure CN120336497A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a question-answering method and product for a question-answering large language model based on data-driven regularization. Background Art
[0002] With the rapid development of artificial intelligence technology, question-answering large language model (LLMs) products are widely used in multiple fields such as natural language processing, intelligent customer service, medical health, and education and training. Due to the popularization of artificial intelligence products, question-answering large language model products are deployed in lightweight devices (such as mobile phones, tablets, etc.) to meet the usage needs of users. However, when users use question-answering large language model products on lightweight devices for question reasoning, there are problems such as insufficient device hardware configuration, high memory occupancy rate, large consumption of computing resources, and too long question reasoning response time, which seriously affect the user experience.
[0003] Based on this, it is necessary to find a method that can improve the question-answering efficiency of question-answering large language model products on lightweight devices, reduce resource overhead, and improve the user experience. Summary of the Invention
[0004] In view of this, the present application aims to propose a question-answering method and product for a question-answering large language model based on data-driven regularization, so as to improve the question-answering efficiency of question-answering large language model products on lightweight devices, reduce resource overhead, and improve the user experience.
[0005] To achieve the above object, the technical solution of the present application is as follows: The first aspect of the embodiments of the present application provides a question-answering method for a question-answering large language model based on data-driven regularization, and the method includes: Obtain the device parameter values of the target device for deploying the target question-answering large language model, where the device parameter values include at least one of the following: storage space size, operation speed; Determine the channels to be pruned of the pre-trained original question-answering large language model according to the device parameter values, where the pre-trained original question-answering large language model is used to reason about questions from the client to obtain answers; Based on the question-answering sample data, use a loss function with a regularization loss term to update the model parameters of the pre-trained original question-answering large language model, where the regularization loss term is determined according to the parameter matrix of the channels to be pruned; Prune the channels of the question-answering large language model after model parameter update according to the channels to be pruned, and obtain the target question-answering large language model based on the question-answering large language model after channel pruning; For the problems from the target device, reasoning is performed through the target question-answering large language model to obtain an answer.
[0006] Optionally, according to the device parameter value, determining the channels to be pruned of the pre-trained original question-answering large language model includes: Determining the number of parameters of the target question-answering large language model according to the device parameter value; Determining the pruning ratio according to the number of parameters and the original number of parameters of the pre-trained original question-answering large language model; Determining the channels to be pruned of the pre-trained original question-answering large language model according to the pruning ratio.
[0007] Optionally, determining the channels to be pruned of the pre-trained original question-answering large language model according to the pruning ratio includes: Constructing a pseudo-index selection matrix R according to the pruning ratio p%, where the pseudo-index selection matrix R is a diagonal matrix, and p% of the elements in the pseudo-index selection matrix R are 1, and 1 - p% of the elements are 0; Determining the regularization loss value corresponding to the regularization loss term according to the pseudo-index selection matrix R and the original parameter matrix W of the pre-trained original question-answering large language model; Updating the model parameters of the pre-trained original question-answering large language model using a loss function with a regularization loss term includes: For the first question in the question-and-answer sample data, obtaining a predicted answer through the pre-trained original question-answering large language model; Determining the original question-and-answer loss value of the pre-trained original question-answering large language model according to the difference between the predicted answer and the first answer in the question-and-answer sample data; Updating the model parameters of the pre-trained original question-answering large language model according to the original question-and-answer loss value and the regularization loss value corresponding to the regularization loss term.
[0008] Optionally, pruning the channels of the question-answering large language model after model parameter update according to the channels to be pruned includes: Constructing a channel pruning matrix according to the pseudo-index selection matrix R and the identity matrix I S , S = I - R ; Obtaining the parameter matrix of the question-answering large language model after channel pruning according to the parameter matrix of the question-answering large language model after model parameter update and the channel pruning matrix S .
[0009] Optionally, selecting the matrix R according to the pseudo-index and the original parameter matrix W of the pre-trained original question-answering large language model, and determining the regular loss value corresponding to the regular loss term, including: According to the pseudo-index selection matrix R and the i original parameter matrix of the layer of the attention mechanism module of the pre-trained original question-answering large language model , , , determine the attention regularization loss value of the i layer; According to the pseudo-index selection matrix R and the upper projection original parameter matrix in the i layer of the feed-forward neural network module of the pre-trained original question-answering large language model and the lower projection original parameter matrix , determine the feed-forward neural network regularization loss value of the i layer; According to the pseudo-index selection matrix R and the original parameter matrix of the remaining modules of the pre-trained original question-answering large language model, determine the regular loss value of the remaining modules.
[0010] Optionally, the method further includes: Obtaining the question-answering sample data set used to pre-train the original question-answering large language model; Randomly extracting the question-answering sample data set to obtain a question-answering sample data subset; Based on the question-answering sample data, using a loss function with a regular loss term to update the model parameters of the pre-trained original question-answering large language model, including: Inputting each question-answering sample data in the question-answering sample data subset into the pre-trained original question-answering large language model, and using a loss function with a regular loss term to update the model parameters of the pre-trained original question-answering large language model; Based on the question-answering large language model after channel pruning, obtaining the target question-answering large language model, including: Using the question-answering sample data subset to fine-tune the question-answering large language model after channel pruning to obtain the target question-answering large language model.
[0011] Optionally, the method further includes: For the test question-answering sample data, obtaining the predicted answer generated by the target question-answering large language model and the predicted answer generated by the pre-trained original question-answering large language model; In the case where the difference between the predicted answer generated by the target question-and-answer large language model and the predicted answer generated by the pre-trained original question-and-answer large language model is greater than or equal to a preset difference, continuously adjust the channels to be pruned, and return to the step of inputting the question-and-answer sample data into the pre-trained original question-and-answer large language model, and use a loss function with a regularization loss term to update the model parameters of the pre-trained original question-and-answer large language model until the difference between the predicted answer generated by the finally obtained target question-and-answer large language model and the predicted answer generated by the pre-trained original question-and-answer large language model is less than the preset difference; For a question from the target device, perform inference through the finally obtained target question-and-answer large language model to obtain an answer.
[0012] Optionally, continuously adjusting the channels to be pruned includes: Continuously adjusting the number of elements with a value of 1 in the pseudo-index selection matrix R, or continuously adjusting the positions of the elements with a value of 1 in the pseudo-index selection matrix R.
[0013] According to the second aspect of the embodiments of the present application, there is provided a question-and-answer system for a question-and-answer large language model based on data-driven regularization, which is used to implement the steps in the method provided in the first aspect of the embodiments of the present application. The system includes: A preprocessing module, configured to obtain device parameter values of a target device for deploying a target question-and-answer large language model, where the device parameter values include at least one of the following: storage space size, operation speed; according to the device parameter values, determine channels to be pruned of a pre-trained original question-and-answer large language model, and the pre-trained original question-and-answer large language model is used to perform inference on questions from a client to obtain answers; An information migration module, configured to update the model parameters of the pre-trained original question-and-answer large language model based on question-and-answer sample data by using a loss function with a regularization loss term, and the regularization loss term is determined according to the parameter matrix of the channels to be pruned; A pruning module, configured to perform channel pruning on the question-and-answer large language model after model parameter update according to the channels to be pruned, and obtain the target question-and-answer large language model based on the question-and-answer large language model after channel pruning; An inference module, configured to perform inference on questions from the target device through the target question-and-answer large language model to obtain answers.
[0014] According to the third aspect of the embodiments of the present application, there is provided an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps in the method provided in the first aspect of the embodiments of the present application are implemented.
[0015] Using the question-answering method of the question-answering large language model based on data-driven regularization provided by this application, first determine the device parameter values (storage space size, computing speed, etc.) of the target device (i.e., lightweight device) where the pruned target question-answering large language model needs to be deployed finally, so as to determine the channels that need to be pruned for the pre-trained original model according to the device parameter values, so that the pruned model can run normally on the target device. After determining the channels to be pruned, determine the regularization loss term according to the parameter matrix of the channels to be pruned, and use the loss function with the regularization loss term to train the pre-trained original model to update the model parameters, so that the important information of the model is transferred to the remaining channels. Then, prune the original question-answering large language model according to the channels to be pruned to obtain the target question-answering large language model. Deploy the pruned lightweight model on the lightweight target device to use the pruned lightweight model to reason about the questions raised by the user in the target device to obtain answers.
[0016] The lightweight model pruned by this application can run efficiently on lightweight devices, greatly reducing the question reasoning response time on the target device, and at the same time reducing the requirements for the hardware configuration and computing resource overhead of the target device. And, because the model is regularized in advance to transfer important information to the channels that need to be retained, the information loss caused by parameter removal during the pruning process is reduced, and the pruned model can maintain the reasoning performance comparable to the original large language model. In addition, because the model performance is maintained, there is no need to use a large amount of data for debugging to restore the model performance after pruning, further saving the time cost and computing resources for model deployment. Brief Description of the Drawings
[0017] In order to more clearly illustrate the technical solutions of the embodiments of this application, the following will briefly introduce the drawings required to be used in the description of the embodiments of this application. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to these drawings.
[0018] Figure 1 is a flowchart of the question-answering method of the question-answering large language model based on data-driven regularization proposed in an embodiment of this application; Figure 2 is a flowchart of the regularization process of each layer of the large language model in an embodiment of this application; Figure 3 is a schematic diagram of migrating important information through the regularization process in an embodiment of this application; Figure 4 is a flowchart of pruning the large language model proposed in an embodiment of this application; Figure 5It is a schematic diagram of a question-answering system for a question-answering large language model based on data-driven regularization proposed in an embodiment of the present application; Figure 6 It is a schematic diagram of an electronic device proposed in an embodiment of the present application. Detailed implementation manners
[0019] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0020] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of the present application. Therefore, the appearances of "in one embodiment" or "in an embodiment" throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0021] In various embodiments of the present application, it should be understood that the order numbers of the following processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0022] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are only examples of devices and methods consistent with some aspects detailed in the present application.
[0023] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other.
[0024] Currently, when deploying large language models for question answering on lightweight devices, there is a problem that the inference response time is too long, and it has high requirements for the device's hardware configuration, memory occupancy, and computational overhead, seriously affecting the user experience. To improve the response efficiency of the model, this application prunes the model to compress the model size, reduce the computational overhead of running the large language model for question answering on lightweight devices, improve the response efficiency of question inference, and thus improve the user experience. Traditional model pruning methods, such as structured pruning, when reducing the size of a large model, follow the steps of first selecting the channels or layers to be pruned according to the designed metrics, and then performing recovery fine-tuning. Since important information may be located in the selected parts to be pruned, direct pruning removes valuable information from the model, resulting in an irreversible decline in model performance. If the pruning ratio is relatively high, it will also lead to model performance collapse. To restore the model performance after pruning, a large amount of training data is required for recovery fine-tuning (RFT), increasing the time cost and computational overhead, resulting in insignificant model compression effects and limiting the effectiveness of the pruning scheme in reducing latency and improving throughput.
[0025] The solution of this application adopts the method of first migrating valuable information to the remaining part of the model and then pruning. By transferring important information to the remaining part of the model in advance, the information loss caused by removing some channels during the pruning process is reduced, the model size is compressed while maintaining the language modeling ability of the model, thereby improving the question inference efficiency of the large language model for question answering deployed on lightweight devices, and reducing the requirements of the model for device hardware configuration and computational overhead, and improving the user experience. The following will describe this application in detail with reference to the drawings and in combination with embodiments.
[0026] Figure 1 It is a flowchart of a question answering method for a large language model for question answering based on data-driven regularization proposed in an embodiment of this application. As Figure 1 shown, the method includes: S1: Obtain the device parameter values of the target device for deploying the target large language model for question answering, where the device parameter values include at least one of the following: storage space size, operation speed; S2: According to the device parameter values, determine the channels to be pruned of the pre-trained original large language model for question answering, where the pre-trained original large language model is used to infer answers to questions from the client; S3: Based on the question and answer sample data, use a loss function with a regularization loss term to update the model parameters of the pre-trained original large language model for question answering, where the regularization loss term is determined according to the parameter matrix of the channels to be pruned; S4: Prune the channels of the large language model for question answering after model parameter update according to the channels to be pruned, and obtain the target large language model for question answering based on the large language model for question answering after channel pruning; S5: For the problem from the target device, perform inference through the target question-and-answer large language model to obtain an answer.
[0027] In this embodiment, according to the device parameter value of the target device, determine the channels that need to be pruned in the original large language model. Specifically, according to at least one of the actual available storage space size and operation speed in the target device, determine the channels that need to be pruned in the original question-and-answer large language model.
[0028] The original question-and-answer large language model is a model pre-trained using a widely used benchmark dataset. In this embodiment, determine the question-and-answer sample data based on the benchmark dataset for subsequent regularization. During the regularization process, construct a loss function with a regularization loss term according to the parameter matrix of the channels that need to be pruned, and use the question-and-answer sample data to update the parameters of the original question-and-answer large language model, that is, apply regularization to the channels that need to be pruned using the constructed regularization loss term, which redistributes the important information in the channels that need to be pruned to the unpruned components. After performing the regularization process, perform channel pruning on the question-and-answer large language model after parameter update according to the previously determined channels that need to be pruned to obtain the target question-and-answer large language model. Deploy the target question-and-answer large language model to the target device, and then use the target question-and-answer large language model to perform inference on the questions from the target device and output answers (i.e., inference results). Through the pruning operation, the scale of the original large language model and the computational amount of question inference are reduced, and the pruned model is adapted to the hardware configuration, storage space size, and operation speed of the target device. Therefore, it can achieve efficient inference on the questions from the target device, reduce the response time of the question-and-answer, and improve the user experience.
[0029] Since the important information of the model is transferred from the part to be pruned to the remaining part of the model through regularization in advance, the information loss in the subsequent pruning process is reduced, and the performance of the pruned model is maintained to a great extent. Subsequently, there is no need to use a large amount of training data to perform performance recovery fine-tuning (RFT) on the model, thereby saving the computational overhead and time cost of model recovery fine-tuning and improving the efficiency of deploying the question-and-answer large language model on lightweight devices. This solution can also maintain the inference performance of the pruned model equivalent to that of the original model at a relatively high pruning ratio, meet the deployment requirements of the large language model in resource-constrained scenarios, enable the model to achieve a relatively high sparsity while maintaining performance, and improve the inference speed and data throughput efficiency of the question-and-answer large language model in lightweight devices.
[0030] As an implementation manner of the present application, determining the channels that need to be pruned in the pre-trained original question-and-answer large language model according to the device parameter value includes: According to the device parameter value, determine the number of parameters of the target question-and-answer large language model; Determine a pruning ratio based on the parameter quantity and the original parameter quantity of the pre-trained original question-answering large language model; Determine the channels that need to be pruned from the pre-trained original question-answering large language model according to the pruning ratio.
[0031] In this embodiment, the device parameter values include the storage space size and the operation speed. To determine the channels that need to be pruned according to the device parameter values, first, based on the device parameter values, determine the parameter quantity of the target question-answering large language model that needs to be deployed on the target device. Determine the channels that need to be pruned based on the original question-answering large language model according to the parameter quantity, so that the lightweight model (target question-answering large language model) obtained after pruning matches the device parameter values of the target device, ensuring that the target device can run the target question-answering large language model normally.
[0032] The parameter quantity of the target question-answering large language model includes: storage parameter quantity and operation parameter quantity. For the storage parameter quantity, it is determined according to the actual available storage space size of the target device. Based on the memory size occupied by each parameter in the model and the actual available storage space size of the target device, determine the storage parameter quantity of the target question-answering large language model.
[0033] For the operation parameter quantity, it is determined according to the operation speed of the target device. Based on the operation speed of the target device and the calculation amount of the model for a single inference, determine the operation parameter quantity of the target question-answering large language model.
[0034] In this embodiment, at least one of the storage parameter quantity and the operation parameter quantity is used to determine the channels that need to be pruned. Optionally, compare the storage parameter quantity and the operation parameter quantity, subtract the smaller value of the two from the original parameter quantity of the original question-answering large language model, and then determine the pruning ratio to ensure the optimal operation efficiency of the target question-answering large language model on the target device. Determine the channels that need to be pruned in the original question-answering large language model according to the pruning ratio.
[0035] As an implementation manner of the present application, determining the channels that need to be pruned from the pre-trained original question-answering large language model according to the pruning ratio includes: Construct a pseudo-index selection matrix R according to the pruning ratio p%, where the pseudo-index selection matrix R is a diagonal matrix, and p% of the elements in the pseudo-index selection matrix R are 1, and 1 - p% of the elements are 0; Determine the regularization loss value corresponding to the regularization loss term according to the pseudo-index selection matrix R and the original parameter matrix W of the pre-trained original question-answering large language model; Use the loss function with the regularization loss term to update the model parameters of the pre-trained original question-answering large language model, including: For the first question in the Q&A sample data, obtain a predicted answer through the pre-trained original Q&A large language model; Determine the original Q&A loss value of the pre-trained original Q&A large language model according to the difference between the predicted answer and the first answer in the Q&A sample data; Update the model parameters of the pre-trained original Q&A large language model according to the original Q&A loss value and the regularization loss value corresponding to the regularization loss term.
[0036] In one embodiment, determine the pruning ratio p% of the pre-trained original Q&A large language model according to the device parameter value, and then determine the channels to be pruned according to the pruning ratio. In this embodiment, the range of p in the pruning ratio p% is [10, 60]. Construct a pseudo-index selection matrix R according to the pruning ratio p%, and use this pseudo-index selection matrix to construct the regularization loss term in the loss function. The pseudo-index selection matrix R is a diagonal matrix. According to the pruning ratio p% , the elements on the diagonal of this matrix p% are 1, and the remaining elements on the diagonal (i.e., the elements of 1 -p% ) are 0. Determine the regularization loss term in the loss function according to the pseudo-index selection matrix R and the parameter matrix W of the original Q&A large language model, and construct a loss function with a regularization loss term. Train the original Q&A large language model based on this loss function to perform the transfer of important information. During the training process, for the first question in the Q&A sample data, perform inference through the original Q&A large language model to obtain a predicted answer. Further, determine the original Q&A loss value of the original Q&A large language model according to the difference between the predicted answer and the first answer to the first question in the Q&A sample data. Update the model parameters of the original Q&A large language model based on this original Q&A loss value.
[0037] As an implementation manner of the present application, determining the regularization loss value corresponding to the regularization loss term according to the pseudo-index selection matrix R and the original parameter matrix W of the pre-trained original Q&A large language model includes: According to the pseudo-index selection matrix R and the original parameter matrix of the i th layer of the attention mechanism module of the pre-trained original Q&A large language model , , , , determine the attention regularization loss value of the i th layer; According to the pseudo-index selection matrix R and the upper projection original parameter matrix in the i th layer of the feed-forward neural network module of the pre-trained original Q&A large language model and the lower projection original parameter matrix , determine the regularization loss value of the feedforward neural network for the i th layer; According to the pseudo-index selection matrix R and the original parameter matrix of the remaining modules of the pre-trained original question-answering large language model, determine the regularization loss value of the remaining modules.
[0038] In one embodiment, the loss function during the regularization process of the model adopts a combination of language modeling loss and regularization loss, where the regularization loss includes the loss brought by the regularization of the attention mechanism module, the forward propagation network, and the remaining part of the model. Figure 2 is a flowchart of the regularization process for each layer of the large language model in an embodiment of the present application. As Figure 2 shown, during the regularization process of each layer of the model, the regularization loss of the attention module , the regularization loss of the forward propagation network in the FFN (Feedforward Neural Network), and the regularization loss of the remaining part of the model (including layer normalization) are included.
[0039] The expression of the loss function is specifically as follows: ; where, is the language modeling loss of the model, used to measure the difference between the predicted answer and the true answer of the model; is the regularization loss of the attention mechanism module for the i th layer of the model; is the regularization loss of the forward propagation network for the i th layer of the model; is the regularization loss of the remaining part of the model except for the pruning part; is the number of layers of the model.
[0040] Specifically, the expression of the regularization loss of the attention mechanism module for the th layer of the model is as follows: ; where R represents the pseudo-index selection matrix (diagonal matrix). In the case where the pruning ratio is , R has elements equal to 1 on the diagonal, and elements equal to 0; represents a certain vector norm applied to the matrix column-wise, , , , They are the query matrix, key matrix, value matrix, and output matrix in the i layer respectively.
[0041] The regularization loss of the FFN forward propagation network in the i layer of the model is expressed as follows: ; where , are the upper projection matrix and lower projection matrix in the i layer respectively.
[0042] The regularization loss of the remaining part of the model except the pruning part is expressed as follows: ; where represents the data feature matrix, represents the position encoding matrix, represents the layer normalization vector, represents the language modeling matrix.
[0043] As an implementation manner of this application, according to the channels to be pruned, channel pruning is performed on the question-answering large language model after model parameter update, including: Construct a channel pruning matrix S , S = I - R ; According to the parameter matrix of the question-answering large language model after model parameter update and the channel pruning matrix S, obtain the parameter matrix of the question-answering large language model after channel pruning.
[0044] In this embodiment, after completing the regularization process of the model, according to the channels to be pruned determined before, channel pruning is performed on the question-answering large language model after parameter update, and the model parameter channels that have been regularized are pruned. Specifically, based on the pseudo-index selection matrix R , define the matrix , where I is the identity matrix. Use the matrix S for pruning. The attention module in the i layer of the model is pruned through the following expressions: ; ; ; .
[0045] Among them, , , , are the query matrix, key matrix, value matrix, and output matrix in the th layer after the model parameters are updated, respectively; , , , are the query matrix, key matrix, value matrix, and output matrix in the th layer of the pruned model, respectively.
[0046] The FFN forward propagation network of the th layer of the model is pruned through the following expression: ; ; Among them, , are the upper projection matrix and lower projection matrix in the i th layer after the model parameters are updated, respectively; , are the upper projection matrix and lower projection matrix in the i th layer of the pruned model, respectively.
[0047] The remaining part of the model is pruned through the following expression: ; ; ; ; Among them, is the data feature matrix after the model parameters are updated, is the position encoding matrix after the model parameters are updated, is the i th layer normalization vector after the model parameters are updated, is the language modeling matrix after the model parameters are updated; is the data feature matrix of the pruned model, is the position encoding matrix of the pruned model, is the i th layer normalization vector of the pruned model, is the language modeling matrix of the pruned model.
[0048] Figure 3 is a schematic diagram of migrating important information through the regularization process in an embodiment of the present application. As Figure 3As shown, in this embodiment, before pruning the model, a regularization process is performed. Based on the Q&A sample data in advance, the parameters of the model are updated using a loss function with a regularization loss term, so that the important information ( Figure 3 the part marked with an exclamation mark in
[0049] Figure 3 ) contained in the channels to be pruned is migrated to the remaining part of the model. After the information migration is completed, the pruning process is then carried out, and the channels to be pruned are removed using the S matrix. Since the information of the channels to be pruned is migrated to other channels in advance, the pruned model can still retain the key information of the original Q&A large language model.
[0050] As an implementation manner of this application, the method further includes: Obtain the Q&A sample data set used to obtain the original Q&A large language model through pre-training; Randomly sample the Q&A sample data set to obtain a Q&A sample data subset; Based on the Q&A sample data, update the model parameters of the pre-trained original Q&A large language model using a loss function with a regularization loss term, including: Input each Q&A sample data in the Q&A sample data subset into the pre-trained original Q&A large language model, and update the model parameters of the pre-trained original Q&A large language model using a loss function with a regularization loss term; Based on the Q&A large language model after channel pruning, obtain the target Q&A large language model, including: Use the Q&A sample data subset to fine-tune the Q&A large language model after channel pruning to obtain the target Q&A large language model.
[0051] In one embodiment, before updating the parameters of the original question-and-answer large language model using a loss function with a regularization loss term, a question-and-answer sample data set of the original question-and-answer large language model (e.g., a benchmark data set) is obtained, and the question-and-answer sample data set is randomly sampled to obtain a subset of question-and-answer sample data composed of a small amount of question-and-answer sample data. The regularization process of the model is performed using the extracted subset of question-and-answer sample data. The question-and-answer sample data in the subset of question-and-answer sample data is input into the original question-and-answer large language model, and the parameters of the model are updated using a loss function with a regularization loss term to transfer important information to the remaining channels that do not need to be pruned. Then, pruning operations are performed on the large model with updated parameters.
[0052] In this embodiment, after pruning the question-and-answer large language model with updated parameters, the pruned model is further fine-tuned using the data in the training phase (i.e., the subset of question-and-answer sample data) to obtain a target question-and-answer large language model that can be finally used for deployment. Since this solution first performs regularization to transfer important information and then performs model pruning, important information can be retained in the components of the pruned model, and the performance of the model is hardly affected. Therefore, compared with traditional pruning schemes, in this embodiment, only a very small amount of training data is needed to perform recovery fine-tuning (RFT) on the pruned model to meet the performance requirements. Furthermore, the model compression efficiency is greatly improved, and the computational overhead and time cost in the regularization and recovery fine-tuning processes are saved.
[0053] Figure 4 It is a flowchart of pruning a large language model proposed in an embodiment of the present application. As Figure 4 shown, in this embodiment, first, the parameters of the original question-and-answer large language model are updated using a loss function with a regularization loss term to transfer the important information of the channels to be pruned in the model to the remaining part of the model, and then the model is pruned to obtain a pruned question-and-answer large language model. Further, using the subset of question-and-answer sample data, the pruned question-and-answer large language model is fine-tuned to obtain a target question-and-answer large language model (i.e., the compressed model) finally deployed to the target device, completing the compression of the original question-and-answer large language model while maintaining the performance of the model.
[0054] As an implementation manner of the present application, the method further includes: For the test question-and-answer sample data, obtaining the predicted answer generated by the target question-and-answer large language model and the predicted answer generated by the pre-trained original question-and-answer large language model; In the case where the difference between the predicted answer generated by the target question-and-answer large language model and the predicted answer generated by the pre-trained original question-and-answer large language model is greater than or equal to a preset difference, continuously adjust the channels to be pruned, and return to the step of inputting the question-and-answer sample data into the pre-trained original question-and-answer large language model, and use a loss function with a regularization loss term to update the model parameters of the pre-trained original question-and-answer large language model until the difference between the predicted answer generated by the finally obtained target question-and-answer large language model and the predicted answer generated by the pre-trained original question-and-answer large language model is less than the preset difference; For the question from the target device, perform inference through the finally obtained target question-and-answer large language model to obtain an answer.
[0055] In one embodiment, input the test question-and-answer sample data into the target question-and-answer large language model and the original question-and-answer large language model respectively, obtain the predicted answers generated by the two models respectively, and compare the difference between the two predicted answers with the preset difference. If the difference between the two predicted answers is greater than the preset difference, return to re-determine the channels for pruning the original large model, and reconstruct a loss function with a regularization loss term based on the adjusted pruning channels, use the reconstructed loss function to update the parameters of the original question-and-answer large language model, then perform pruning and fine-tuning to obtain a new target question-and-answer large language model, and again compare the difference between the two predicted answers output by the new target question-and-answer large language model and the original question-and-answer large language model with the preset difference.
[0056] Repeat the above steps until the difference between the two predicted answers output by the target question-and-answer large language model and the original question-and-answer large language model is less than the preset difference, then it is determined that the performance of the target question-and-answer large language model is equivalent to the performance of the original question-and-answer large language model, that is, the compression of the original question-and-answer large language model is completed, and a lightweight model that can be deployed on the target device is obtained. Deploy the target question-and-answer large language model on the target device, and use this model to perform inference on the questions submitted by the user on the target device to obtain an answer.
[0057] In this embodiment, the inference performance of the target question-and-answer large language model is measured by a preset difference. In the case where the preset difference requirement is not met, the target question-and-answer large language model is adjusted by returning to modify the pruning channels to ensure that the lightweight model after pruning has an inference performance equivalent to that of the original question-and-answer large language model.
[0058] As an implementation manner of the present application, continuously adjusting the channels to be pruned includes: Continuously adjust the number of elements with a value of 1 in the pseudo-index selection matrix R, or continuously adjust the positions of the elements with a value of 1 in the pseudo-index selection matrix R.
[0059] In one embodiment, if the difference between the predicted answer generated by the target question-and-answer large language model and the predicted answer generated by the pre-trained original question-and-answer large language model is greater than or equal to a preset difference, it is necessary to adjust the pruned channels. Adjusting the pruned channels can adjust the positions of the pruned channels or the number of pruned channels. Specifically, by changing the number of elements with a value of 1 on the diagonal of the pseudo-index selection matrix R, the adjustment of the number of pruned channels is achieved. The more elements with a value of 1, the more channels to be pruned; by changing the positions of the elements with a value of 1 on the diagonal of the pseudo-index selection matrix R, the adjustment of the corresponding positions of the pruned channels is achieved.
[0060] Optionally, when the inference performance of the target question-and-answer large language model and the original question-and-answer large language model is not equivalent (i.e., the difference between the predicted answers is not less than the preset difference), at least one of the following operations is performed to regenerate a new target question-and-answer large language model: (1) Randomly adjust the positions of the elements with a value of 1 in the pseudo-index selection matrix R; (2) Randomly adjust the number of elements with a value of 1 in the pseudo-index selection matrix R.
[0061] Based on the same inventive concept, an embodiment of the present application provides a question-and-answer system for a question-and-answer large language model based on data-driven regularization. Figure 5 It is a schematic diagram of a question-and-answer system 100 for a question-and-answer large language model based on data-driven regularization proposed in an embodiment of the present application. As Figure 5 shown, the system includes: A preprocessing module 101, configured to obtain device parameter values of a target device for deploying a target question-and-answer large language model, where the device parameter values include at least one of the following: storage space size, operation speed; according to the device parameter values, determine the channels to be pruned in a pre-trained original question-and-answer large language model, and the pre-trained original question-and-answer large language model is used to infer answers to questions from a client; An information migration module 102, configured to update the model parameters of the pre-trained original question-and-answer large language model based on question-and-answer sample data by using a loss function with a regularization loss term, and the regularization loss term is determined according to the parameter matrix of the channels to be pruned; A pruning module 103, configured to prune the channels of the question-and-answer large language model after model parameter update according to the channels to be pruned, and obtain the target question-and-answer large language model based on the question-and-answer large language model after channel pruning; An inference module 104, configured to infer answers to questions from the target device through the target question-and-answer large language model.
[0062] As an implementation manner of the present application, the preprocessing module 101 is configured to determine the channels to be pruned of the pre-trained original question-and-answer large language model according to the device parameter values, specifically including: Determine the number of parameters of the target question-and-answer large language model according to the device parameter values; Determine the pruning ratio according to the number of parameters and the original number of parameters of the pre-trained original question-and-answer large language model; Determine the channels to be pruned of the pre-trained original question-and-answer large language model according to the pruning ratio.
[0063] As an implementation manner of the present application, the preprocessing module 101 is configured to determine the channels to be pruned of the pre-trained original question-and-answer large language model according to the pruning ratio, specifically including: Construct a pseudo-index selection matrix R according to the pruning ratio p%, the pseudo-index selection matrix R is a diagonal matrix, and p% of the elements in the pseudo-index selection matrix R are 1, and 1 - p% of the elements are 0; Determine the regularization loss value corresponding to the regularization loss term according to the pseudo-index selection matrix R and the original parameter matrix W of the pre-trained original question-and-answer large language model; The information transfer module 102 is configured to update the model parameters of the pre-trained original question-and-answer large language model by using a loss function with a regularization loss term, specifically including: For the first question in the question-and-answer sample data, obtain a predicted answer through the pre-trained original question-and-answer large language model; Determine the original question-and-answer loss value of the pre-trained original question-and-answer large language model according to the difference between the predicted answer and the first answer in the question-and-answer sample data; Update the model parameters of the pre-trained original question-and-answer large language model according to the original question-and-answer loss value and the regularization loss value corresponding to the regularization loss term.
[0064] As an implementation manner of the present application, the pruning module 103 is configured to perform channel pruning on the question-and-answer large language model after model parameter update according to the channels to be pruned, specifically including: Construct a channel pruning matrix according to the pseudo-index selection matrix R and the identity matrix I S , S = I - R ; According to the parameter matrix of the question-and-answer large language model after model parameter update and the channel pruning matrix S, obtain the parameter matrix of the question-and-answer large language model after channel pruning.
[0065] As an implementation manner of the present application, the preprocessing module 101 is configured to determine the regularization loss value corresponding to the regularization loss term according to the pseudo-index selection matrix R and the original parameter matrix W of the pre-trained original question-answering large language model, including: According to the pseudo-index selection matrix R and the original parameter matrix of the i layer of the attention mechanism module of the pre-trained original question-answering large language model , , , , determine the attention regularization loss value of the i layer; According to the pseudo-index selection matrix R and the upper projection original parameter matrix in the i layer of the feed-forward neural network module of the pre-trained original question-answering large language model and the lower projection original parameter matrix , determine the feed-forward neural network regularization loss value of the i layer; According to the pseudo-index selection matrix R and the original parameter matrix of the remaining module of the pre-trained original question-answering large language model, determine the remaining module regularization loss value.
[0066] As an implementation manner of the present application, the preprocessing module 101 is further configured to perform the following operations: obtain the question-answering sample data set used to pre-train the original question-answering large language model; randomly sample the question-answering sample data set to obtain a question-answering sample data subset; The information migration module 102 is configured to update the model parameters of the pre-trained original question-answering large language model based on the question-answering sample data by using a loss function with a regularization loss term, specifically including: Input each question-answering sample data in the question-answering sample data subset into the pre-trained original question-answering large language model, and use a loss function with a regularization loss term to update the model parameters of the pre-trained original question-answering large language model; The pruning module 103 is configured to obtain the target question-answering large language model based on the question-answering large language model after channel pruning, specifically including: Use the question-answering sample data subset to fine-tune the question-answering large language model after channel pruning to obtain the target question-answering large language model.
[0067] As an implementation manner of the present application, the system further includes a callback module, which is configured to perform the following steps: For the test question-answering sample data, obtain the predicted answer generated by the target question-answering large language model and the predicted answer generated by the pre-trained original question-answering large language model; In the case where the difference between the predicted answer generated by the target question-answering large language model and the predicted answer generated by the pre-trained original question-answering large language model is greater than or equal to a preset difference, continuously adjust the channels to be pruned, and return to the step: input the question-answering sample data into the pre-trained original question-answering large language model, and use a loss function with a regularization loss term to update the model parameters of the pre-trained original question-answering large language model until the difference between the predicted answer generated by the finally obtained target question-answering large language model and the predicted answer generated by the pre-trained original question-answering large language model is less than the preset difference; The inference module 104 is further configured to perform inference on the question from the target device through the finally obtained target question-answering large language model to obtain an answer.
[0068] As an implementation manner of the present application, the callback module is configured to continuously adjust the channels to be pruned, specifically including: Continuously adjust the number of elements with a value of 1 in the pseudo-index selection matrix R, or continuously adjust the positions of the elements with a value of 1 in the pseudo-index selection matrix R.
[0069] Based on the same inventive concept, an embodiment of the present application provides an electronic device. Figure 6 It is a schematic diagram of the electronic device proposed in an embodiment of the present application. As Figure 6 shown, the electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps in the question-answering method of the question-answering large language model based on data-driven regularization as described in any of the above embodiments of the present application.
[0070] Regarding the system in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0071] The above are only the preferred embodiments of the present application, and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
[0072] For the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and components involved are not necessarily essential to the present application.
[0073] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, an apparatus, or a computer program product. Therefore, the embodiments of the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0074] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing terminal devices generate means for realizing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.
[0075] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.
[0076] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable terminal device provide steps for realizing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.
[0077] Although the preferred embodiments of the embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concept. Therefore, the present application is construed as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present application.
[0078] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the said element.
[0079] The above has introduced in detail the question-answering method and product of the question-answering large language model based on data-driven regular expressions provided by this application. Specific examples are used in this text to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation on this application.
Claims
1. A question-answering method for a question-answering large language model based on data-driven regular expressions, characterized in that, The method includes: Obtaining the device parameter values of the target device for deploying the target question-and-answer large language model, where the device parameter values include at least one of the following: storage space size, computing speed; Determining the channels to be pruned of the pre-trained original question-and-answer large language model according to the device parameter values, where the pre-trained original question-and-answer large language model is used to reason about questions from a client to obtain answers; Based on the question-and-answer sample data, using a loss function with a regularization loss term to update the model parameters of the pre-trained original question-and-answer large language model, where the regularization loss term is determined according to the parameter matrix of the channels to be pruned; Pruning the channels of the question-and-answer large language model after model parameter update according to the channels to be pruned, and obtaining the target question-and-answer large language model based on the question-and-answer large language model after channel pruning; For questions from the target device, reasoning through the target question-and-answer large language model to obtain answers.
2. The question-answering method of the question-answering large language model based on data-driven regularization according to claim 1, characterized in that, Determining the channels to be pruned of the pre-trained original question-and-answer large language model according to the device parameter values includes: Determining the number of parameters of the target question-and-answer large language model according to the device parameter values; Determining the pruning ratio according to the number of parameters and the original number of parameters of the pre-trained original question-and-answer large language model; Determining the channels to be pruned of the pre-trained original question-and-answer large language model according to the pruning ratio.
3. The question-answering method of the question-answering large language model based on data-driven regularization according to claim 2, characterized in that, Determining the channels to be pruned of the pre-trained original question-and-answer large language model according to the pruning ratio includes: Constructing a pseudo-index selection matrix R according to the pruning ratio p%, where the pseudo-index selection matrix R is a diagonal matrix, and p% of the elements in the pseudo-index selection matrix R are 1, and 1 - p% of the elements are 0; Determining the regularization loss value corresponding to the regularization loss term according to the pseudo-index selection matrix R and the original parameter matrix W of the pre-trained original question-and-answer large language model; Using a loss function with a regularization loss term to update the model parameters of the pre-trained original question-and-answer large language model includes: For the first question in the question-and-answer sample data, obtaining a predicted answer through the pre-trained original question-and-answer large language model; Determining the original question-and-answer loss value of the pre-trained original question-and-answer large language model according to the difference between the predicted answer and the first answer in the question-and-answer sample data; Updating the model parameters of the pre-trained original question-and-answer large language model according to the original question-and-answer loss value and the regularization loss value corresponding to the regularization loss term.
4. The question-answering method of the question-answering large language model based on data-driven regularization according to claim 3, wherein, Pruning the channels of the question-and-answer large language model after model parameter update according to the channels to be pruned includes: Construct a channel pruning matrix by selecting the pseudo-index selection matrix R and the identity matrix I S , S = I - R ; According to the parameter matrix of the question-and-answer large language model after model parameter update and the channel pruning matrix S, obtain the parameter matrix of the question-and-answer large language model after channel pruning .
5. The question-answering method of the question-answering large language model based on data-driven regularization according to claim 3, characterized in that, Determining the regularization loss value corresponding to the regularization loss term according to the pseudo-index selection matrix R and the original parameter matrix W of the pre-trained original question-and-answer large language model includes: According to the pseudo-index selection matrix R and the original parameter matrix of the i th layer of the attention mechanism module of the pre-trained original question-answering large language model , , , , determine the attention regularization loss value of the i th layer; Select the upper projection original parameter matrix in the i layer of the feed-forward neural network module of the pre-trained original question-answering large language model according to the pseudo-index selection matrix R and the lower projection original parameter matrix to determine the feed-forward neural network regularization loss value of the i layer; Determining the remaining module regularization loss value according to the pseudo-index selection matrix R and the original parameter matrix of the remaining modules of the pre-trained original question-and-answer large language model.
6. The question-answering method of the question-answering large language model based on data-driven regularization according to claim 1, characterized in that, The method further includes: Obtaining the question-and-answer sample data set used to pre-train the original question-and-answer large language model; Randomly extract from the Q&A sample data set to obtain a Q&A sample data subset; Based on the Q&A sample data, use a loss function with a regularization loss term to update the model parameters of the pre-trained original Q&A large language model, including: Input each Q&A sample data in the Q&A sample data subset into the pre-trained original Q&A large language model, and use a loss function with a regularization loss term to update the model parameters of the pre-trained original Q&A large language model; Based on the Q&A large language model after channel pruning, obtain the target Q&A large language model, including: Use the Q&A sample data subset to fine-tune the Q&A large language model after channel pruning to obtain the target Q&A large language model.
7. The question-answering method of the question-answering large language model based on data-driven regularization according to claim 3, wherein The method further includes: For test Q&A sample data, obtain the predicted answer generated by the target Q&A large language model and the predicted answer generated by the pre-trained original Q&A large language model; In the case where the difference between the predicted answer generated by the target Q&A large language model and the predicted answer generated by the pre-trained original Q&A large language model is greater than or equal to a preset difference, continuously adjust the channels to be pruned, and return to the step: input the Q&A sample data into the pre-trained original Q&A large language model, and use a loss function with a regularization loss term to update the model parameters of the pre-trained original Q&A large language model until the difference between the predicted answer generated by the finally obtained target Q&A large language model and the predicted answer generated by the pre-trained original Q&A large language model is less than the preset difference; For a question from the target device, perform inference through the finally obtained target Q&A large language model to obtain an answer.
8. The question-answering method of the question-answering large language model based on data-driven regularization according to claim 7, wherein Continuously adjusting the channels to be pruned includes: Continuously adjust the number of elements with a value of 1 in the pseudo-index selection matrix R, or continuously adjust the positions of the elements with a value of 1 in the pseudo-index selection matrix R.
9. A question-answering system for a question-answering large language model based on data-driven regular expressions, characterized in that, For executing the method according to any one of claims 1-8, including: A preprocessing module, configured to obtain the device parameter values of the target device for deploying the target Q&A large language model, where the device parameter values include at least one of the following: storage space size, operation speed; according to the device parameter values, determine the channels to be pruned of the pre-trained original Q&A large language model, and the pre-trained original Q&A large language model is used to perform inference on questions from the client to obtain answers; An information migration module, configured to update the model parameters of the pre-trained original Q&A large language model based on the Q&A sample data by using a loss function with a regularization loss term, and the regularization loss term is determined according to the parameter matrix of the channels to be pruned; A pruning module, configured to perform channel pruning on the Q&A large language model after model parameter update according to the channels to be pruned, and obtain the target Q&A large language model based on the Q&A large language model after channel pruning; An inference module, configured to perform inference on a question from the target device through the target Q&A large language model to obtain an answer.
10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps in the method according to any one of claims 1-8 are implemented.
Citation Information
Patent Citations
Input pruning acceleration method for question and answer task model
CN113849601A
Large-scale pre-training language model compression method based on hardware perception
CN116822593A
Cloud-integrated embedded large language model training method and language question and answer method
CN117689041A
Network parameter processing method, task processing method and corresponding devices
CN119067182A
Large model pruning method, and response method, device and equipment based on large model
CN119227767A