Parameter setting method, text generation method, device, equipment, program product
By adaptively setting the low-rank conversion parameters of the pre-trained language model network layer and dynamically adjusting the parameters according to similarity, the information loss problem caused by low-rank conversion in the existing technology is solved, and the efficiency and performance of model fine-tuning are improved.
Patent Information
- Application Number
- CN202411876272.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-12-19
AI Technical Summary
The prior art uses fixed low-rank conversion parameters when performing low-rank conversion on the network layer weight matrix of the pre-trained language model, resulting in redundancy of shallow network parameters or loss of deep network information, affecting the efficiency and performance of model fine-tuning.
By adaptively setting low-rank conversion parameters for the network layer in the pre-trained language model, the low-rank conversion parameters are dynamically adjusted according to the similarity between adjacent network layers, ensuring that the deeper the network layer, the larger the low-rank conversion parameters are, thereby avoiding information loss.
It improves the efficiency and performance of model fine-tuning training, reduces the problem of deep network information loss, and ensures the improvement of low-rank conversion effect.
Smart Images

Figure CN119337240B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and particularly to a parameter setting method, a text generation method, an apparatus, a device, and a program product. Background Art
[0002] With the continuous development of artificial intelligence technology, large-scale pre-trained language models have gradually been applied in many fields. To ensure that the pre-trained language model can adapt to different domain tasks, a common model training method is to pre-train a basic language model on a huge general dataset, and then apply it to downstream domain tasks through fine-tuning technology. However, the huge number of model parameters brings great challenges to the subsequent fine-tuning process, such as resource requirements, training time, etc.
[0003] In related technologies, in order to reduce the resource consumption during the fine-tuning process of the model and improve the fine-tuning efficiency, a low-rank approximation method can be used to perform a low-rank transformation on the weight matrix of the model network layer, and use multiple matrices with lower dimensions to approximate the parameter changes of the original weight. However, when performing a low-rank transformation on the weight matrix of the model network layer in related technologies, fixed low-rank transformation parameters are usually used for processing, which easily leads to problems such as redundancy of shallow network parameters or serious loss of information in deep networks, thus affecting the efficiency and performance of model fine-tuning. Summary of the Invention
[0004] The purpose of the present invention is to provide a parameter setting method, a text generation method, an apparatus, a device, and a program product, which can adaptively set low-rank transformation parameters for network layers in a pre-trained language model, avoid using fixed low-rank transformation parameters, and thus improve the fine-tuning training effect of the model.
[0005] To solve the above technical problems, the present invention provides a parameter setting method, including:
[0006] Obtain a pre-trained language model and preset low-rank transformation parameters;
[0007] For network layers corresponding to the same network layer type in the pre-trained language model, set the preset low-rank transformation parameters as the low-rank transformation parameters of the top network layer, and determine the similarity between adjacent network layers;
[0008] If the similarity is not less than a preset threshold, use the low-rank transformation parameters of the upper network layer in the adjacent network layers as the low-rank transformation parameters of the lower network layer in the adjacent network layers;
[0009] If the similarity is less than the preset threshold, an adjustment coefficient is determined according to the similarity, and the low-rank transformation parameter of the upper network layer in the adjacent network layer is increased by using the adjustment coefficient to obtain the low-rank transformation parameter of the lower network layer in the adjacent network layer.
[0010] Optionally, determining an adjustment coefficient according to the similarity and increasing the low-rank transformation parameter of the upper network layer in the adjacent network layer by using the adjustment coefficient includes:
[0011] Determine a correlation factor according to the similarity; wherein, the similarity has a positive correlation with the correlation factor;
[0012] Take the reciprocal of the correlation factor as the adjustment coefficient, and increase the low-rank transformation parameter of the upper network layer in the adjacent network layer by using the adjustment coefficient.
[0013] Optionally, obtaining a preset low-rank transformation parameter includes:
[0014] Determine the network layer types corresponding to each network layer in the pre-trained language model;
[0015] Obtain the preset low-rank transformation parameters corresponding to each network layer type.
[0016] Optionally, for the network layers in the pre-trained language model corresponding to the same network layer type, setting the preset low-rank transformation parameter to the low-rank transformation parameter of the top network layer includes:
[0017] For the network layers in the pre-trained language model corresponding to the same network layer type, set the preset low-rank transformation parameter corresponding to the network layer type to the low-rank transformation parameter of the top network layer corresponding to the network layer type.
[0018] Optionally, it further includes:
[0019] Create different groups of preset low-rank transformation parameters; wherein, each group of preset low-rank transformation parameters stores the preset low-rank transformation parameters corresponding to each network layer type, and different groups of preset low-rank transformation parameters store different preset low-rank transformation parameters;
[0020] Determine different low-rank transformation parameters for the network layers in the pre-trained language model based on the different groups of preset low-rank transformation parameters;
[0021] Fine-tune and train each network layer in the pre-trained language model based on different low-rank transformation parameters to determine the model performance values corresponding to different groups of preset low-rank transformation parameters;
[0022] Determine the best group of preset low-rank transformation parameters according to the model performance values.
[0023] Optionally, determining the similarity between adjacent network layers includes:
[0024] Determining the feature distance between the weight matrices of the adjacent network layers and using the feature distance as the similarity.
[0025] Optionally, determining the feature distance between the weight matrices of the adjacent network layers includes:
[0026] Performing singular value decomposition on the first weight matrix of the upper network layer in the adjacent network layers to obtain a first non - negative diagonal matrix and a first right singular matrix, and determining the rank of the first non - negative diagonal matrix;
[0027] Determining a first screening quantity according to the rank of the first non - negative diagonal matrix, and screening the first singular vectors with the largest number of the first screening quantity in the first right singular matrix;
[0028] Performing singular value decomposition on the second weight matrix corresponding to the lower network layer in the adjacent network layers to obtain a second non - negative diagonal matrix and a second right singular matrix, and determining the rank of the second non - negative diagonal matrix;
[0029] Determining a second screening quantity according to the rank of the second non - negative diagonal matrix, and screening the second singular vectors with the largest number of the second screening quantity in the second right singular matrix;
[0030] Determining the similarity using the first singular vectors, the rank of the first non - negative diagonal matrix, the second singular vectors, and the rank of the second non - negative diagonal matrix.
[0031] Optionally, the pre - trained language model is composed of multiple transformer network layers connected in series.
[0032] Optionally, it further includes:
[0033] Performing matrix decomposition on the weight matrix of the network layer according to the low - rank transformation parameter to obtain the low - rank transformation weight matrix corresponding to the network layer;
[0034] Performing fine - tuning training on the low - rank transformation weight matrix of the network layer.
[0035] Optionally, performing fine - tuning training on the low - rank transformation weight matrix of the network layer includes:
[0036] Obtaining a preset training text;
[0037] Process the preset training text layer by layer using each network layer in the pre-trained language model to obtain model output data; wherein, the output of the network layer is the fused output data obtained by fusing the first output data and the second output data, the first output data is the processing result of the weight matrix of the network layer on the preset training text, and the second output data is the processing result of the low-rank transformation weight matrix of the network layer on the preset training text;
[0038] Determine a loss value based on the model output data, and update the parameters of the low-rank transformation weight matrix of each network layer according to the loss value.
[0039] Optionally, processing the preset training text layer by layer using each network layer in the pre-trained language model to obtain model output data includes:
[0040] Take the first network layer in the pre-trained language model as the current network layer, take the preset training text as the data to be processed, and input the data to be processed into the current network layer;
[0041] Process the data to be processed using the weight matrix of the current network layer to obtain the first output data;
[0042] Process the data to be processed using the low-rank transformation weight matrix of the current network layer to obtain the second output data, and perform scaling processing on the second output data using a scaling parameter;
[0043] Superimpose the first output data and the scaled second output data to obtain the fused output data;
[0044] Determine whether there is a lower network layer that has not processed the preset training text;
[0045] If so, take the fused output data of the current network layer as the data to be processed, update the current network layer to the lower network layer of the current network layer, and enter the step of inputting the data to be processed into the current network layer;
[0046] If not, take the fused output data of the current network layer as the model output data.
[0047] Optionally, before performing scaling processing on the second output data using the scaling parameter, it further includes:
[0048] Determine the scaling parameter corresponding to the current network layer using the low-rank transformation parameter and the preset scaling factor of the current network layer;
[0049] The step of performing scaling processing on the second output data using the scaling parameter includes:
[0050] Scale the second output data by using the scaling parameter corresponding to the current network layer.
[0051] The present invention further provides a text generation method, including:
[0052] Obtain user input text and a pre-trained language model; wherein, the network layers in the pre-trained language model include a weight matrix and a low-rank transformation weight matrix, the low-rank transformation weight matrix is generated by using the weight matrix and the low-rank transformation parameter corresponding to the network layer, and the low-rank transformation parameter is generated according to the parameter setting method as described above;
[0053] Process the user input text layer by layer by using each network layer in the pre-trained language model to obtain model output data; wherein, the output of the network layer is a fusion output data obtained by fusing a first output data and a second output data, the first output data is the processing result of the weight matrix of the network layer on the user input text, and the second output data is the processing result of the low-rank transformation weight matrix of the network layer on the user input text;
[0054] Generate text based on the model output data.
[0055] Optionally, processing the user input text layer by layer by using each network layer in the pre-trained language model to obtain model output data includes:
[0056] Take the first network layer in the pre-trained language model as the current network layer, take the user input text as the data to be processed, and input the data to be processed into the current network layer;
[0057] Process the data to be processed by using the weight matrix of the current network layer to obtain the first output data;
[0058] Process the data to be processed by using the low-rank transformation weight matrix of the current network layer to obtain the second output data, and scale the second output data by using the scaling parameter;
[0059] Superimpose the first output data and the scaled second output data to obtain the fusion output data;
[0060] Determine whether there is a lower network layer that has not processed the user input text;
[0061] If so, take the fusion output data of the current network layer as the data to be processed, update the current network layer to the lower network layer of the current network layer, and enter the step of inputting the data to be processed into the current network layer;
[0062] If not, use the fused output data of the current network layer as the model output data.
[0063] Optionally, before scaling the second output data using the scaling parameter, it further includes:
[0064] Determine the scaling parameter corresponding to the current network layer using the low-rank transformation parameter of the current network layer and a preset scaling factor;
[0065] The scaling the second output data using the scaling parameter includes:
[0066] Scale the second output data using the scaling parameter corresponding to the current network layer.
[0067] The present invention also provides a parameter setting device, including:
[0068] An acquisition module, configured to acquire a pre-trained language model and preset low-rank transformation parameters;
[0069] A network layer processing module, configured to set the preset low-rank transformation parameter as the low-rank transformation parameter of the top network layer for network layers in the pre-trained language model corresponding to the same network layer type, and determine the similarity between adjacent network layers;
[0070] A first low-rank transformation parameter setting module, configured to, if the similarity is not less than a preset threshold, use the low-rank transformation parameter of the upper network layer in the adjacent network layers as the low-rank transformation parameter of the lower network layer in the adjacent network layers;
[0071] A second low-rank transformation parameter setting module, configured to, if the similarity is less than the preset threshold, determine an adjustment coefficient according to the similarity, and increase the low-rank transformation parameter of the upper network layer in the adjacent network layers using the adjustment coefficient to obtain the low-rank transformation parameter of the lower network layer in the adjacent network layers.
[0072] The present invention also provides a text generation device, including:
[0073] An acquisition module, configured to acquire user input text and a pre-trained language model; wherein, the network layers in the pre-trained language model include a weight matrix and a low-rank transformation weight matrix, the low-rank transformation weight matrix is generated using the weight matrix and the low-rank transformation parameter corresponding to the network layer, and the low-rank transformation parameter is generated according to the parameter setting method as described above;
[0074] A model processing module, configured to layer - by - layer process the user input text by using each network layer in the pre - trained language model to obtain model output data; wherein, the output of each network layer is the fused output data obtained by fusing the first output data and the second output data, the first output data is the processing result of the weight matrix of the network layer on the user input text, and the second output data is the processing result of the low - rank transformation weight matrix of the network layer on the user input text;
[0075] A text generation module, configured to generate text based on the model output data.
[0076] The present invention also provides an electronic device, including:
[0077] A memory, configured to store a computer program;
[0078] A processor, configured to implement the parameter setting method or the text generation method as described above when executing the computer program.
[0079] The present invention also provides a computer program product, including a computer program or instruction, where the computer program or instruction implements the parameter setting method or the text generation method as described above when being executed by a processor.
[0080] The present invention also provides a non - volatile computer - readable storage medium, where computer - executable instructions are stored in the computer - readable storage medium, and when the computer - executable instructions are loaded and executed by a processor, the parameter setting method or the text generation method as described above is implemented.
[0081] The present invention provides a parameter setting method, including: obtaining a pre - trained language model and preset low - rank transformation parameters; for network layers corresponding to the same network layer type in the pre - trained language model, setting the preset low - rank transformation parameters as the low - rank transformation parameters of the top - layer network layer, and determining the similarity between adjacent network layers; if the similarity is not less than a preset threshold, using the low - rank transformation parameters of the upper network layer in the adjacent network layers as the low - rank transformation parameters of the lower network layer in the adjacent network layers; if the similarity is less than the preset threshold, determining an adjustment coefficient according to the similarity, and increasing the low - rank transformation parameters of the upper network layer in the adjacent network layers by using the adjustment coefficient to obtain the low - rank transformation parameters of the lower network layer in the adjacent network layers.
[0082] The beneficial effects of the present invention are as follows: First, the present invention can obtain a pre-trained language model and preset low-rank transformation parameters. For the network layers corresponding to the same network layer type in the pre-trained language model, the preset low-rank transformation parameters can be set as the low-rank transformation parameters of the top-layer network layer, and the similarity between adjacent network layers can be determined. Subsequently, the present invention determines whether the similarity between adjacent network layers is less than a preset threshold. If the similarity is not less than the preset threshold, it indicates that these adjacent network layers are relatively similar, and the same low-rank transformation parameters can be used for low-rank transformation. Furthermore, the low-rank transformation parameters of the upper network layer in the adjacent network layers can be used as the low-rank transformation parameters of the lower network layer in the adjacent network layers. If the similarity is less than the preset threshold, it indicates that these adjacent network layers are not similar, and larger low-rank transformation parameters need to be set for the lower network layer to retain information. Furthermore, an adjustment coefficient can be determined according to the similarity, and the low-rank transformation parameters of the upper network layer in the adjacent network layers can be increased by using the adjustment coefficient to obtain the low-rank transformation parameters of the lower network layer in the adjacent network layers. In this way, the present invention can adaptively set low-rank transformation parameters for the lower network layer according to the similarity between adjacent network layers, and generally ensure that the deeper the network layer, the larger the low-rank transformation parameters. This can effectively ensure the low-rank transformation effect of the model and reduce the problem of information loss in the deep network, thereby improving the efficiency and performance of model fine-tuning training. The present invention also provides a text generation method, a parameter setting device, a text generation device, an electronic device, a computer program product, and a computer-readable storage medium, which have the above beneficial effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0084] Figure 1 It is a flowchart of a parameter setting method provided by an embodiment of the present invention;
[0085] Figure 2 It is a schematic diagram of matrix decomposition provided by an embodiment of the present invention;
[0086] Figure 3 It is a flowchart of a text generation method provided by an embodiment of the present invention;
[0087] Figure 4 It is a schematic diagram of a parameter setting process provided by an embodiment of the present invention;
[0088] Figure 5 It is a structural block diagram of a parameter setting device provided by an embodiment of the present invention;
[0089] Figure 6 A structural block diagram of a text generation device provided by an embodiment of the present invention;
[0090] Figure 7 A structural block diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0091] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0092] A pre-trained language model is a machine learning model trained with a large amount of text data, mainly used for parsing and generating natural language text. The common application scenarios of this model are: receiving the question text input by the user and outputting the output text that meets the user's needs according to the question text. With the continuous development of artificial intelligence technology, large-scale pre-trained language models (such as LLM models, Large Language Models) have gradually been applied in many fields. To ensure that the pre-trained language model can adapt to different domain tasks, the common model training method is to pre-train a basic language model on a huge general dataset and then apply it to downstream domain tasks through fine-tuning technology. However, the huge number of model parameters brings great challenges to the subsequent fine-tuning process, such as resource requirements, training time, and so on.
[0093] In the related art, in order to reduce the resource consumption during the fine-tuning process of the model and improve the fine-tuning efficiency, the method of low-rank approximation can be used to perform low-rank transformation on the weight matrix of the model network layer, and multiple matrices with lower dimensions are used to approximate the parameter changes of the original weight. However, when performing low-rank transformation on the weight matrix of the model network layer in the related art, fixed low-rank transformation parameters are usually used for processing, which easily leads to the problem of redundant parameters in the shallow network or serious information loss in the deep network, thus affecting the efficiency and performance of model fine-tuning.
[0094] In view of this, aiming at the technical problem of how to better determine the low-rank transformation parameters for each network layer of the pre-trained language model, the present invention can provide a parameter setting method, which can adaptively set the low-rank transformation parameters for the network layers in the pre-trained language model, avoiding the use of fixed low-rank transformation parameters, thereby improving the fine-tuning training effect of the model.
[0095] For ease of understanding, please refer to Figure 1 , Figure 1 which is a flowchart of a parameter setting method provided by an embodiment of the present invention. This method may include:
[0096] S101. Obtain a pre-trained language model and preset low-rank transformation parameters.
[0097] In this step, the pre-trained language model refers to a language model that has been trained with a large amount of general data and needs to be fine-tuned. The specific structure of the pre-trained language model is not limited in this embodiment. For example, it may be composed of multiple cascaded Transformer network layers. Among them, the core part of the Transformer network layer includes an AttentionLayer and a Feed-Forward Neural Network (FFN), and the FFN occupies 70% of the model's parameters. By stacking the Transformer network layer N times, the model can extract data features multiple times and understand semantics.
[0098] Furthermore, this step also obtains preset low-rank transformation parameters. Different from the related technology, the preset low-rank transformation parameters obtained here are only used as the low-rank transformation parameters of the top network layer, rather than the low-rank transformation parameters of the remaining network layers. The low-rank transformation parameters of the remaining network layers will be adjusted based on the preset low-rank transformation parameters. It should be noted that the specific values of the preset low-rank transformation parameters are not limited in this embodiment and can be set according to actual application requirements.
[0099] S102. For the network layers corresponding to the same network layer type in the pre-trained language model, set the preset low-rank transformation parameters as the low-rank transformation parameters of the top network layer, and determine the similarity between adjacent network layers.
[0100] In this step, the low-rank transformation parameters will be set in the network layers of the same network layer type. Specifically, in this embodiment, the preset low-rank transformation parameters can be set as the low-rank transformation parameters of the top network layer in the network layers of the same network layer type; subsequently, the similarity between adjacent network layers can be determined to set the low-rank transformation parameters for each network layer below the top network layer according to this similarity. It should be pointed out here that the "adjacent network layers" here are hierarchically adjacent within the same category and have a logical adjacent relationship, and they are not adjacent in the actual model structure. For example, the network layers in the pre-trained language model can be classified as follows in this embodiment:
[0101] Table 1 Example list of low-rank approximation fine-tuning modules
[0102]
[0103] For category c1, there are N module sequences in total. The modules module i c1 and module (i+1) c1 are adjacent network layers, but they are not adjacent in the overall model structure because there will be network layers of other categories in between.
[0104] Furthermore, since the essence of low-rank transformation is to compress the weight matrix, and the uses and importance of different network layer types are different, different intensities can be adopted for low-rank transformation. For this reason, this embodiment can also set corresponding preset low-rank transformation parameters for different network layers. Specifically, this embodiment can determine the network layer types corresponding to each network layer in the pre-trained language model for classification; subsequently, it can obtain the preset low-rank transformation parameters corresponding to each network layer type, where the preset low-rank transformation parameters corresponding to different network layer types can be different. Furthermore, this embodiment can set the preset low-rank transformation parameter corresponding to the network layer type as the low-rank transformation parameter of the top-level network layer corresponding to this network layer type, so as to perform low-rank transformation on the network layers of different network layer types with different intensities.
[0105] Based on this, obtaining the preset low-rank transformation parameter can include:
[0106] Step 11: Determine the network layer types corresponding to each network layer in the pre-trained language model;
[0107] Step 12: Obtain the preset low-rank transformation parameters corresponding to each network layer type.
[0108] Based on this, for the network layers in the pre-trained language model that correspond to the same network layer type, setting the preset low-rank transformation parameter as the low-rank transformation parameter of the top-level network layer includes:
[0109] Step 21: For the network layers in the pre-trained language model that correspond to the same network layer type, set the preset low-rank transformation parameter corresponding to this network layer type as the low-rank transformation parameter of the top-level network layer corresponding to this network layer type.
[0110] Furthermore, this embodiment does not limit how to calculate the similarity between network layers. For example, the feature distance between the weight matrices of adjacent network layers can be determined and used as the similarity.
[0111] Based on this, determining the similarity between adjacent network layers can include:
[0112] Step 31: Determine the feature distance between the weight matrices of the adjacent network layers and use the feature distance as the similarity.
[0113] It should be noted that this embodiment does not limit the specific feature distance type, and can be set according to actual application requirements. For example, Euclidean distance, Chebyshev distance, Manhattan distance, Minkowski distance, Mahalanobis distance, symmetric point distance, cosine similarity, etc. can be adopted. In this embodiment, in order to effectively indicate the similarity relationship between subspaces corresponding to the intrinsic rank of the weight matrix, the Grassmann distance between the weight matrices can be calculated, and the Grassmann distance is used as the similarity between adjacent network layers. The specific calculation process of Grassmann is introduced below.
[0114] Based on this, determining the characteristic distance between the weight matrices of the adjacent network layers may include:
[0115] Step 41: Perform singular value decomposition on the first weight matrix of the upper network layer in the adjacent network layer to obtain a first non-negative diagonal matrix and a first right singular matrix, and determine the rank of the first non-negative diagonal matrix.
[0116] Specifically, the first weight matrix M1 can be subjected to singular value decomposition SVD(M1)= ,in denotes the first nonnegative diagonal matrix, and V denotes the first right singular matrix.
[0117] Step 42: Determine a first screening quantity according to the rank of the first non-negative diagonal matrix, and screen the first right singular matrix for the first singular vectors whose number is the first screening quantity.
[0118] In this step, you can The rank of the first right singular matrix is screened top i dimensional data (i.e., the first singular vector), i is The rank of the first weight matrix contains the most abundant information in the top i In a dimension.
[0119] Step 43: Perform singular value decomposition on the second weight matrix corresponding to the lower network layer in the adjacent network layer to obtain a second non-negative diagonal matrix and a second right singular matrix, and determine the rank of the second non-negative diagonal matrix.
[0120] Step 44: Determine a second screening quantity according to the rank of the second non-negative diagonal matrix, and screen the first second singular vectors whose number is the second screening quantity in the second right singular matrix.
[0121] Steps 43 and 44 are similar to steps 41 and 42.
[0122] Step 45: Determine the similarity using the first singular vectors, the rank of the first non-negative diagonal matrix, the second singular vectors, and the rank of the second non-negative diagonal matrix.
[0123] Specifically, the similarity between the weight matrices of adjacent network layers can be expressed as:
[0124] ;
[0125] where sim represents similarity. When sim = 0, the two operators (i.e., weight matrices) are orthogonal; when sim = 1, the two operators represent the same subspace, and the value range of similarity is [0, 1]. GD represents the Grassmann distance, and M1 and M2 represent two weight matrices. represents the top i dimensional data of the right singular matrix of matrix M1. If matrix M1 is subjected to singular value decomposition SVD(M1) = , then top k is equal to 's rank. The size of the rank represents the highest dimensional value where the information represented by matrix M1 lies, that is, the most abundant information in M1 is contained in the top i dimensions. Similarly is the top j dimensional data of the right singular matrix of matrix M2.
[0126] S103. Determine whether the similarity is less than a preset threshold.
[0127] As described above, in this embodiment, low-rank transformation parameters will be set for each network layer below the top network layer according to the similarity. In this step, it is possible to determine whether adjacent network layers are similar based on the relationship between the similarity and the preset threshold, and different low-rank transformation parameter setting means are adopted in the two cases of similarity and dissimilarity. It should be noted that the specific value of the preset threshold is not limited in this embodiment and can be set according to actual application requirements.
[0128] S104. If the similarity is not less than the preset threshold, then use the low-rank transformation parameter of the upper network layer in the adjacent network layers as the low-rank transformation parameter of the lower network layer in the adjacent network layers.
[0129] In this step, if it is determined that the similarity between adjacent network layers is not less than the preset threshold, it can be determined that these two network layers are relatively similar. Furthermore, the subspaces described by the weight matrices of these two network layers are relatively close, and the same low-rank transformation parameter can be used for low-rank transformation. Therefore, in this step, the low-rank transformation parameter of the upper network layer in the adjacent network layers can be directly used as the low-rank transformation parameter of the lower network layer in the adjacent network layers.
[0130] S105. If the similarity is less than the preset threshold, determine an adjustment coefficient according to the similarity, and use the adjustment coefficient to increase the low-rank transformation parameter of the upper network layer in the adjacent network layer to obtain the low-rank transformation parameter of the lower network layer in the adjacent network layer.
[0131] In this step, if it is determined that the similarity between adjacent network layers is less than the preset threshold, it can be determined that these two network layers are not similar. Furthermore, the subspaces described by the weight matrices of these two network layers are not close. If the same low-rank transformation parameter is used for processing, it is easy to cause information loss in the lower network layer. Therefore, in this step, an adjustment coefficient can be determined according to this similarity, and the adjustment coefficient is used to increase the low-rank transformation parameter of the upper network layer in the adjacent network layer to obtain the low-rank transformation parameter of the lower network layer in the adjacent network layer. After the low-rank transformation parameter is increased, the compression intensity received by the lower network layer is reduced, and thus more information can be retained, thereby avoiding the problem of information loss in the lower network layer.
[0132] It is also worth pointing out that, from the perspective of the overall model, the low-rank transformation parameter gradually increases as the network layer deepens. Thus, appropriate low-rank transformation parameters can be set for each layer in the deep network, and therefore the problem of serious information loss in the deep network easily caused by setting fixed low-rank transformation parameters in the related art can be avoided.
[0133] It should be noted that the calculation method of the adjustment coefficient is not limited in this embodiment and can be set according to actual application requirements, as long as it can ensure that the adjustment coefficient can increase the low-rank transformation parameter of the upper network layer in the adjacent network layer. The similarity can have a negative correlation with the adjustment coefficient, that is, the smaller the similarity, the larger the adjustment coefficient. In this way, it can be ensured that the greater the difference between adjacent network layers, the larger the low-rank transformation parameter obtained by the lower network layer and the more information retained. The following introduces a calculation method of the adjustment coefficient.
[0134] Based on this, determining the adjustment coefficient according to the similarity and using the adjustment coefficient to increase the low-rank transformation parameter of the upper network layer in the adjacent network layer can include:
[0135] Step 51: Determine a correlation factor according to the similarity; wherein, the similarity has a positive correlation with the correlation factor;
[0136] Step 52: Take the reciprocal of the correlation factor as the adjustment coefficient, and use the adjustment coefficient to increase the low-rank transformation parameter of the upper network layer in the adjacent network layer.
[0137] Specifically, the adjustment coefficient can be expressed as:
[0138] ;
[0139] Among them, R(L(i - 1)) represents the low-rank transformation parameter of the (i - 1)-th layer; represents the correlation factor, round() represents the rounding function, for example, round(1.5) = 2; sim i represents the similarity between the i-th network layer and the (i - 1)-th network layer. It can be seen that the similarity is positively correlated with the correlation factor. The smaller the similarity, the smaller the correlation factor. The correlation factor is negatively correlated with the adjustment coefficient. The smaller the correlation factor, the larger the adjustment coefficient. This can ensure that if the similarity is lower, the low-rank transformation parameter of the i-th network layer is larger than that of the (i - 1)-th layer.
[0140] Steps S104 and S105 can be represented by the Adaptive Function (AF) as:
[0141] ;
[0142] Among them, T represents the preset threshold; F1 is an equivalent function and can be represented as: , that is, the low-rank transformation parameter of the i-th operator is equal to that of the (i - 1)-th layer. F2 can be represented as:
[0143] .
[0144] When the similarity sim i is greater than or equal to the threshold, the adaptive selection function AF is F1. When the similarity sim i is less than the threshold, the adaptive selection function AF is F2.
[0145] Based on the above embodiments, the present invention can first obtain a pre-trained language model and preset low-rank transformation parameters. For the network layers corresponding to the same network layer type in the pre-trained language model, the preset low-rank transformation parameters can be set as the low-rank transformation parameters of the top network layer, and the similarity between adjacent network layers can be determined. Subsequently, the present invention determines whether the similarity between adjacent network layers is less than a preset threshold. If the similarity is not less than the preset threshold, it means that these two adjacent network layers are relatively similar, and the same low-rank transformation parameters can be used for low-rank transformation. Furthermore, the low-rank transformation parameters of the upper network layer in the adjacent network layers can be used as the low-rank transformation parameters of the lower network layer in the adjacent network layers. If the similarity is less than the preset threshold, it means that these two adjacent network layers are not similar, and a larger low-rank transformation parameter needs to be set for the lower network layer to retain information. Furthermore, an adjustment coefficient can be determined according to the similarity, and the low-rank transformation parameters of the upper network layer in the adjacent network layers can be increased by using the adjustment coefficient to obtain the low-rank transformation parameters of the lower network layer in the adjacent network layers. In this way, the present invention can adaptively set the low-rank transformation parameters for the lower network layer according to the similarity between adjacent network layers, and generally ensure that the deeper the network layer, the larger the low-rank transformation parameter. This can effectively ensure the low-rank transformation effect of the model and reduce the problem of information loss in the deep network, thereby improving the efficiency and performance of model fine-tuning training.
[0146] Based on the above embodiments, the uses of the low-rank transformation parameters and the fine-tuning process of the pre-trained language model are introduced below. In one possible case, the method may further include:
[0147] S201. Perform matrix decomposition on the weight matrix of the network layer according to the low-rank transformation parameters to obtain a low-rank transformation weight matrix corresponding to the network layer.
[0148] For ease of understanding, please refer to Figure 2 , Figure 2 which is a schematic diagram of a matrix decomposition provided by an embodiment of the present invention. x represents the input matrix, h represents the output matrix, W represents the weight matrix before decomposition, A and B represent the weight matrices after decomposition. The A matrix is initialized with a random Gaussian distribution (A = N(0, δ 2 ))), and the B matrix is initialized with all 0s (B = 0). In this embodiment, the weight matrix W of the network layer can be d×k decomposed into two matrices A d×r and B r×k, the two are connected by a low-rank transformation parameter r, where the rank r is much smaller than d and k. The number of training parameters is reduced from the original d*k to (d*r + r*k), greatly reducing the number of training parameters. At the same time, the number of gradient values to be saved is much less, thus saving a large amount of video memory. At the same time, since the low-rank transformation parameters corresponding to different network layers are different, generally speaking, the deeper the network layer, the larger the corresponding low-rank transformation parameter. In this way, on the basis of reducing the number of parameters, it can be ensured that the deeper network layer retains more information, thus avoiding the loss of key information.
[0149] S202. Fine-tune and train the low-rank transformation weight matrix of the network layer.
[0150] In this step, the original weight matrix of the network layer can be frozen, and only the low-rank transformation weight matrix is fine-tuned and trained, so as to achieve the effect of efficiently fine-tuning the model. The specific process of fine-tuning and training will be introduced below.
[0151] Based on this, fine-tuning and training the low-rank transformation weight matrix of the network layer may include:
[0152] S301. Obtain a preset training text.
[0153] In this step, the preset training text can come from a knowledge base in a specific field. The specific content and quantity of the preset training text are not limited in this embodiment and can be set according to actual application requirements.
[0154] S302. Use each network layer in the pre-trained language model to process the preset training text layer by layer to obtain model output data; wherein, the output of the network layer is the fusion output data obtained by fusing the first output data and the second output data, the first output data is the processing result of the weight matrix of the network layer on the preset training text, and the second output data is the processing result of the low-rank transformation weight matrix of the network layer on the preset training text.
[0155] In this step, the preset training text can be input into the pre-trained language model, and each network layer in the pre-trained language model is used to process the preset training text layer by layer to obtain model output data. It should be noted that in each network layer, the features corresponding to the preset training text need to be processed by the original weight matrix and the low-rank transformation weight matrix of the network layer respectively to obtain the first output data processed by the weight matrix and the second output data processed by the low-rank transformation weight matrix, and then the first output data and the second output data are fused to obtain the fusion output data output by this network layer.
[0156] The processing process of the preset training text in each network layer is introduced below. Based on this, each network layer in the pre-trained language model is used to process the preset training text layer by layer to obtain model output data, which may include:
[0157] Step 51: Take the first network layer in the pre-trained language model as the current network layer, take the preset training text as the data to be processed, and input the data to be processed into the current network layer.
[0158] Step 52: Process the data to be processed using the weight matrix of the current network layer to obtain the first output data.
[0159] Step 53: Process the data to be processed using the low-rank transformation weight matrix of the current network layer to obtain the second output data, and perform scaling processing on the second output data using a scaling parameter.
[0160] Step 54: Superimpose the first output data and the scaled second output data to obtain the fused output data.
[0161] In Steps 53 and 54, after obtaining the second output data, the second output data can be scaled using the scaling parameter, and the first output data and the scaled second output data are superimposed, which is equivalent to performing weighted superposition on the first output data and the second output data to obtain the fused output data.
[0162] It should be noted that the specific value of the scaling parameter is not limited in this embodiment and can be set according to actual application requirements. In addition, the scaling parameter can also be a value adaptively set for each network layer. For example, the low-rank transformation parameter of each network layer can be obtained, and the scaling parameter corresponding to the network layer can be determined using the low-rank transformation parameter and a preset scaling factor.
[0163] Based on this, before performing scaling processing on the second output data using the scaling parameter, it further includes:
[0164] Step 61: Determine the scaling parameter corresponding to the current network layer using the low-rank transformation parameter of the current network layer and a preset scaling factor;
[0165] The scaling processing of the second output data using the scaling parameter includes:
[0166] Step 62: Perform scaling processing on the second output data using the scaling parameter corresponding to the current network layer.
[0167] It can be seen that the present application can adaptively generate scaling parameters for each network layer according to the low-rank transformation parameters of each network layer, so as to ensure that the superposition process of the first output data and the second output data is also related to the network hierarchy. The superposition process of the first output data and the second output data can be expressed as:
[0168] ;
[0169] where h represents the final output result, x represents the input data, W represents the original weight matrix, Wx represents the first output data, represents the second output data, k represents the scaling factor, B and A represent the dot product of the low-rank transformation weight matrices, α represents the scaling parameter, and r represents the low-rank transformation parameter.
[0170] Step 55: Determine whether there is a lower network layer that has not processed the preset training text.
[0171] Step 56: If there is, use the fusion output data of the current network layer as the data to be processed, update the current network layer to the lower network layer of the current network layer, and enter the step of inputting the data to be processed into the current network layer.
[0172] Step 57: If not, use the fusion output data of the current network layer as the model output data.
[0173] S303: Determine a loss value based on the model output data, and update the parameters of the low-rank transformation weight matrices of each network layer according to the loss value.
[0174] In this step, a loss value can be determined based on the model output data, and the parameters of the low-rank transformation weight matrices of each network layer can be updated according to the loss value, so that the information added by the preset training text can be saved by using the parameter values of the low-rank transformation weight matrices. Furthermore, in actual inference, the general knowledge saved by the original weight matrices of the network layers and the fine-tuning knowledge saved by the low-rank transformation weight matrices of the network layers can be jointly used for text parsing and generation, thereby ensuring the performance of the model.
[0175] It should be noted that the specific determination process of the loss value is not limited in this embodiment, and relevant technologies of artificial intelligence and language models can be referred to.
[0176] Based on the above embodiments, to effectively evaluate the impact of different preset low-rank transformation parameters on the model performance, this embodiment may also preset multiple different groups of preset low-rank transformation parameters, determine low-rank transformation parameters for the network layers in the pre-trained language model based on each group of preset low-rank transformation parameters, and perform fine-tuning training on the pre-trained language model, so as to determine the model performance values corresponding to different groups of preset low-rank transformation parameters, and then select the best group of preset low-rank transformation parameters. The selection process of the preset low-rank transformation parameter groups is introduced below. In one possible case, the method may further include:
[0177] S401. Create different groups of preset low-rank transformation parameters; wherein, each group of preset low-rank transformation parameters stores the preset low-rank transformation parameters corresponding to each type of network layer, and different groups of preset low-rank transformation parameters store different preset low-rank transformation parameters.
[0178] S402. Determine different low-rank transformation parameters for the network layers in the pre-trained language model based on the different groups of preset low-rank transformation parameters.
[0179] S403. Perform fine-tuning training on each network layer in the pre-trained language model based on different low-rank transformation parameters, and determine the model performance values corresponding to different groups of preset low-rank transformation parameters.
[0180] S404. Determine the best group of preset low-rank transformation parameters according to the model performance values.
[0181] In this embodiment, multiple different model versions can be fine-tuned using different groups of preset low-rank transformation parameters, and the best group of preset low-rank transformation parameters can be selected according to the model performance values of each model version, so as to improve the model fine-tuning performance. It should be noted that in this embodiment, only the preset low-rank transformation parameters need to be determined for the top-level network layer, which can effectively reduce the manual debugging difficulty of the preset low-rank transformation parameters.
[0182] It should be noted that this embodiment does not limit how to determine the model performance values. For example, metrics such as accuracy, precision, recall, F1-score, AUC-ROC, and cross-entropy loss can be used to determine the model performance values.
[0183] Based on the above embodiments, the process of generating text according to the user input text using the fine-tuned pre-trained language model is introduced below. Please refer to Figure 3 , Figure 3 which is a flowchart of a text generation method provided by an embodiment of the present invention. In one possible case, the method may further include:
[0184] S501. Obtain the user input text and the pre-trained language model; wherein, the network layer in the pre-trained language model includes a weight matrix and a low-rank transformation weight matrix, and the low-rank transformation weight matrix is generated by using the weight matrix and the low-rank transformation parameters corresponding to the network layer, and the low-rank transformation parameters are generated according to the parameter setting method described above.
[0185] In this step, the user input text and the pre-trained language model can be obtained. Among them, the user input text contains the user's intention. The pre-trained language model has been fine-tuned, and each network layer therein includes an original weight matrix and a low-rank transformation weight matrix. The weight parameters of the weight matrix are used to store general knowledge, and the weight parameters of the low-rank transformation weight matrix are used to store domain-specific knowledge. In this way, when generating text, the general knowledge stored in the original weight matrix of the network layer and the domain-specific knowledge stored in the low-rank transformation weight matrix of the network layer can be jointly used for text parsing and generation, thereby ensuring the model performance.
[0186] S502. Use each network layer in the pre-trained language model to process the user input text layer by layer to obtain model output data; wherein, the output of the network layer is the fused output data obtained by fusing the first output data and the second output data, the first output data is the processing result of the weight matrix of the network layer on the user input text, and the second output data is the processing result of the low-rank transformation weight matrix of the network layer on the user input text.
[0187] In this step, similar to the fine-tuning training, when each network layer processes the user input text, it is necessary to first perform separate processing by the weight matrix and the low-rank transformation weight matrix in the network layer to obtain the first output data processed by the weight matrix and the second output data processed by the low-rank transformation weight matrix, and then fuse the first output data and the second output data to obtain the fused output data output by the network layer.
[0188] The processing process of the user input text in each network layer is introduced below. Based on this, using each network layer in the pre-trained language model to process the user input text layer by layer to obtain model output data may include:
[0189] Step 61: Use the first network layer in the pre-trained language model as the current network layer, use the user input text as the data to be processed, and input the data to be processed into the current network layer.
[0190] Step 62: Use the weight matrix of the current network layer to process the data to be processed to obtain the first output data.
[0191] Step 63: Process the data to be processed by using the low-rank transformation weight matrix of the current network layer to obtain the second output data, and perform scaling processing on the second output data by using a scaling parameter.
[0192] Step 64: Superimpose the first output data and the second output data after scaling processing to obtain the fused output data.
[0193] In Steps 63 and 64, after obtaining the second output data, the second output data can be scaled by using a scaling parameter, and the first output data and the second output data after scaling processing are superimposed, which is equivalent to performing weighted superposition on the first output data and the second output data to obtain the fused output data.
[0194] Furthermore, the scaling parameter can be a value adaptively set for each network layer. For example, the low-rank transformation parameter of each network layer can be obtained, and the scaling parameter corresponding to the network layer can be determined by using the low-rank transformation parameter and a preset scaling factor.
[0195] Based on this, before performing scaling processing on the second output data by using the scaling parameter, it may further include:
[0196] Step 71: Determine the scaling parameter corresponding to the current network layer by using the low-rank transformation parameter of the current network layer and a preset scaling factor;
[0197] The step of performing scaling processing on the second output data by using the scaling parameter includes:
[0198] Step 72: Perform scaling processing on the second output data by using the scaling parameter corresponding to the current network layer.
[0199] It can be seen that the present application can adaptively generate scaling parameters for each network layer according to the low-rank transformation parameters of each network layer, so as to ensure that the superimposing process of the first output data and the second output data is also related to the network level.
[0200] Step 65: Determine whether there is a lower network layer that has not processed the user input text;
[0201] Step 66: If so, use the fused output data of the current network layer as the data to be processed, update the current network layer to the lower network layer of the current network layer, and enter the step of inputting the data to be processed into the current network layer;
[0202] Step 67: If not, use the fused output data of the current network layer as the model output data.
[0203] S503: Generate text based on the model output data.
[0204] It should be noted that this embodiment does not limit how to generate text based on the model output data. For example, text sampling can be performed based on the model output data to select each word in the generated text one by one. The specific content of the language model can be referred to.
[0205] Based on the above embodiments, the above parameter setting method will be introduced below based on specific examples and schematic diagrams. This process may include:
[0206] (1) Load the pre-trained model and the fine-tuning dataset.
[0207] (2) Pre-define the low-rank approximation module. Define the linear layer Linear in the feed-forward network of the pre-trained model as the adaptive low-rank approximation module. Assume that N has a total of 24 layers.
[0208] (3) Adaptive low-rank approximation fine-tuning module. In this method, for all modules of each category, first calculate their respective low-rank transformation parameters. Decompose the pre-trained weight W ∈ into the dot product BA of two matrices, where B ∈ , A ∈ . Define the similarity threshold T = 0.5. Please refer to Figure 4 , Figure 4 which is a schematic diagram of a parameter setting process provided by an embodiment of the present invention, where i is the serial number of the linear layer. The specific operation process is as follows:
[0209] A. For the first linear layer Linear, set the low-rank transformation parameter r1 = 2 of its operator W1.
[0210] B. Calculate the similarity value between the operator W2 of the second linear layer and W1, sim1 = 0.3.
[0211] C. Since sim1 is less than the threshold T (0.5), the adaptive selection function AF = F2 is selected, so r2 = 2 / 2 (-2) = 8.
[0212] D. Use the same method to calculate the low-rank transformation parameters of 24 Linear layers in sequence, denoted as r = [2, 8, 16, 16, 32, ……, 128, 128, 256, 512]. It can be seen that the low-rank transformation parameters of the deep network are getting larger and larger, indicating that it is not suitable to use low-rank approximation and more complete information needs to be retained.
[0213] (4) On the code dataset, fine-tune the pre-trained model based on 24 low-rank transformation parameters and save it.
[0214] (5)Model application. During the actual inference process, for the input data x, the output of the original weight and the output after passing through matrices A and B are superimposed as follows to obtain the final output result:
[0215] ;
[0216] where h represents the final output result, W represents the original weight matrix, represents the low-rank transformation weight matrix, k represents the scaling factor, α represents the scaling parameter, and r represents the low-rank transformation parameter.
[0217] Next, the parameter setting device, text generation method, electronic device, computer program product, and non-volatile computer-readable storage medium provided by the embodiments of the present invention will be introduced. The parameter setting device, text generation method, electronic device, computer program product, and non-volatile computer-readable storage medium described below can be correspondingly referred to the parameter setting method and text generation method described above.
[0218] Please refer to Figure 5 , Figure 5 which is the structural block diagram of a parameter setting device provided by an embodiment of the present invention. The device may include:
[0219] An acquisition module 501, configured to acquire a pre-trained language model and a preset low-rank transformation parameter;
[0220] A network layer processing module 502, configured to set the preset low-rank transformation parameter as the low-rank transformation parameter of the top network layer for the network layers corresponding to the same network layer type in the pre-trained language model, and determine the similarity between adjacent network layers;
[0221] A first low-rank transformation parameter setting module 503, configured to, if the similarity is not less than a preset threshold, use the low-rank transformation parameter of the upper network layer in the adjacent network layer as the low-rank transformation parameter of the lower network layer in the adjacent network layer;
[0222] A second low-rank transformation parameter setting module 504, configured to, if the similarity is less than the preset threshold, determine an adjustment coefficient according to the similarity, and increase the low-rank transformation parameter of the upper network layer in the adjacent network layer by using the adjustment coefficient to obtain the low-rank transformation parameter of the lower network layer in the adjacent network layer.
[0223] Optionally, the second low-rank transformation parameter setting module 504 may include:
[0224] An association factor determination sub-module, configured to determine an association factor according to the similarity; wherein, the similarity has a positive correlation with the association factor;
[0225] An adjustment sub-module, configured to use the reciprocal of the association factor as the adjustment coefficient, and increase the low-rank transformation parameters of the upper network layer in the adjacent network layers by using the adjustment coefficient.
[0226] Optionally, the obtaining module 501 may include:
[0227] A classification sub-module, configured to determine the network layer types corresponding to each network layer in the pre-trained language model;
[0228] An obtaining sub-module, configured to obtain the preset low-rank transformation parameters corresponding to each network layer type.
[0229] Optionally, the network layer processing module 502 may include:
[0230] A parameter setting sub-module, configured to set the preset low-rank transformation parameters corresponding to the network layer type as the low-rank transformation parameters of the top network layer corresponding to the network layer type for the network layers in the pre-trained language model that correspond to the same network layer type.
[0231] Optionally, the apparatus may further include:
[0232] A creation module, configured to create different groups of preset low-rank transformation parameters; wherein, each group of preset low-rank transformation parameters stores the preset low-rank transformation parameters corresponding to each network layer type, and different groups of preset low-rank transformation parameters store different preset low-rank transformation parameters;
[0233] A parameter setting module, configured to determine different low-rank transformation parameters for the network layers in the pre-trained language model based on the different groups of preset low-rank transformation parameters;
[0234] A multi-version fine-tuning module, configured to fine-tune and train each network layer in the pre-trained language model based on different low-rank transformation parameters, and determine the model performance values corresponding to different groups of preset low-rank transformation parameters;
[0235] A selection module, configured to determine the best group of preset low-rank transformation parameters according to the model performance values.
[0236] Optionally, the network layer processing module 502 may include:
[0237] A similarity calculation sub-module, configured to determine the feature distance between the weight matrices of the adjacent network layers, and use the feature distance as the similarity.
[0238] Optionally, the similarity calculation sub-module may include:
[0239] A first decomposition unit is used to perform singular value decomposition on a first weight matrix of an upper network layer in an adjacent network layer to obtain a first non-negative diagonal matrix and a first right singular matrix, and determine the rank of the first non-negative diagonal matrix;
[0240] A first screening unit is used to determine a first screening quantity according to the rank of the first non-negative diagonal matrix, and to screen the first right singular matrix for the first singular vectors whose number is the first screening quantity;
[0241] A second decomposition unit is used to perform singular value decomposition on a second weight matrix corresponding to a lower network layer in an adjacent network layer to obtain a second non-negative diagonal matrix and a second right singular matrix, and determine the rank of the second non-negative diagonal matrix;
[0242] A second screening unit is used to determine a second screening quantity according to the rank of the second non-negative diagonal matrix, and screen the front second singular vectors whose number is the second screening quantity in the second right singular matrix;
[0243] A computing unit is used to determine the similarity using the first singular vector, the rank of the first non-negative diagonal matrix, the second singular vector, and the rank of the second non-negative diagonal matrix.
[0244] Optionally, the device may further include:
[0245] A low-rank conversion weight matrix generation module is used to perform matrix decomposition on the weight matrix of the network layer according to the low-rank conversion parameters to obtain a low-rank conversion weight matrix corresponding to the network layer;
[0246] The fine-tuning module is used to perform fine-tuning training on the low-rank transformation weight matrix of the network layer.
[0247] Optionally, the fine-tuning module may include:
[0248] The acquisition submodule is used to obtain the preset training text;
[0249] A processing submodule, used to process the preset training text layer by layer using each network layer in the pre-trained language model to obtain model output data; wherein the output of the network layer is fused output data obtained by fusion processing of the first output data and the second output data, the first output data is the processing result of the preset training text by the weight matrix of the network layer, and the second output data is the processing result of the preset training text by the low-rank transformation weight matrix of the network layer;
[0250] A parameter updating module is used to determine a loss value based on the model output data, and to update the parameters of the low-rank transformation weight matrix of each network layer according to the loss value.
[0251] Optionally, the processing sub-module may include:
[0252] A setting unit, configured to use the first network layer in the pre-trained language model as the current network layer, use the preset training text as the data to be processed, and input the data to be processed into the current network layer;
[0253] A first processing unit, configured to process the data to be processed by using the weight matrix of the current network layer to obtain the first output data;
[0254] A second processing unit, configured to process the data to be processed by using the low-rank transformation weight matrix of the current network layer to obtain the second output data, and perform scaling processing on the second output data by using a scaling parameter;
[0255] A fusion unit, configured to superimpose the first output data and the scaled second output data to obtain the fused output data;
[0256] A judgment unit, configured to judge whether there is a lower network layer that has not processed the preset training text;
[0257] A loop control unit, configured to, if so, use the fused output data of the current network layer as the data to be processed, update the current network layer to the lower network layer of the current network layer, and enter the step of inputting the data to be processed into the current network layer;
[0258] An output unit, configured to, if not, use the fused output data of the current network layer as the model output data.
[0259] Optionally, the device may further include:
[0260] A scaling parameter determination module, configured to determine the scaling parameter corresponding to the current network layer by using the low-rank transformation parameter of the current network layer and a preset scaling factor;
[0261] The second processing unit may include:
[0262] A scaling sub-unit, configured to perform scaling processing on the second output data by using the scaling parameter corresponding to the current network layer.
[0263] Please refer to Figure 6 , Figure 6 For the structural block diagram of a text generation device provided by an embodiment of the present invention, the device may include:
[0264] An acquisition module 601, configured to acquire user input text and a pre-trained language model; wherein, the weight matrix and the low-rank transformation weight matrix included in the network layer of the pre-trained language model, the low-rank transformation weight matrix is generated by using the weight matrix and the low-rank transformation parameters corresponding to the network layer, and the low-rank transformation parameters are generated according to the parameter setting method described above;
[0265] A model processing module 602, configured to perform layer-by-layer processing on the user input text by using each network layer in the pre-trained language model to obtain model output data; wherein, the output of each network layer is the fused output data obtained by performing fusion processing on the first output data and the second output data, the first output data is the processing result of the weight matrix of the network layer on the user input text, and the second output data is the processing result of the low-rank transformation weight matrix of the network layer on the user input text;
[0266] A text generation module 603, configured to generate text based on the model output data.
[0267] Optionally, the model processing module 602 may include:
[0268] A setting sub-module, configured to use the first network layer in the pre-trained language model as the current network layer, use the user input text as the data to be processed, and input the data to be processed into the current network layer;
[0269] A first processing sub-module, configured to process the data to be processed by using the weight matrix of the current network layer to obtain the first output data;
[0270] A second processing sub-module, configured to process the data to be processed by using the low-rank transformation weight matrix of the current network layer to obtain the second output data, and perform scaling processing on the second output data by using a scaling parameter;
[0271] A fusion sub-module, configured to superimpose the first output data and the second output data after scaling processing to obtain the fused output data;
[0272] A judgment sub-module, configured to judge whether there is a lower network layer that has not processed the user input text;
[0273] A loop control sub-module, configured to, if so, use the fused output data of the current network layer as the data to be processed, update the current network layer to the lower network layer of the current network layer, and enter the step of inputting the data to be processed into the current network layer;
[0274] An output sub-module, configured to, if not, use the fused output data of the current network layer as the model output data.
[0275] Optionally, the device may further include:
[0276] A scaling parameter determination module, configured to determine a scaling parameter corresponding to the current network layer by using the low-rank transformation parameter of the current network layer and a preset scaling factor;
[0277] A second processing sub-module, which may include:
[0278] A scaling unit, configured to perform scaling processing on the second output data by using the scaling parameter corresponding to the current network layer.
[0279] Please refer to Figure 7 , Figure 7 , which is a structural block diagram of an electronic device provided by an embodiment of the present invention. An embodiment of the present invention provides an electronic device 10, including a processor 11 and a memory 12; wherein, the memory 12 is used to store a computer program; the processor 11 is configured to execute the parameter setting method and the text generation method provided in the foregoing embodiment when executing the computer program.
[0280] For the specific processes of the foregoing parameter setting method and text generation method, reference may be made to the corresponding content provided in the foregoing embodiment, and details are not described herein again.
[0281] Moreover, as a carrier for resource storage, the memory 12 may be a read-only memory, a random access memory, a magnetic disk, or an optical disc, etc., and the storage method may be temporary storage or permanent storage.
[0282] In addition, the electronic device 10 further includes a power supply 13, a communication interface 14, an input / output interface 15, and a communication bus 16; wherein, the power supply 13 is used to provide a working voltage for each hardware device on the electronic device 10; the communication interface 14 can create a data transmission channel between the electronic device 10 and an external device, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present invention, and specific limitations are not imposed thereon herein; the input / output interface 15 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application requirements, and specific limitations are not imposed herein.
[0283] An embodiment of the present invention further provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the parameter setting method and the text generation method described in the foregoing embodiment are implemented.
[0284] Since the embodiments of the computer program product part correspond to the embodiments of the parameter setting method and the text generation method part, please refer to the description of the embodiments of the parameter setting method and the text generation method part for the embodiments of the computer program product part, and details are not described herein again.
[0285] An embodiment of the present invention also provides a non-volatile computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the parameter setting method and the text generation method described in the above embodiments are implemented.
[0286] Since the embodiments of the non-volatile computer-readable storage medium part correspond to the embodiments of the parameter setting method and the text generation method part, please refer to the descriptions of the embodiments of the parameter setting method and the text generation method part for the embodiments of this part, and details will not be repeated here.
[0287] The embodiments in the specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the device disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple. For the relevant parts, please refer to the description of the method part.
[0288] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0289] The steps of the method or algorithm described in combination with the embodiments disclosed in this article can be directly implemented by hardware, a software module executed by a processor, or a combination of the two. The software module can be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0290] The above has introduced in detail the parameter setting method, text generation method, device, equipment, and program product provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and modifications can still be made to the present invention, and these improvements and modifications also fall within the protection scope of the present invention.
Claims
1. A parameter setting method, characterized in that: include: Get the pre-trained language model and preset low-rank transformation parameters; For network layers corresponding to the same network layer type in the pre-trained language model, the preset low-rank transformation parameters are set as low-rank transformation parameters of the top network layer, and the similarity between adjacent network layers is determined; the similarity represents the similarity of subspaces described by weight matrices of the adjacent network layers; If the similarity is not less than a preset threshold, using the low-rank transformation parameters of the upper network layer in the adjacent network layers as the low-rank transformation parameters of the lower network layer in the adjacent network layers; If the similarity is less than the preset threshold, determining an adjustment coefficient according to the similarity, and using the adjustment coefficient to increase a low-rank transformation parameter of an upper network layer in the adjacent network layer to obtain a low-rank transformation parameter of a lower network layer in the adjacent network layer; Determining an adjustment coefficient according to the similarity, and increasing a low-rank transformation parameter of an upper network layer in the adjacent network layer by using the adjustment coefficient, including: Determining a correlation factor according to the similarity; wherein the similarity is positively correlated with the correlation factor; The reciprocal of the correlation factor is used as the adjustment coefficient, and the low-rank transformation parameter of the upper network layer in the adjacent network layer is increased by using the adjustment coefficient.
2. The parameter setting method according to claim 1, characterized in that: Get the preset low-rank transformation parameters, including: Determine the network layer type corresponding to each network layer in the pre-trained language model; Obtain preset low-rank transformation parameters corresponding to each of the network layer types.
3. The parameter setting method according to claim 2, characterized in that: For the network layers corresponding to the same network layer type in the pre-trained language model, setting the preset low-rank conversion parameter as the low-rank conversion parameter of the top network layer includes: For network layers corresponding to the same network layer type in the pre-trained language model, the preset low-rank transformation parameters corresponding to the network layer type are set as low-rank transformation parameters of the top network layer corresponding to the network layer type.
4. The parameter setting method according to claim 2, characterized in that: Also includes: Creating different preset low-rank transformation parameter groups; wherein the preset low-rank transformation parameter groups store the preset low-rank transformation parameters corresponding to each of the network layer types, and different preset low-rank transformation parameter groups store different preset low-rank transformation parameters; Determining different low-rank transformation parameters for the network layer in the pre-trained language model based on the different preset low-rank transformation parameter groups; Fine-tune each network layer in the pre-trained language model based on different low-rank transformation parameters to determine model performance values corresponding to different preset low-rank transformation parameter groups; An optimal preset low-rank transformation parameter set is determined based on the model performance value.
5. The parameter setting method according to claim 1, characterized in that: Determine the similarity between adjacent network layers, including: Determine the characteristic distance between the weight matrices of the adjacent network layers, and use the characteristic distance as the similarity.
6. The parameter setting method according to claim 5, characterized in that: Determining the characteristic distance between the weight matrices of the adjacent network layers includes: Performing singular value decomposition on a first weight matrix of an upper network layer in an adjacent network layer to obtain a first non-negative diagonal matrix and a first right singular matrix, and determining the rank of the first non-negative diagonal matrix; Determine a first screening quantity according to the rank of the first non-negative diagonal matrix, and screen the first right singular matrix for the first singular vectors whose number is the first screening quantity; Performing singular value decomposition on a second weight matrix corresponding to a lower network layer in an adjacent network layer to obtain a second non-negative diagonal matrix and a second right singular matrix, and determining the rank of the second non-negative diagonal matrix; Determine a second screening number according to the rank of the second non-negative diagonal matrix, and screen the first second singular vectors whose number is the second screening number in the second right singular matrix; The similarity is determined using the first singular vectors, the rank of the first non-negative diagonal matrix, the second singular vectors, and the rank of the second non-negative diagonal matrix.
7. The parameter setting method according to claim 1, characterized in that: The pre-trained language model is composed of multiple transformer network layers connected in series.
8. The parameter setting method according to any one of claims 1 to 7, characterized in that: Also includes: Performing matrix decomposition on the weight matrix of the network layer according to the low-rank transformation parameter to obtain a low-rank transformation weight matrix corresponding to the network layer; The low-rank transformation weight matrix of the network layer is fine-tuned and trained.
9. The parameter setting method according to claim 8, characterized in that: Fine-tuning the low-rank transformation weight matrix of the network layer includes: Get the preset training text; The preset training text is processed layer by layer using each network layer in the pre-trained language model to obtain model output data; wherein the output of the network layer is fused output data obtained by fusion processing of the first output data and the second output data, the first output data is the processing result of the preset training text by the weight matrix of the network layer, and the second output data is the processing result of the preset training text by the low-rank transformation weight matrix of the network layer; A loss value is determined based on the model output data, and parameters of the low-rank transformation weight matrix of each network layer are updated according to the loss value.
10. The parameter setting method according to claim 9, characterized in that: The preset training text is processed layer by layer using each network layer in the pre-trained language model to obtain model output data, including: Using the first network layer in the pre-trained language model as the current network layer, using the preset training text as the data to be processed, and inputting the data to be processed into the current network layer; Processing the data to be processed using the weight matrix of the current network layer to obtain the first output data; Using the low-rank transformation weight matrix of the current network layer to process the data to be processed to obtain the second output data, and using the scaling parameter to scale the second output data; Superimposing the first output data with the second output data after scaling processing to obtain the fused output data; Determining whether there is a lower network layer that has not processed the preset training text; If yes, the fused output data of the current network layer is used as the data to be processed, the current network layer is updated to the lower network layer of the current network layer, and the step of inputting the data to be processed into the current network layer is entered; If not, the fused output data of the current network layer is used as the model output data.
11. The parameter setting method according to claim 10, characterized in that: Before scaling the second output data using the scaling parameter, the method further includes: Determine a scaling parameter corresponding to the current network layer by using a low-rank transformation parameter of the current network layer and a preset scaling factor; The scaling process of the second output data by using the scaling parameter includes: The second output data is scaled using a scaling parameter corresponding to the current network layer.
12. A text generation method, characterized in that: include: Obtain user input text and a pre-trained language model; wherein the network layer in the pre-trained language model comprises a weight matrix and a low-rank conversion weight matrix, the low-rank conversion weight matrix is generated using the weight matrix and the low-rank conversion parameters corresponding to the network layer, and the low-rank conversion parameters are generated according to the parameter setting method according to any one of claims 1 to 11; Using each network layer in the pre-trained language model to process the user input text layer by layer to obtain model output data; wherein the output of the network layer is fused output data obtained by fusion processing of the first output data and the second output data, the first output data is the result of the weight matrix of the network layer processing the user input text, and the second output data is the result of the low-rank transformation weight matrix of the network layer processing the user input text; Text generation is performed based on the model output data.
13. The text generation method according to claim 12, characterized in that: The user input text is processed layer by layer using each network layer in the pre-trained language model to obtain model output data, including: Using the first network layer in the pre-trained language model as the current network layer, using the user input text as data to be processed, and inputting the data to be processed into the current network layer; Processing the data to be processed using the weight matrix of the current network layer to obtain the first output data; Using the low-rank transformation weight matrix of the current network layer to process the data to be processed to obtain the second output data, and using the scaling parameter to scale the second output data; Superimposing the first output data with the second output data after scaling processing to obtain the fused output data; Determining whether there is a lower network layer that has not processed the user input text; If yes, the fused output data of the current network layer is used as the data to be processed, the current network layer is updated to the lower network layer of the current network layer, and the step of inputting the data to be processed into the current network layer is entered; If not, the fused output data of the current network layer is used as the model output data.
14. The text generation method according to claim 13, characterized in that: Before scaling the second output data using the scaling parameter, the method further includes: Determine a scaling parameter corresponding to the current network layer by using a low-rank transformation parameter of the current network layer and a preset scaling factor; The scaling process of the second output data by using the scaling parameter includes: The second output data is scaled using a scaling parameter corresponding to the current network layer.
15. A parameter setting device, characterized in that: include: An acquisition module is used to obtain a pre-trained language model and preset low-rank conversion parameters; A network layer processing module, for setting the preset low-rank transformation parameters as low-rank transformation parameters of the top network layer for network layers corresponding to the same network layer type in the pre-trained language model, and determining the similarity between adjacent network layers; the similarity represents the similarity of subspaces described by weight matrices of the adjacent network layers; A first low-rank conversion parameter setting module, configured to use the low-rank conversion parameters of the upper network layer in the adjacent network layers as the low-rank conversion parameters of the lower network layer in the adjacent network layers if the similarity is not less than a preset threshold; A second low-rank conversion parameter setting module is used to determine an adjustment coefficient according to the similarity if the similarity is less than the preset threshold, and use the adjustment coefficient to increase the low-rank conversion parameter of the upper network layer in the adjacent network layer to obtain the low-rank conversion parameter of the lower network layer in the adjacent network layer; The second low-rank conversion parameter setting module includes: A correlation factor determination submodule, used to determine the correlation factor according to the similarity; wherein the similarity is positively correlated with the correlation factor; The adjustment submodule is used to use the reciprocal of the correlation factor as the adjustment coefficient, and use the adjustment coefficient to increase the low-rank conversion parameter of the upper network layer in the adjacent network layer.
16. A text generation device, characterized in that: include: An acquisition module, used to acquire user input text and a pre-trained language model; wherein the network layer in the pre-trained language model comprises a weight matrix and a low-rank conversion weight matrix, wherein the low-rank conversion weight matrix is generated using the weight matrix and the low-rank conversion parameters corresponding to the network layer, and the low-rank conversion parameters are generated according to the parameter setting method according to any one of claims 1 to 11; A model processing module, used to process the user input text layer by layer using each network layer in the pre-trained language model to obtain model output data; wherein the output of each network layer is fused output data obtained by fusion processing of first output data and second output data, the first output data is the processing result of the user input text by the weight matrix of the network layer, and the second output data is the processing result of the user input text by the low-rank transformation weight matrix of the network layer; A text generation module is used to generate text based on the model output data.
17. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the parameter setting method as described in any one of claims 1 to 11 or the text generation method as described in any one of claims 12 to 14 when executing the computer program.
18. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the parameter setting method according to any one of claims 1 to 11 or the text generation method according to any one of claims 12 to 14 is implemented.
19. A non-volatile computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are loaded and executed by the processor, the parameter setting method according to any one of claims 1 to 11 or the text generation method according to any one of claims 12 to 14 is implemented.