Parameter optimization method and device and computer readable medium
The parameter optimization method using low-rank matrices with sparse constraints addresses inefficiencies in LoRA by ensuring selective and precise parameter updates, improving training stability and efficiency while maintaining domain-specific knowledge.
Patent Information
- Application Number
- CN202510381468.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-15
AI Technical Summary
The existing LoRA algorithm has a singularization in feature representation, resulting in a coarse granularity of parameter updates, huge computing resources consumption, difficulty in accurately editing and retaining knowledge in specific tasks, and easy to cause knowledge conflicts and forgetting.
By introducing sparseness constraints, redefining the loss function, combining the gradient descent algorithm iteratively updates the low-rank matrices A and B, and pruning is performed based on sensitivity, optimizing the pre-trained weight matrix W to achieve sparseness and refined parameter updates.
Improve the stability and efficiency of model training, avoid knowledge conflicts and forgetting, ensure the integrity of knowledge in specific domains, and improve model performance.
Smart Images

Figure CN120317293A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information technology, and in particular, to a parameter optimization method, device, and computer-readable medium that combines LoRA and sparsity constraints. Background Art
[0002] With the development of large-scale pre-trained language models, these models have become the core of NLP (Natural Language Processing) tasks and are widely used in fields such as natural language understanding and generation. Although these models show strong generalization ability in general tasks, fine-tuning for specific tasks usually involves updating all model parameters, consuming huge computing resources. PEFT (Parameter-Efficient Fine-tuning) methods, such as the LoRA (Low-Rank Adaptation) algorithm, etc., by introducing low-rank matrices, without modifying the model architecture, only train a small number of parameters, thus greatly reducing the computational overhead.
[0003] However, there are problems of single feature representation and coarse control granularity of parameter updates in existing LoRA algorithms. Specifically, traditional LoRA algorithms mainly rely on the existing general language representations in the pre-trained model and lack fine-grained feature representations for specific tasks. In tasks that require knowledge editing or domain fine-tuning, such general representations may lead to problems such as knowledge conflicts or insufficient context understanding, thus affecting model performance. Since this algorithm updates all pre-trained parameters through dense matrices, it is impossible to selectively modify specific internal knowledge. This will cause all parameters to be updated during training, resulting in unnecessary computational overhead, thus reducing computational efficiency. And it is difficult to only update the knowledge of a specific domain, or the original knowledge is easily over-modified or forgotten, resulting in inaccurate knowledge editing and poor knowledge retention. Summary of the Invention
[0004] An object of this application is to provide a parameter optimization method, device, and computer-readable medium to solve the problems of low computational efficiency and poor performance of the optimized model in the existing solutions.
[0005] To achieve the above object, an embodiment of this application provides a parameter optimization method, and the method includes:
[0006] Obtain the initial weight matrix W0 to be fine-tuned, and initialize the first low-rank matrix A and the second low-rank matrix B, and determine to obtain the pre-trained weight matrix W = W0 + B·A, where, W0 ∈ R d1×d2 , A ∈ R r×d2 , B ∈ R d1×r , and r << min{d1, d2};
[0007] Introduce a sparsity constraint into the loss function and redefine the loss function. Among them, the optimization objective corresponding to the loss function is to minimize the loss function by updating the first low-rank matrix A and the second low-rank matrix B while keeping W0 frozen and under the sparsity constraint. The sparsity constraint is that the proportion of non-zero elements in the matrix B·A is less than or equal to the sparsity rate.
[0008] Based on the loss function, use the gradient descent algorithm to iteratively update the first low-rank matrix A and the second low-rank matrix B. During the iterative update process, calculate the sensitivity of each element in the pre-trained weight matrix W to the loss function, and perform pruning on the elements in the first low-rank matrix A and the second low-rank matrix B based on the sensitivity.
[0009] When the optimization termination condition is met, stop the iterative update, and obtain the optimized pre-trained weight matrix W according to the current first low-rank matrix A and the second low-rank matrix B.
[0010] Furthermore, the redefined loss function L is:
[0011]
[0012] Among them, D is the training data set used for iterative update. is the sparsity constraint, τ is the sparsity rate, indicating that the proportion of non-zero elements in the matrix B·A is less than or equal to the sparsity rate.
[0013] Furthermore, the sparsity constraint includes the row sparsity constraint on the row elements of the first low-rank matrix A and the column sparsity constraint on the column elements of the second low-rank matrix B. The redefined loss function L is:
[0014]
[0015] A i* represents the elements of the i-th row of the first low-rank matrix A, B *i is the element of the i-th column of the second low-rank matrix B. The row sparsity constraint is that the proportion of non-zero elements in the i-th row of the first low-rank matrix A is less than or equal to the sparsity rate, and the column sparsity constraint is that the proportion of non-zero elements in the i-th column of the second low-rank matrix B is less than or equal to the sparsity rate.
[0016] Furthermore, when calculating the sensitivity of each element in the pre-trained weight matrix W to the loss function, the following calculation formula is used:
[0017]
[0018] Among them, I(W ij) is the sensitivity of the element in the \(i\)-th row and \(j\)-th column of the pre-trained weight matrix \(W\) to the loss function, \(W\). ij is the weight value of the element in the \(i\)-th row and \(j\)-th column of the pre-trained weight matrix \(W\). is the gradient of the element in the \(i\)-th row and \(j\)-th column of the pre-trained weight matrix \(W\) with respect to the loss function.
[0019] Further, the method further includes:
[0020] According to the sensitivity calculated in the previous iteration update process, smooth the sensitivity of each element in the pre-trained weight matrix \(W\) obtained in this calculation to the loss function.
[0021] Further, pruning the elements in the first low-rank matrix \(A\) and the second low-rank matrix \(B\) based on the sensitivity includes:
[0022] For each row of elements in the first low-rank matrix \(A\), sort them according to the sensitivity respectively;
[0023] In each row of the first low-rank matrix \(A\), retain the elements with a higher sensitivity in a preset proportion according to the sorting result, and set the other elements to zero, where the preset proportion is determined according to the sparsity rate;
[0024] For each column of elements in the second low-rank matrix \(B\), sort them according to the sensitivity respectively;
[0025] In each column of the second low-rank matrix \(B\), retain the elements with a higher sensitivity in a preset proportion according to the sorting result, and set the other elements to zero.
[0026] Further, the method further includes:
[0027] During the iterative update process, dynamically adjust the sparsity rate, and the specific method is as follows:
[0028]
[0029] where \(\tau\) (t) represents the sparsity rate of the \(t\)-th iteration, \(t\) i and \(t\) f represent the starting iteration number and the ending iteration number of the dynamic adjustment stage respectively, \(T\) represents the total number of iterations, and \(\tau\) represents the preset sparsity rate threshold.
[0030] Further, the optimization termination condition includes that the loss function converges or reaches the total number of iterations.
[0031] Some embodiments of the present application also provide a computing device, where the device includes a memory for storing computer program instructions and a processor for executing the computer program instructions. When the computer program instructions are executed by the processor, the device is triggered to execute the foregoing parameter optimization method.
[0032] Some other embodiments of the present application also provide a computer-readable medium, on which computer program instructions are stored. The computer program instructions can be executed by a processor to implement the parameter optimization method.
[0033] Compared with the prior art, the embodiments of the present application provide a parameter optimization scheme. The scheme first obtains an initial weight matrix W0 to be fine-tuned, initializes a first low-rank matrix A and a second low-rank matrix B, and determines to obtain a pre-trained weight matrix W = W0 + B·A. Then, a sparsity constraint is introduced into the loss function, the loss function is redefined, and the optimization objective corresponding to the loss function is defined as minimizing the loss function by updating the first low-rank matrix A and the second low-rank matrix B while keeping W0 frozen and under the sparsity constraint; based on the loss function, the gradient descent algorithm is used to iteratively update the first low-rank matrix A and the second low-rank matrix B. During the iterative update process, the sensitivity of each element in the pre-trained weight matrix W to the loss function is calculated, and the elements in the first low-rank matrix A and the second low-rank matrix B are pruned based on the sensitivity; when the optimization termination condition is met, the iterative update is stopped, and according to the current first low-rank matrix A and the second low-rank matrix B, the optimized pre-trained weight matrix W is obtained. Thus, by combining the low-rank matrix with the sparsity constraint, the low-rank matrix is pruned based on the sensitivity during the iterative update process, thereby realizing the sparsification of the matrix, retaining important parameters with higher sensitivity, making the parameter update in the pre-trained weight matrix selective and refined, improving the stability and efficiency of training, and ensuring the integrity of specific domain knowledge, effectively avoiding knowledge conflicts and forgetting, and making the performance of the trained model better. Description of the Drawings
[0034] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objectives, and advantages of the present application will become more apparent:
[0035] Figure 1 It is a processing flow chart of a parameter optimization method provided by an embodiment of the present application;
[0036] The same or similar reference numerals in the drawings represent the same or similar components. Detailed Embodiments
[0037] The present application will be further described in detail below with reference to the drawings.
[0038] In a typical configuration of the present application, both the terminal and the devices of the service network include one or more processors (CPUs), input / output interfaces, network interfaces, and memories.
[0039] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash RAM. The memory is an example of computer-readable media.
[0040] Computer-readable media includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer program instructions, data structures, program devices, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device.
[0041] The embodiment of the present application provides a parameter optimization method. The method first obtains an initial weight matrix W0 to be fine-tuned, initializes a first low-rank matrix A and a second low-rank matrix B, and determines to obtain a pre-trained weight matrix W = W0 + B·A. Then, a sparsity constraint is introduced into the loss function, the loss function is redefined, and the optimization objective corresponding to the loss function is defined as minimizing the loss function by updating the first low-rank matrix A and the second low-rank matrix B while keeping W0 frozen and under the sparsity constraint; based on the loss function, the gradient descent algorithm is used to iteratively update the first low-rank matrix A and the second low-rank matrix B. During the iterative update process, the sensitivity of each element in the pre-trained weight matrix W to the loss function is calculated, and the elements in the first low-rank matrix A and the second low-rank matrix B are pruned based on the sensitivity; when the optimization termination condition is met, the iterative update is stopped, and according to the current first low-rank matrix A and the second low-rank matrix B, the optimized pre-trained weight matrix W is obtained. Thus, in this solution, by combining the low-rank matrix with the sparsity constraint, the low-rank matrix is pruned based on the sensitivity during the iterative update process, thereby realizing the sparsification of the matrix, retaining important parameters with higher sensitivity, making the parameter update in the pre-trained weight matrix selective and refined, improving the stability and efficiency of training, and ensuring the integrity of specific domain knowledge, effectively avoiding knowledge conflicts and forgetting, and making the performance of the trained model better.
[0042] In an actual scenario, the execution subject of this method can be a user device, a network device, or a device formed by integrating the user device and the network device through a network, or it can also be an application program running on the above devices. The user device includes but is not limited to various terminal devices such as computers, mobile phones, and tablet computers; the network device includes but is not limited to being implemented by a network host, a single network server, a set of multiple network servers, or a computer set based on cloud computing. Here, the cloud consists of a large number of hosts or network servers based on cloud computing (Cloud Computing), where cloud computing is a type of distributed computing and consists of a virtual computer formed by a group of loosely coupled computer sets.
[0043] Figure 1 The processing flow of a parameter optimization method provided by the embodiment of the present application is shown, which at least includes the following steps:
[0044] Step S101, obtain an initial weight matrix W0 to be fine-tuned, initialize a first low-rank matrix A and a second low-rank matrix B, and determine to obtain a pre-trained weight matrix W = W0 + B·A.
[0045] The initial weight matrix W0 refers to the pre-trained weight matrix in the initial state of this solution. This initial weight matrix is in a frozen state during the optimization process of the solution of this application, that is, the parameters in it are not adjusted, but the optimization of the pre-trained weight matrix W is achieved by adjusting the parameters in the first low-rank matrix A and the second low-rank matrix B. The numerical values of the elements in the above matrices are the parameters that need to be updated and optimized in this solution, that is, the weight values of the pre-trained weight matrix.
[0046] where, W0 ∈ R d1×d2 , A ∈ R r×d2 , B ∈ R d1×r , and r << min{d1, d2}. Specifically, the first low-rank matrix A satisfies A ∈ R r×d2 , which means that the elements in the first low-rank matrix A belong to the real number field, that is, the value of each element in the matrix is a real number, and the number of rows of the first low-rank matrix A is r, and the number of columns is d2. The second low-rank matrix B satisfies B ∈ R d1×r , which means that the elements in the second low-rank matrix B belong to the real number field, that is, the value of each element in the matrix is a real number, and the number of rows of the second low-rank matrix B is d1, and the number of columns is r. Thus, B·A ∈ R d1×d2 , which is consistent with the dimension of W0, so the result W of adding the two also satisfies W ∈ R d1×d2 . In addition, r << min{d1, d2} means that r is much smaller than d1 and d2, that is, the ranks of matrix A and matrix B are much smaller than the initial weight matrix W0.
[0047] During the training and optimization process, W0 is frozen. By iteratively updating the parameters in the two low-rank matrices A and B, the computational amount of training can be effectively reduced, the computational load can be reduced, and the processing efficiency can be improved. For example, assume that the initial weight matrix is W0 ∈ R 512 ×1024 , and its number of parameters is 512 × 1024 = 524288. If the solution of this application is adopted and the rank r of the first low-rank matrix A and the second low-rank matrix B is set to 8, then A ∈ R 8×1024 , B ∈ R 512×8 , then the number of parameters that need to be trained and optimized is 8192 + 4096 = 12288, which is much smaller than the number of parameters 524288 of the initial weight matrix.
[0048] Step S102, introduce sparsity constraints into the loss function and redefine the loss function.
[0049] In an actual scenario, the optimization objective corresponding to the loss function in the conventional LoRA algorithm is: while keeping W0 frozen, minimize the loss function by updating the first low-rank matrix A and the second low-rank matrix B. In the solution of the present application, on the basis of the conventional LoRA algorithm, a sparsity constraint is further introduced. The specific content of this sparsity constraint is: the proportion of non-zero elements in the matrix B·A is less than or equal to the sparsity rate. Thus, after introducing the sparsity constraint, the optimization objective of the loss function can be redefined as: while keeping W0 frozen and under the sparsity constraint, minimize the loss function by updating the first low-rank matrix A and the second low-rank matrix B.
[0050] Specifically, the redefined loss function L is expressed in the following way:
[0051]
[0052] where D is the training dataset used for iterative update, and the content after s.t. is the sparsity constraint, and τ is the sparsity rate, indicating that the proportion of non-zero elements in the matrix B·A is less than or equal to the sparsity rate.
[0053] Since the above loss function involves l0 optimization in the process of iterative update, and since the l0 optimization problem is NP-Hard (Nondeterministic Polynomial Hard), that is, it is difficult to find an exact solution in polynomial time, which may lead to a reduction in the computational efficiency and accuracy of the present solution. Therefore, to solve the l0 optimization problem in the calculation process, the solution of the embodiment of the present application can adopt a row and column sparsity constraint form when introducing the sparsity constraint, that is, by respectively imposing sparsity constraints on the rows of the first low-rank matrix A and the columns of the second low-rank matrix B, so as to more accurately control the position and quantity of parameter update, and improve the computational efficiency and accuracy.
[0054] Specifically, when adopting the row and column sparsity constraint form, the sparsity constraint may include a row sparsity constraint on the row elements of the first low-rank matrix A and a column sparsity constraint on the column elements of the second low-rank matrix B. At this time, the redefined loss function can be converted into the following content:
[0055]
[0056] where A i* represents the elements of the i-th row of the first low-rank matrix A, B *i is the elements of the i-th column of the second low-rank matrix B, is the current sparsity constraint, including a row sparsity constraint and a column sparsity constraint. Among them, the specific content of the row sparsity constraint is It is indicated that the proportion of non-zero elements in the $i$-th row of the first low-rank matrix $A$ is less than or equal to the sparsity rate. The specific content of the column sparsity constraint is It is indicated that the proportion of non-zero elements in the $i$-th column of the second low-rank matrix $B$ is less than or equal to the sparsity rate.
[0057] Step S103: Based on the loss function, use the gradient descent algorithm to iteratively update the first low-rank matrix $A$ and the second low-rank matrix $B$. During the iterative update process, calculate the sensitivity of each element in the pre-trained weight matrix $W$ to the loss function, and perform pruning processing on the elements in the first low-rank matrix $A$ and the second low-rank matrix $B$ based on the sensitivity.
[0058] In the way of iterative update, under the condition of imposing the sparsity constraint, adjust the elements in the first low-rank matrix $A$ and the second low-rank matrix $B$, so as to realize the parameter optimization of the pre-trained weight matrix $W$. The iterative update process can include two main steps: using the gradient descent algorithm to iteratively update the first low-rank matrix $A$ and the second low-rank matrix $B$, and performing pruning processing on the elements in the first low-rank matrix $A$ and the second low-rank matrix $B$ based on the sensitivity.
[0059] In some embodiments of the present application, the SGD (Stochastic Gradient Descent) algorithm can be used to update the matrices $A$ and $B$. The specific method can be expressed as follows:
[0060]
[0061] where $A$ (t) and $B$ (t) are the current values of the first low-rank matrix $A$ and the second low-rank matrix $B$ at the $t$-th iteration, and are the values of the first low-rank matrix $A$ and the second low-rank matrix $B$ after the $t$-th iteration update. $\eta$ is the learning rate, which is used to control the update step size during the gradient update process. and are the gradients of the loss function $L$ with respect to the first low-rank matrix $A$ and the second low-rank matrix $B$.
[0062] During the iterative update process, after each update of the first low-rank matrix $A$ and the second low-rank matrix $B$ is completed, the sensitivity of each element in the pre-trained weight matrix $W$ to the loss function can be calculated, and the elements in the first low-rank matrix $A$ and the second low-rank matrix $B$ can be pruned based on the sensitivity. Among them, the sensitivity is used to represent the importance of the parameter to the loss function. The higher the sensitivity of a certain parameter, the greater its impact on the loss function.
[0063] Specifically, when calculating the sensitivity of each element in the pre-trained weight matrix W to the loss function, the following calculation formula is adopted:
[0064]
[0065] where I(W ij ) is the sensitivity of the element in the i-th row and j-th column of the pre-trained weight matrix W to the loss function, and W ij is the weight value of the element in the i-th row and j-th column of the pre-trained weight matrix W, is the gradient of the element in the i-th row and j-th column of the pre-trained weight matrix W with respect to the loss function. The larger the absolute value of the product of the weight value and the gradient value, the greater the influence of the parameter on the loss function and the higher its importance.
[0066] In an actual scenario, since the sensitivity calculated during a single iteration update may have large fluctuations, it may lead to a decrease in the stability of subsequent pruning processing. Therefore, in order to improve the stability of the pruning process, when calculating the sensitivity, the sensitivity can be smoothed so that the calculated sensitivity is smoother and the decrease in pruning stability caused by excessive fluctuations is avoided. Specifically, in the solution provided in this embodiment of the application, the sensitivity of each element in the pre-trained weight matrix W obtained in this calculation to the loss function can be smoothed according to the sensitivity obtained during the previous iteration update process.
[0067] In this embodiment, when performing the smoothing process, the EMA (Exponential Moving Average) algorithm can be adopted, and its specific calculation formula is as follows:
[0068]
[0069] where represents the finally obtained sensitivity after smoothing at the t-th iteration, represents the sensitivity obtained during the (t - 1)-th iteration (i.e., the previous iteration), I (t) (W ij ) represents the sensitivity that has not been smoothed at the t-th iteration, and β is a smoothing coefficient used to determine and the weights during the smoothing process, thereby controlling the smoothing degree of the exponential moving average. Usually, it can be set to a value close to 1. For example, in this embodiment, it can be set to 0.9, etc.
[0070] In the process of iterative update, after calculating the sensitivity in the above manner, the elements in the first low-rank matrix A and the second low-rank matrix B can be pruned based on the sensitivity. In this solution, the principle of pruning is to retain the parameters in the matrix that have a greater impact on the loss function (i.e., the elements with higher sensitivity in the matrix), and set the remaining parameters (the elements with lower sensitivity in the matrix) to zero. The retention ratio of the parameters can be determined according to the sparsity rate.
[0071] Thus, in some embodiments of the present application, when pruning the elements in the first low-rank matrix A and the second low-rank matrix B based on the sensitivity, the rows of the first low-rank matrix A and the columns of the second low-rank matrix B can be pruned respectively, so as to achieve row and column sparsification. Specifically, for each row of elements in the first low-rank matrix A, they can be sorted according to the sensitivity respectively, and then in each row of the first low-rank matrix A, a preset ratio of elements with higher sensitivity are retained according to the sorting result, and the other elements are set to zero, thereby achieving the pruning process of the rows of the first low-rank matrix A. At the same time, for each column of elements in the second low-rank matrix B, they can be sorted according to the sensitivity respectively, and in each column of the second low-rank matrix B, a preset ratio of elements with higher sensitivity are retained according to the sorting result, and the other elements are set to zero, thereby achieving the pruning process of the rows of the second low-rank matrix B.
[0072] Among them, the preset ratio is determined according to the sparsity rate. For example, when the sparsity rate is 10%, the elements in the first 10% of the sensitivity sorting in each row of the first low-rank matrix A and the elements in the first 10% of the sensitivity sorting in each column of the second low-rank matrix B are retained according to the sorting result, and the elements in the matrix area are set to zero, thereby realizing the row and column sparsification of the matrices A and B, and finally controlling the sparsity of the matrix B·A.
[0073] In an actual scenario, pruning the rows of the first low-rank matrix A and the columns of the second low-rank matrix B can be expressed by the following formula:
[0074]
[0075] Among them, t represents the current iteration number, represents the i-th row of the first low-rank matrix A at the (t + 1)-th iteration, represents when the condition I (t) (A ij ) is in Top-τ (t) , the value of the element in the i-th row and j-th column of the first low-rank matrix A updated by gradient descent after the t-th iteration, I (r) (A ij ) is in Top-τ (t)It means that the sensitivity ranking corresponding to the element in the \(i\)-th row and \(j\)-th column of the first low-rank matrix \(A\) after \(t\) iterations is in the top \(\tau\) of that row. In other cases, that is, when the sensitivity ranking corresponding to the element in the \(i\)-th row and \(j\)-th column of the first low-rank matrix \(A\) after \(t\) iterations is not in the top \(\tau\) of that column, the element in the \(i\)-th row and \(j\)-th column of the first low-rank matrix \(A\) is set to zero, thus completing the pruning. For example, when the sparsity rate \(\tau\) is 10%, if the ranking of the sensitivity \(I\) (t) (A 23 ) is in the top 5% of that row, lower than the proportion corresponding to the current sparsity rate \(\tau\), the value of the element in the 2nd row and 3rd column obtained by gradient descent update after the \(t\)-th iteration can be retained as the value for the next iteration. Otherwise, if the ranking of the sensitivity \(I\) of the element in the 2nd row and 3rd column of the first low-rank matrix \(A\) after \(t\) iterations (t) (A 23 ) is not in the top 10% of that row, for example, after the top 20%, then the corresponding element in the first low-rank matrix \(A\) will be set to zero as the value for the next iteration.
[0076] Similarly, for the second low-rank matrix \(B\), it represents the \(i\)-th column of the second low-rank matrix \(B\) at the \((t + 1)\)-th iteration, it represents that when the condition \(I\) (t) (B ji ) is in Top-\(\tau\) (t) , the value of the element in the \(j\)-th row and \(i\)-th column of the second low-rank matrix \(B\) obtained by gradient descent update after the \(t\)-th iteration, \(I\) (t) (B ji ) in Top-\(\tau\) (t) represents that the sensitivity ranking corresponding to the element in the \(j\)-th row and \(i\)-th column of the second low-rank matrix \(B\) after \(t\) iterations is in the top \(\tau\) of that column.
[0077] In some embodiments of the present application, a dynamic adjustment mechanism for the sparsity rate is also introduced to dynamically adjust the sparsity rate during the iterative update process, so as to control the pruning progress through a preset threshold, ensure the training stability, and optimize the expression ability of the model. Specifically, the scheme of this embodiment adopts a cubic function strategy to dynamically adjust the sparsity rate \(\tau\) (t) , where \(\tau\) (t) represents the sparsity rate determined by the \(t\)-th iteration and can be expressed by the following formula:
[0078]
[0079] where \(\tau\) (t) represents the sparsity rate of the \(t\)-th iteration, \(t\) i and \(t\) fdenote the starting iteration number and the ending iteration number of the dynamic adjustment phase respectively, T denotes the total iteration number, and τ0 denotes the preset sparsity rate threshold.
[0080] It can be seen from this that in the initial stage of this dynamic adjustment process (the iteration number t is less than or equal to t i ), the sparsity rate is always 1, that is, pruning is not performed in this stage. When entering the dynamic adjustment phase (the iteration number t is greater than t i , and less than or equal to t f ), the sparsity rate τ (t) will change with the change of the iteration number. When entering the final stage (the iteration number t is greater than t f , and less than the total iteration number T), the sparsity rate τ (t) is fixed to the preset sparsity rate threshold.
[0081] Among them, the learning rate η, the sparsity rate threshold τ0, the smoothing coefficient β, and the starting iteration number and the ending iteration number t i and t f in this scheme are all pre-set hyperparameters, and appropriate values can be set according to the needs of the actual application scenario.
[0082] Step S104, when the optimization termination condition is met, stop the iterative update, and obtain the optimized pre-trained weight matrix W according to the current first low-rank matrix A and second low-rank matrix B.
[0083] Among them, the optimization termination condition may include that the loss function converges or reaches the total iteration number. Thus, during the iterative update, if the loss function converges or reaches the total iteration number, the iterative update can be performed, and according to the current first low-rank matrix A and second low-rank matrix B, the finally optimized pre-trained weight matrix W = W0 + B·A is obtained. Among them, the matrix B·A has achieved the expected fine update effect through the pruning process of sparsification during the iterative update, thereby realizing the parameter optimization of the pre-trained weight matrix in the pre-trained language model.
[0084] To sum up, this scheme combines the low-rank matrix with the sparsity constraint, performs pruning on the low-rank matrix based on sensitivity during the iterative update process, thereby realizing the sparsification of the matrix, retaining the important parameters with higher sensitivity, making the parameter update in the pre-trained weight matrix selective and refined, improving the stability and efficiency of training, and ensuring the integrity of specific domain knowledge, effectively avoiding knowledge conflicts and forgetting, and making the performance of the trained model better. This scheme mainly has the following advantages:
[0085] 1. By combining LoRA and sparsity constraints, the optimization efficiency and performance of the pre-trained language model on specific tasks are improved, ensuring that large-scale fine-tuning of the pre-trained model can still be carried out under limited computing resources, which is particularly suitable for resource-constrained environments or edge devices.
[0086] 2. The sensitivity metric is used to calculate the impact of each parameter on the loss function, and sensitivity-based pruning is performed on the rows and columns of the low-rank matrix during each iteration update process. This approach dynamically optimizes the pruning ratio during model training and reduces the volatility of pruning, thereby improving the stability and performance of model training.
[0087] 3. The dynamic adjustment mechanism of the sparsity rate provides a smooth pruning process and controls the pruning progress through a preset threshold to ensure training stability while optimizing the model's expressive ability.
[0088] Based on another aspect of the present application, an embodiment of the present application further provides a computing device, which includes a memory for storing computer program instructions and a processor for executing the computer program instructions. When the computer program instructions are executed by the processor, the device is triggered to execute the foregoing parameter optimization method.
[0089] In particular, the method and / or embodiment in the embodiment of the present application can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium. The computer program includes program code for executing the method shown in the flowchart. When the computer program is executed by a processing unit, the above functions defined in the method of the present application are executed.
[0090] It should be noted that the computer-readable medium described in the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable medium can be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, apparatus, or device.
[0091] In this application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination of the foregoing.
[0092] Computer program code for performing the operations of this application may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, execute as a stand-alone software package, execute partially on the user's computer and partially on a remote computer, or execute entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0093] The flowcharts or block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of devices, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0094] As another aspect, the present application also provides a computer-readable medium, which may be included in the devices described in the above embodiments; or may exist separately without being assembled into the devices. The above computer-readable medium carries one or more computer program instructions, and the computer program instructions can be executed by a processor to implement the methods and / or technical solutions of the foregoing multiple embodiments of the present application.
[0095] It should be noted that the present application can be implemented in software and / or a combination of software and hardware. For example, it can be implemented using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device. In some embodiments, the software program of the present application can be executed by a processor to implement the above steps or functions. Similarly, the software program (including related data structures) of the present application can be stored in a computer-readable recording medium, such as a RAM memory, a magnetic or optical drive, or a floppy disk and the like. In addition, some steps or functions of the present application can be implemented using hardware, for example, as a circuit that cooperates with the processor to execute each step or function.
[0096] For those skilled in the art, it is obvious that the present application is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present application. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present application is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present application. Any reference signs in the claims should not be regarded as limiting the claims involved. In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices described in the apparatus claims can also be implemented by one unit or device through software or hardware. The words such as "first" and "second" are used to denote names and do not represent any particular order.
Claims
1. A parameter optimization method, characterized in that, The method includes: Obtain the initial weight matrix W0 that needs to be fine-tuned, and initialize the first low-rank matrix A and the second low-rank matrix B. Determine to obtain the pre-trained weight matrix W = W0 + B·A, where W0 ∈ R d1×d2 , A ∈ R r×d2 , B ∈ R d1×r , and r << min{d1, d2}; Introducing a sparsity constraint into the loss function and redefining the loss function, where the optimization objective corresponding to the loss function is to minimize the loss function by updating the first low-rank matrix A and the second low-rank matrix B while keeping W0 frozen and under the sparsity constraint, and the sparsity constraint is that the proportion of non-zero elements in the matrix B·A is less than or equal to the sparsity rate; Based on the loss function, using the gradient descent algorithm to iteratively update the first low-rank matrix A and the second low-rank matrix B. During the iterative update process, calculate the sensitivity of each element in the pre-trained weight matrix W to the loss function, and perform pruning on the elements in the first low-rank matrix A and the second low-rank matrix B based on the sensitivity; When the optimization termination condition is satisfied, stop the iterative update, and obtain the optimized pre-trained weight matrix W according to the current first low-rank matrix A and the second low-rank matrix B.
2. The method according to claim 1, wherein The redefined loss function L is: where D is the training dataset used for iterative update, is the sparsity constraint, τ is the sparsity rate, indicating that the proportion of non-zero elements in the matrix B·A is less than or equal to the sparsity rate.
3. The method according to claim 2, wherein The sparsity constraint includes a row sparsity constraint on the row elements of the first low-rank matrix A and a column sparsity constraint on the column elements of the second low-rank matrix B. The redefined loss function L is: A i* represents the elements of the i-th row of the first low-rank matrix A, B *i is the element of the i-th column of the second low-rank matrix B. The row sparsity constraint is that the proportion of non-zero elements in the i-th row of the first low-rank matrix A is less than or equal to the sparsity rate, and the column sparsity constraint is that the proportion of non-zero elements in the i-th column of the second low-rank matrix B is less than or equal to the sparsity rate.
4. The method according to claim 1, wherein When calculating the sensitivity of each element in the pre-trained weight matrix W to the loss function, the following calculation formula is used: Among them, I(W ij ) is the sensitivity of the element in the \(i\)-th row and \(j\)-th column of the pre-trained weight matrix \(W\) to the loss function, \(W ij is the weight value of the element in the \(i\)-th row and \(j\)-th column of the pre-trained weight matrix \(W\), is the gradient of the element in the \(i\)-th row and \(j\)-th column of the pre-trained weight matrix \(W\) with respect to the loss function.
5. The method according to claim 4, wherein The method further includes: Smoothing the sensitivity of each element in the pre-trained weight matrix W obtained in this calculation to the loss function according to the sensitivity calculated in the previous iterative update process.
6. The method according to claim 1, wherein Performing pruning on the elements in the first low-rank matrix A and the second low-rank matrix B based on the sensitivity includes: Sorting the elements in each row of the first low-rank matrix A according to the sensitivity respectively; In each row of the first low-rank matrix A, retaining the elements with a higher sensitivity of a preset proportion according to the sorting result, and setting the other elements to zero, where the preset proportion is determined according to the sparsity rate; Sorting the elements in each column of the second low-rank matrix B according to the sensitivity respectively; In each column of the second low-rank matrix B, retaining the elements with a higher sensitivity of a preset proportion according to the sorting result, and setting the other elements to zero.
7. The method according to claim 1, wherein The method further includes: During the iterative update process, dynamically adjusting the sparsity rate, and the specific method is as follows: Among them, τ (t) represents the sparsity rate of the t-th iteration, where t i and t f represent the starting iteration number and the ending iteration number of the dynamic adjustment phase respectively, T represents the total number of iterations, and τ represents the preset sparsity rate threshold.
8. The method according to claim 1, wherein The optimization termination condition includes that the loss function converges or the total number of iterations is reached.
9. A computing device, wherein, The device includes a memory for storing computer program instructions and a processor for executing the computer program instructions. When the computer program instructions are executed by the processor, the device is triggered to execute the method according to any one of claims 1 to 8.
10. A computer-readable medium, on which computer program instructions are stored, and the computer program instructions can be executed by a processor to implement the method according to any one of claims 1 to 8.