Method and system for reducing model fine tuning parameter redundancy based on parameter subspace similarity
Through supervised fine-tuning and singular value decomposition, fine-tuning parameters with maximum subspace similarity are selected, which solves the problems of redundancy and forgetfulness in model fine-tuning, and improves the performance of the model's downstream tasks.
Patent Information
- Application Number
- CN202510366506.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-08
AI Technical Summary
Although the existing high-efficiency fine-tuning method of parameter reductions in computing resources, there is still redundancy, which leads to insufficient learning of the model in downstream tasks and interferes with the pre-trained knowledge, affecting the model performance.
Through supervised fine-tuning and efficient parameter fine-tuning, combining grid search and singular value decomposition, parameter redundancy is calculated and fine-tuning parameters with maximum subspace similarity are selected, and combined into final inference parameters, ensuring that the model learns downstream task knowledge and reduces forgetting.
It effectively reduces the redundancy of model fine-tuning parameters, avoids parameter forgetting problems, and improves the accuracy and professionalism of the model's downstream tasks.
Smart Images

Figure CN120277428A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular, to a method and system for reducing the redundancy of fine-tuning parameters of a model based on the similarity of parameter subspaces. Background Art
[0002] Existing parameter-efficient fine-tuning methods can significantly reduce the computational resources required for fine-tuning, thus achieving efficient fine-tuning. However, even though the number of updated parameters is very small, the parameters still do not completely contain the content required for downstream tasks. That is to say, there is still redundancy in the parameters of efficient fine-tuning, which will cause fine-tuning to not only bring knowledge of downstream tasks, but also interfere with pre-trained knowledge, resulting in catastrophic forgetting. Existing solutions usually focus on avoiding learning redundancy during training, but this method will cause the model to be insufficiently explored during training, thus reducing the possibility of finding the optimal parameters during training.
[0003] Patent application document CN117520842A, application number CN202311478476.1, this invention discloses a downstream task fine-tuning method and system based on incremental learning, belonging to the field of natural language processing. By replaying important samples, imposing regularization on important parameters, and separating parameters for different network architectures, the original large model is fine-tuned, so as to ensure that the fine-tuned large model can learn the knowledge of important downstream tasks while reducing the forgetting of important parameters in the original large model, thereby improving the accuracy and professionalism of the large model in specific tasks. However, data replay and parameter regularization will limit the exploration space of model parameters, which will cause the parameters learned by this method not to be the optimal parameters for downstream tasks. Summary of the Invention
[0004] Aiming at the defects in the prior art, the purpose of the present invention is to provide a method and system for reducing the redundancy of fine-tuning parameters of a model based on the similarity of parameter subspaces.
[0005] According to the method for reducing the redundancy of fine-tuning parameters of a model based on the similarity of parameter subspaces provided by the present invention, it includes:
[0006] Step S1: Fine-tune the instruction fine-tuning base model on the downstream task dataset using supervised fine-tuning and parameter-efficient fine-tuning methods to obtain the fine-tuned low-rank parameters;
[0007] Step S2: Use grid search, set the parameter redundancy measurement step size and search space, and determine a list of low-rank parameters after removing different redundancy components for each parameter-efficient fine-tuning parameter;
[0008] Step S3: For a set of candidate low-rank parameters, use singular value decomposition to obtain the left singular matrix, and calculate the Grassmann distance between the left singular matrix and the initial base model respectively, which is used to measure the similarity of the parameter subspace after removing redundancy;
[0009] Step S4: Select the low-rank parameters with the maximum Grassmann distance, and merge them with the parameters at the corresponding positions of the base model to obtain the fine-tuning parameters actually used in the inference stage.
[0010] Preferably, the step S1 includes:
[0011] Using the downstream fine-tuning data, the fine-tuned parameters are obtained by training the base model fine-tuned based on one instruction with parameter-efficient fine-tuning. The formula is:
[0012]
[0013] where, is the negative log-likelihood loss, represents the expected value on the training dataset D train , x is the query in the dataset, y represents the standard output response in the training set, |y| is the length of y, y i is the i-th token for the base model to learn the standard output, θ is the model parameter, p θ (y i ∣y <i , x) is the probability that the model generates the next token y <i given the prefix y i and the input x;
[0014] After training with the negative log-likelihood loss, the trainable parameters in the base model are updated.
[0015] Preferably, the step S2 includes:
[0016] For a given metric step size s and search space τ, calculate the component sizes to be retained: c = {τ·s, (τ + 1)·s,..., r}, where r is the rank of W. For a certain parameter W = B′A′ in the fine-tuning parameters, through random singular value decomposition, set the components to be retained in the decomposition process to a certain c value, obtain the left singular matrix U′, singular value Σ′, and right singular matrix V′ that retain the largest c components, and finally recombine these components into the form of two matrices multiplied. The formula is: B′ = U′Σ′,
[0017] Preferably, the step S3 includes:
[0018] To ensure that the loss of parameters after removing redundancy does not result in the loss of effective knowledge for downstream model learning, the Grassmann distance is used to measure the similarity between the subspace of a certain parameter combination and the subspace of the base model parameters. For a set of parameters obtained in step S2, the left singular matrix is obtained through singular value decomposition: Using the same singular value decomposition, the left singular matrix U of the parameters at the same parameter position in the base model is obtained. The similarity between different subspaces and the base model subspace is measured by calculating the Grassmann distance between the subspaces of the singular matrices composed of the largest r components. The formula is:
[0019]
[0020] where U r is the matrix composed of the singular vectors corresponding to the largest r singular values of the base model parameters, and U cr is the matrix composed of the singular vectors corresponding to the largest r singular values of the fine-tuning parameters retaining c components. represents the Frobenius norm.
[0021] Preferably, step S4 includes:
[0022] Among the components corresponding to the c Grassmann distances obtained in step S3, the parameter component corresponding to the largest Grassmann distance, that is, the largest subspace similarity, is selected. The formula is:
[0023] c,B′ c ,A′ c =arg maxφ c
[0024] This set of parameters is used to restore the parameter update amount that can be merged back into the base model for downstream inference. The formula is:
[0025] ΔW′=B′ c A′ c
[0026] where c,B′ c ,A′ c is the parameter combination that maximizes φ c .
[0027] According to the system for reducing the redundancy of model fine-tuning parameters based on the similarity of parameter subspaces provided by the present invention, it includes:
[0028] Module M1: Fine-tune the instruction fine-tuned base model on the downstream task dataset using supervised fine-tuning and parameter-efficient fine-tuning methods to obtain the fine-tuned low-rank parameters;
[0029] Module M2: Use grid search to set the parameter redundancy measurement step size and search space, and determine a list of low-rank parameters after removing components with different degrees of redundancy for each efficient fine-tuning parameter;
[0030] Module M3: For a set of candidate low-rank parameters, singular value decomposition is used to obtain the left singular matrix, and the Grassmann distance between the left singular matrix and the initial basis model is calculated respectively to measure the similarity of the parameter subspace after removing redundancy;
[0031] Module M4: Select the low-rank parameters with the largest Grassmann distance and merge them with the parameters at the corresponding position of the base model to obtain the fine-tuning parameters actually used in the inference stage.
[0032] Preferably, the module M1 comprises:
[0033] The fine-tuned base model based on an instruction using downstream fine-tuning data is trained using parameter efficient fine-tuning to obtain the fine-tuned parameters, and the formula is:
[0034]
[0035] in, is the negative log-likelihood loss, Indicates that in the training data set D train , x is the query in the dataset, y represents the standard output response in the training set, |y| is the length of y, and y i is the i-th word unit of the base model learning standard output, θ is the model parameter, and p θ (y i ∣y <i ,x) is the model under a given prefix y <i and generate the next word y when input x i The probability of
[0036] After training with negative log-likelihood loss, the trainable parameters in the base model are updated.
[0037] Preferably, the module M2 comprises:
[0038] For a given metric step size s and search space τ, calculate the size of the components that need to be retained: c = {τ·s, (τ+1)·s, …, r}, where r is the rank of W. For a parameter W = B′A′ in the fine-tuning parameters, set the component that needs to be retained in the decomposition process to a certain c value through random singular value decomposition, and obtain the left singular matrix U′, singular value Σ′ and right singular matrix V′ that retain the maximum c components. Finally, recombine these components into the form of two matrix multiplications, and the formula is: B′ = U′Σ′,
[0039] Preferably, the module M3 comprises:
[0040] To ensure that the loss of parameters after removing redundancy does not cause the loss of effective knowledge for downstream learning of the model, the Grassmann distance is used to measure the similarity between the subspace of a certain parameter combination and the subspace of the base model parameters. For a set of parameters obtained in the module M2, its left singular matrix is obtained through singular value decomposition: Using the same singular value decomposition, the left singular matrix U of the parameters at the same parameter position in the base model is obtained. The similarity between different subspaces and the subspace of the base model is measured by calculating the Grassmann distance between the subspaces of the singular matrices composed of the largest r components. The formula is:
[0041]
[0042] where U r is the matrix composed of the singular vectors corresponding to the largest r singular values of the base model parameters, and U cr is the matrix composed of the singular vectors corresponding to the largest r singular values of the fine-tuning parameters retaining c components. represents the Frobenius norm.
[0043] Preferably, the module M4 includes:
[0044] Among the components corresponding to the c Grassmann distances obtained in the module M3, the parameter components corresponding to the largest Grassmann distance, that is, the largest subspace similarity, are selected. The formula is:
[0045] c,B′ c ,A′ c =arg maxφ c
[0046] This set of parameters is used to restore the parameter update amount that can be merged back into the base model for downstream inference. The formula is:
[0047] ΔW′=B′ c A′ c
[0048] where c,B′ c ,A′ c is the parameter combination that makes φ c maximum.
[0049] Compared with the prior art, the present invention has the following beneficial effects:
[0050] The present invention proposes a method for reducing the redundancy of fine-tuning parameters of a model based on the similarity of parameter subspaces during runtime without training. By comparing the similarity between the fine-tuning parameters and the pre-trained parameter subspaces, the fine-tuning parameter components with the maximum subspace similarity are selected as the final parameters during inference, enabling the model to learn effective knowledge of downstream tasks while avoiding the parameter forgetting problem caused by excessive distribution differences between the fine-tuning parameters and the pre-trained parameters. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Other features, objects, and advantages of the present invention will become more apparent by reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0052] Figure 1 It is a flowchart of a method for reducing the redundancy of fine-tuning parameters of a model based on the similarity of parameter subspaces during runtime without training provided by the present invention;
[0053] Figure 2 It is a schematic structural diagram of a system for reducing the redundancy of fine-tuning parameters of a model based on the similarity of parameter subspaces during runtime without training provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several changes and improvements without departing from the concept of the present invention. These all fall within the protection scope of the present invention.
[0055] Embodiment 1
[0056] The present invention provides a fine-tuning parameter optimization method based on the maximum optimization of subspace similarity, as Figure 1 shown, including:
[0057] Step 1: Train a dialogue model on the downstream task training set using a parameter-efficient fine-tuning method based on a base model, and train it through a negative log loss function to obtain preliminary fine-tuning parameters;
[0058] Step 2: According to the search space with a given grid search step size, use random singular value decomposition to determine a set of fine-tuning parameter candidates with some components cropped for each fine-tuning parameter;
[0059] Step 3: For each parameter in the candidate set, calculate its Grassmann distance from the homologous parameter of the base model to measure the subspace similarity between each candidate fine-tuning parameter and the base model;
[0060] Step 4: Select the candidate fine-tuning parameters with the maximum subspace similarity and restore them to the parameters used during the operation of the base model, thus ensuring that the fine-tuning parameters learn the knowledge of the downstream task while maintaining the distribution of the base model.
[0061] Specifically, first, using the downstream fine-tuning data, the fine-tuned parameters are obtained through parameter-efficient fine-tuning training of a base model fine-tuned based on an instruction. Its formula can be written as:
[0062]
[0063] where x is the query in the dataset, y represents the standard output response in the training set, and y i is the i-th token for the base model to learn the standard output. θ represents the fine-tuning model with θ as the trainable parameter. After training with the negative log-likelihood loss, the trainable parameters in the base model are updated.
[0064] Next, grid search is needed to determine a set of fine-tuning parameters that can be pruned for each fine-tuning parameter, for subsequent selection of the optimal runtime parameters from this set of candidates. For a given search step size s and search space τ, calculate the size of the components to be retained: c = {τ·s, (τ + 1)·s,..., r}, where r is the rank of W. For a certain parameter W = BA in the fine-tuning parameters, through random singular value decomposition, set the components to be retained during the decomposition process to a certain c value to obtain the left singular matrix U′, singular value Σ′, and right singular matrix V′ that retain the largest c components. Finally, recombine these components into the form of two matrices multiplied, and its formula can be written as:
[0065] B′ = U′Σ′
[0066]
[0067] To ensure that the loss of parameters after removing redundancy does not cause the model to lose effective knowledge for downstream learning, the Grassmann distance is used to measure the similarity between the subspace of a certain parameter combination and the subspace of the base model parameters. For a set of parameters obtained in the step S2, its left singular matrix is obtained through singular value decomposition: Using the same singular value decomposition, obtain the left singular matrix U of the parameters at the same parameter position in the base model. By calculating the Grassmann distance between the subspaces of the singular matrices composed of the largest r components, the similarity between different subspaces and the base model subspace is measured, and its formula can be written as:
[0068]
[0069] where c is the set of component retention sizes within the search range. U r is the matrix composed of the singular vectors corresponding to the largest r singular values of the base model parameters, Ucr is a matrix composed of the singular vectors corresponding to the largest r singular values of the fine-tuning parameters that retain c components.
[0070] To obtain the optimal parameters during runtime, from the components corresponding to c Grassmann distances, select the parameter component with the largest Grassmann distance, that is, the one corresponding to the largest subspace similarity. Its formula can be written as:
[0071] c, B′ c , A′ c = arg maxφ c
[0072] This set of parameters is used to restore the parameter update amount that can be merged back into the base model for downstream inference. Its formula can be written as: ΔW′ = B′ c A′ c .
[0073] Example 2
[0074] The present invention also provides a system for reducing the redundancy of fine-tuning parameters of a model based on parameter subspace similarity. The system for reducing the redundancy of fine-tuning parameters of a model based on parameter subspace similarity can be implemented by executing the process steps of the method for reducing the redundancy of fine-tuning parameters of a model based on parameter subspace similarity. That is, those skilled in the art can understand the method for reducing the redundancy of fine-tuning parameters of a model based on parameter subspace similarity as a preferred implementation manner of the system for reducing the redundancy of fine-tuning parameters of a model based on parameter subspace similarity.
[0075] Such as Figure 2 , according to the system for reducing the redundancy of fine-tuning parameters of a model based on parameter subspace similarity provided by the present invention, it includes: Module M1: Fine-tune the instruction fine-tuning base model on the downstream task dataset using supervised fine-tuning and parameter-efficient fine-tuning methods to obtain the fine-tuned low-rank parameters; Module M2: Use grid search, set the parameter redundancy measurement step size and search space, and determine a list of low-rank parameters after removing components with different redundancy degrees for each efficient fine-tuning parameter; Module M3: For a set of candidate low-rank parameters, use singular value decomposition to obtain the left singular matrix, and calculate the Grassmann distance between the left singular matrix and the initial base model respectively to measure the parameter subspace similarity after removing redundancy; Module M4: Select the low-rank parameter with the largest Grassmann distance and merge it with the parameter at the corresponding position of the base model to obtain the fine-tuning parameters actually used in the inference stage.
[0076] The Module M1 includes: Using the downstream fine-tuning data to train the fine-tuned parameters based on an instruction fine-tuning base model using parameter-efficient fine-tuning. Its formula is:
[0077]
[0078] Among them, is the negative log-likelihood loss, represents the expected value on the training dataset D train . x is the query in the dataset, y represents the standard output response in the training set, |y| is the length of y, and y i is the i-th token for the base model to learn the standard output, θ is the model parameter, and p θ (y i ∣y <i , x) is the probability that the model generates the next token y <i given the prefix y i and the input x; after training with the negative log-likelihood loss, the trainable parameters in the base model are updated.
[0079] The module M2 includes: for a given measurement step size s and search space τ, calculate the component size to be retained: c = {τ·s, (τ + 1)·s,..., r}, where r is the rank of W. For a certain parameter W = B′A′ in the fine-tuning parameters, through random singular value decomposition, set the components to be retained in the decomposition process to a certain c value, obtain the left singular matrix U′, singular value Σ′, and right singular matrix V′ that retain the largest c components, and finally recombine these components into the form of two matrices multiplied, and its formula is: B′ = U′Σ′,
[0080] The module M3 includes: to ensure that the parameter loss after removing redundancy does not lose effective knowledge for the downstream learning of the model, use the Grassmann distance to measure the similarity between the subspace of a certain parameter combination and the subspace of the base model parameters. For a set of parameters obtained in the module M2, obtain its left singular matrix through singular value decomposition: Using the same singular value decomposition, obtain the left singular matrix U of the parameters at the same parameter position in the base model, and measure the similarity between different subspaces and the base model subspace by calculating the Grassmann distance between the subspaces of the singular matrices composed of the largest r components, and its formula is:
[0081]
[0082] Among them, U r is the matrix composed of the singular vectors corresponding to the largest r singular values of the base model parameters, and U cr is the matrix composed of the singular vectors corresponding to the largest r singular values of the fine-tuning parameters that retain c components, represents the Frobenius norm.
[0083] The module M4 includes: among the components corresponding to the c Grassmann distances obtained in the module M3, select the parameter component corresponding to the maximum Grassmann distance, that is, the maximum subspace similarity, and its formula is:
[0084] c,B′ c ,A′ c =arg maxφ c
[0085] This set of parameters is used to restore the parameter update amount that can be merged back into the base model for downstream inference, and its formula is:
[0086] ΔW′=B′ c A′ c
[0087] where c,B′ c ,A′ c is the parameter combination that makes φ c maximum.
[0088] Those skilled in the art know that in addition to implementing the systems, devices, and their respective modules provided by the present invention in the form of pure computer-readable program codes, the method steps can be logically programmed to enable the systems, devices, and their respective modules provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers, etc. to achieve the same program. Therefore, the systems, devices, and their respective modules provided by the present invention can be regarded as a kind of hardware component, and the modules included therein for implementing various programs can also be regarded as the structures within the hardware component; the modules for implementing various functions can also be regarded as both software programs for implementing the method and the structures within the hardware component.
[0089] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined arbitrarily with each other.
Claims
1. A method for reducing the parameter redundancy of model fine-tuning based on the similarity of parameter subspaces, characterized in that, Including: Step S1: Fine-tune the instruction fine-tuning base model on the downstream task dataset using supervised fine-tuning and parameter-efficient fine-tuning methods to obtain the fine-tuned low-rank parameters; Step S2: Use grid search, set the parameter redundancy metric step size and search space, and determine a list of low-rank parameters after removing components with different redundancy degrees for each efficient fine-tuning parameter; Step S3: For a set of candidate low-rank parameters, use singular value decomposition to obtain the left singular matrix, and calculate the Grassmann distance between the left singular matrix and the initial base model respectively to measure the similarity of the parameter subspace after removing redundancy; Step S4: Select the low-rank parameters with the maximum Grassmann distance and merge them with the parameters at the corresponding positions of the base model to obtain the fine-tuning parameters actually used in the inference stage.
2. The method for reducing the parameter redundancy of a model fine-tuning based on the similarity of parameter subspaces according to claim 1, wherein The said Step S1 includes: Using the downstream fine-tuning data, based on an instruction fine-tuning base model, use parameter-efficient fine-tuning to train and obtain the fine-tuned parameters, and its formula is: Among them, is the negative log-likelihood loss, which represents the expected value on the training dataset D train , x is the query in the dataset, y represents the standard output response in the training set, |y| is the length of y, and y i is the i-th token for the base model to learn the standard output, θ is the model parameter, and p θ (y i ∣y <i , x) is the probability that the model generates the next token y <i given the prefix y i and the input x; After training with negative log-likelihood loss, the trainable parameters in the base model are updated.
3. The method for reducing the parameter redundancy of a model fine-tuning based on the similarity of parameter subspaces according to claim 1, wherein The said Step S2 includes: For a given metric step size s and search space τ, calculate the component sizes to be retained: c = {τ·s, (τ + 1)·s, …, r}, where r is the rank of W. For a certain parameter W = B′A′ in the fine-tuning parameters, through random singular value decomposition, set the components to be retained in the decomposition process to a certain c value, obtain the left singular matrix U′, singular value Σ′, and right singular matrix V′ that retain the largest c components. Finally, recombine these components into the form of two matrices multiplied, and its formula is: B′ = U′Σ′, A′ = V′ T [1:c, :].
4. The method for reducing the parameter redundancy of a model fine-tuning based on the similarity of parameter subspaces according to claim 3, characterized in that, The said Step S3 includes: To ensure that the loss of parameters after removing redundancy does not result in the loss of effective knowledge for downstream model learning, the Grassmann distance is used to measure the similarity between the subspace of a certain parameter combination and the subspace of the base model parameters. For a set of parameters obtained in step S2, the left singular matrix is obtained through singular value decomposition: Using the same singular value decomposition, the left singular matrix U of the parameters at the same parameter position in the base model is obtained. The similarity between different subspaces and the base model subspace is measured by calculating the Grassmann distance between the subspaces of the singular matrices composed of the largest r components. The formula is as follows: Among them, U r is a matrix composed of the singular vectors corresponding to the largest r singular values of the base model parameters, and U cr is a matrix composed of the singular vectors corresponding to the largest r singular values of the fine-tuning parameters that retain c components. represents the Frobenius norm.
5. The method for reducing the parameter redundancy of a model fine-tuning based on the similarity of parameter subspaces according to claim 4, characterized in that The said Step S4 includes: Among the components corresponding to the c Grassmann distances obtained in Step S3, select the parameter component with the maximum Grassmann distance, that is, the parameter component corresponding to the maximum subspace similarity, and its formula is: c,B′ v ,A′ v = argmaxφ v This set of parameters is used to restore the parameter update amount that can be merged back into the base model for downstream inference, and its formula is: ΔW′ = B′ c A′ c where c, B' c , A' c is the parameter combination that maximizes φ c to the greatest extent.
6. A system for reducing the redundancy of model fine-tuning parameters based on the similarity of parameter subspaces, characterized in that, Including: Module M1: Fine-tune the instruction fine-tuning base model on the downstream task dataset using supervised fine-tuning and parameter-efficient fine-tuning methods to obtain the fine-tuned low-rank parameters; Module M2: Use grid search, set the parameter redundancy metric step size and search space, and determine a list of low-rank parameters after removing components with different redundancy degrees for each efficient fine-tuning parameter; Module M3: For a set of candidate low-rank parameters, use singular value decomposition to obtain the left singular matrix, and calculate the Grassmann distance between the left singular matrix and the initial base model respectively to measure the similarity of the parameter subspace after removing redundancy; Module M4: Select the low-rank parameters with the maximum Grassmann distance and merge them with the parameters at the corresponding positions of the base model to obtain the fine-tuning parameters actually used in the inference stage.
7. The system for reducing the parameter redundancy of a model fine-tuning based on the similarity of parameter subspaces according to claim 6, wherein The said Module M1 includes: Using the downstream fine-tuning data, based on an instruction fine-tuning base model, use parameter-efficient fine-tuning to train and obtain the fine-tuned parameters, and its formula is: where, is the negative log-likelihood loss, represents the expected value on the training dataset D train where x is the query in the dataset, y represents the standard output response in the training set, |y| is the length of y, y i is the i-th token for the base model to learn the standard output, θ is the model parameter, p θ (y i ∣y <i ,x) is the probability that the model generates the next token y <i given the prefix y i and the input x; After training with negative log-likelihood loss, the trainable parameters in the base model are updated.
8. The system for reducing the parameter redundancy of the model fine-tuning based on the parameter subspace similarity according to claim 6, wherein The said Module M2 includes: For a given metric step size s and search space τ, calculate the component sizes to be retained: c = {τ·s, (τ + 1)·s, …, r}, where r is the rank of W. For a certain parameter W = B′A′ in the fine-tuning parameters, through random singular value decomposition, set the components to be retained in the decomposition process to a certain c value, obtain the left singular matrix U′, singular value Σ′, and right singular matrix V′ that retain the largest c components, and finally recombine these components into the form of two matrices multiplied, and its formula is: B′ = U′Σ′, A′ = V′ T [1:c, :].
9. The system for reducing the parameter redundancy of a model fine-tuning based on the similarity of parameter subspaces according to claim 8, characterized in that The said Module M3 includes: To ensure that the loss of parameters after removing redundancy does not result in the loss of effective knowledge for downstream model learning, the Grassmann distance is used to measure the similarity between the subspace of a certain parameter combination and the subspace of the base model parameters. For a set of parameters obtained in the module M2, its left singular matrix is obtained through singular value decomposition: Using the same singular value decomposition, the left singular matrix U of the parameters at the same parameter position in the base model is obtained. The similarity between different subspaces and the base model subspace is measured by calculating the Grassmann distance between the subspaces of the singular matrices composed of the largest r components. The formula is as follows: Among them, U r is a matrix composed of the singular vectors corresponding to the largest r singular values of the base model parameters, and U cr is a matrix composed of the singular vectors corresponding to the largest r singular values of the fine-tuning parameters retaining c components, represents the Frobenius norm.
10. The system for reducing the parameter redundancy of a model fine-tuning based on the similarity of parameter subspaces according to claim 9, characterized in that The said Module M4 includes: Among the components corresponding to the c Grassmann distances obtained in Module M3, select the parameter component with the maximum Grassmann distance, that is, the parameter component corresponding to the maximum subspace similarity, and its formula is: c,B′ c ,A′ c = argmaxφ c This set of parameters is used to restore the parameter update amount that can be merged back into the base model for downstream inference, and its formula is: ΔW′ = B′ c A′ c where c, B' c , A' c is the parameter combination that maximizes φ c
Citation Information
Patent Citations
Model training method, data processing method, equipment and storage medium
CN117520842A