Task processing method and device based on model parameter adjustment, equipment and medium

Bypassed low-rank matrix related to task is constructed through singular value decomposition and gradient information screening, and only fine-tuning of the matrix is ​​solved, which solves the problem of insufficient initialization of low-rank matrix in the existing technology, and improves the performance and generalization capabilities of the model in complex domain tasks.

CN120068970APending Publication Date: 2025-05-30PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510146011.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, low-rank matrix initialization fails to combine the characteristics of model parameters, resulting in limited performance and insufficient generalization capabilities of the model in complex domain tasks.

Method used

By collecting the auxiliary benchmark dataset, the original weight matrix is ​​extracted and singular value decomposition is performed, the target singular value feature quantity and singular vector are filtered based on the gradient information, the bypass low rank matrix is ​​constructed, and the original weight matrix is ​​frozen, and only the bypass low rank matrix is ​​fine-tuned to adjust the model parameters.

Benefits of technology

减少了低秩矩阵初始化的随机性,降低了计算开销,提高了模型优化效率和泛化能力,使模型在复杂领域任务中表现更为稳定和准确。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068970A_ABST
    Figure CN120068970A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, the field of medical health and the field of financial science and technology, and discloses a task processing method based on model parameter adjustment, which comprises the steps of collecting an auxiliary reference data set and extracting an original weight matrix, decomposing the original weight matrix into a plurality of singular value feature quantities and corresponding singular vectors, and screening target singular value characteristic quantities and corresponding target singular vectors based on gradient information, and constructing a bypass low-rank matrix. And freezing the original weight matrix of the to-be-optimized model, adjusting model parameters by finely tuning the bypass low-rank matrix, generating an optimized target model, and processing a task by using the optimized target model to generate an execution result. According to the method, singular value characteristic quantities and corresponding vectors are screened based on gradient information, a low-rank matrix related to task characteristics is constructed, and the influence of the randomness of initialization of the low-rank matrix on the optimization process is reduced; and the original weight matrix is frozen, and only the bypass low-rank matrix is finely tuned, so that the calculation overhead is reduced, and the model optimization efficiency and generalization ability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence technology, medical and health, and fintech, and particularly relates to a task processing method, device, equipment, and storage medium based on model parameter adjustment. Background Art

[0002] With the wide application of large language models (LLMs) in various tasks, pre-trained language models (PLMs) have shown great potential in the medical and health and financial fields due to their excellent generalization ability. In the medical and health field, these models are widely used in tasks such as electronic medical record analysis, disease prediction, and medical image assisted diagnosis; in the financial field, they are used for risk assessment, market prediction, financial text analysis, etc. However, these complex tasks pose higher requirements for the accuracy and efficiency of the models.

[0003] The traditional full fine-tuning (FT) method adapts to specific tasks by updating all parameters in the model, but its high computational and memory overhead makes it difficult to be practically applied in resource-constrained environments. For example, in the medical and health field, processing massive patient data requires a model with fast response and resource conservation; in the financial field, high-frequency trading and real-time risk assessment tasks require the model to have extremely high real-time performance and resource utilization efficiency. This makes the limitations of full fine-tuning particularly obvious.

[0004] Therefore, the parameter-efficient fine-tuning (PEFT) method has gradually attracted attention. The PEFT method significantly reduces the computational and memory overhead of fine-tuning by only adjusting a small number of parameters while keeping the base model unchanged. This method enables large language models to more efficiently adapt to different tasks, especially having obvious advantages in practical environments with limited resources. As a representative of the PEFT method, LoRA has been widely applied in multiple fields by introducing a bypass low-rank matrix, which greatly reduces the number of parameters to be updated while maintaining the low-latency characteristic of model inference. However, LoRA also has the following significant deficiencies in practical applications:

[0005] The bypass low-rank matrix of LoRA is usually randomly initialized using a Gaussian distribution. This initialization method may result in small gradients or random directions in the initial stage of fine-tuning, making it easy for the optimization process to fall into a local optimal solution. This limitation is particularly obvious in complex tasks. For example, in a disease prediction model in the medical field, it may lead to inaccurate diagnosis results; in a financial risk assessment model, it may cause prediction biases and reduce the reliability of decision-making.

[0006] The initialization of the bypass low-rank matrix of LoRA does not consider the specific requirements of downstream tasks. Therefore, the initialization may contain a large amount of information irrelevant to the target task. This irrelevance weakens the performance of the model in specific tasks. For example, in the medical field, the characteristics of diagnostic data for different diseases vary greatly, and the low-rank matrix without targeted optimization may reduce the accuracy of the model in sub-diseases; in the financial field, the data distribution changes significantly under different market scenarios, and the lack of task-related optimization may lead to poor performance of the model under certain market conditions.

[0007] LoRA differs from the full fine-tuning (FT) method in the direction and magnitude of weight updates, which results in the learning ability and generalization performance of the LoRA model being unable to match that of the FT method in some cases. For example, in the field of healthcare, tasks such as electronic medical record analysis often involve complex context information, and the limitations of LoRA may affect the accurate parsing of long texts; in the market prediction task in the financial field, the performance gap of LoRA may make it difficult to capture long-term trend signals.

[0008] Although existing improvement methods (such as Pissa, DoRA, and LoRA-GA) have tried to solve these problems, there are still deficiencies. These methods mainly focus on optimizing the initialization method of the low-rank matrix or different strategies for decomposing the weight matrix, but most of them fail to effectively combine the specific requirements of downstream tasks. Therefore, in the application of existing methods in complex fields such as healthcare and finance, the stability of model performance and the optimization effect have not yet reached the ideal level. Summary of the Invention

[0009] The main purpose of the present invention is to provide a task processing method, device, equipment, and storage medium based on model parameter adjustment, aiming to solve the technical problems in the prior art that the initialization of the low-rank matrix fails to combine the characteristics of model parameters, resulting in limited performance and insufficient generalization ability of the model in complex field tasks.

[0010] To achieve the above object, the present invention provides a task processing method based on model parameter adjustment, including:

[0011] Collect an auxiliary benchmark data set and extract the original weight matrix from the model to be optimized;

[0012] Decompose the original weight matrix into multiple singular value eigenquantities and corresponding singular vectors for each singular value eigenquantity;

[0013] Based on the auxiliary benchmark data set, determine the gradient information of the multiple singular value eigenquantities;

[0014] Determine the target singular value eigenquantity whose gradient information intensity exceeds a preset threshold and the corresponding target singular vector of the target singular value eigenquantity;

[0015] Construct a bypass low-rank matrix based on the target singular value feature quantity and the target singular vector;

[0016] Freeze the original weight matrix of the model to be optimized, and adjust the parameters of the model to be optimized based on the bypass low-rank matrix to obtain an optimized target model;

[0017] Process the target task based on the optimized target model to generate an execution result of the target task.

[0018] Furthermore, to achieve the above object, the present invention provides a task processing device based on model parameter adjustment, including:

[0019] A data processing module, configured to collect an auxiliary reference data set and extract an original weight matrix from the model to be optimized;

[0020] A matrix decomposition module, configured to decompose the original weight matrix into a plurality of singular value feature quantities and singular vectors corresponding to each singular value feature quantity;

[0021] A gradient analysis module, configured to determine gradient information of the plurality of singular value feature quantities based on the auxiliary reference data set;

[0022] A feature screening module, configured to determine a target singular value feature quantity whose gradient information intensity exceeds a preset threshold and a target singular vector corresponding to the target singular value feature quantity;

[0023] A low-rank matrix construction module, configured to construct a bypass low-rank matrix based on the target singular value feature quantity and the target singular vector;

[0024] A model optimization module, configured to freeze the original weight matrix of the model to be optimized, and adjust the parameters of the model to be optimized based on the bypass low-rank matrix to obtain an optimized target model;

[0025] A task processing module, configured to process the target task based on the optimized target model to generate an execution result of the target task.

[0026] Furthermore, to achieve the above object, the present invention further provides a computer device, the computer device includes a memory, a processor, and a task processing program based on model parameter adjustment stored in the memory and executable on the processor, and when the task processing program based on model parameter adjustment is executed by the processor, the steps of the task processing method based on model parameter adjustment as described above are implemented.

[0027] Further, to achieve the above object, the present invention also provides a computer-readable storage medium, on which a task processing program based on model parameter adjustment is stored. When the task processing program based on model parameter adjustment is executed by a processor, the steps of the task processing method based on model parameter adjustment as described above are implemented.

[0028] Beneficial effects: The present invention relates to the fields of artificial intelligence technology, medical and health fields, and fintech fields. It discloses a task processing method based on model parameter adjustment, including: collecting an auxiliary reference data set and extracting an original weight matrix, decomposing it into multiple singular value feature quantities and corresponding singular vectors, screening target singular value feature quantities and corresponding target singular vectors based on gradient information, and constructing a bypass low-rank matrix. Freeze the original weight matrix of the model to be optimized, adjust the model parameters by fine-tuning the bypass low-rank matrix, generate an optimized target model, and use the optimized target model to process tasks to generate execution results. By screening singular value feature quantities and corresponding vectors based on gradient information, the present invention constructs a low-rank matrix related to task characteristics, reduces the influence of the randomness of low-rank matrix initialization on the optimization process; freezes the original weight matrix and only fine-tunes the bypass low-rank matrix, reducing the computational overhead and improving the model optimization efficiency and generalization ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings:

[0030] Figure 1 is a schematic diagram of an application environment of the task processing method based on model parameter adjustment in an embodiment of the present invention;

[0031] Figure 2 is a schematic flowchart of an embodiment of the task processing method based on model parameter adjustment of the present invention;

[0032] Figure 3 is a schematic diagram of the functional modules of a preferred embodiment of the task processing device based on model parameter adjustment of the present invention;

[0033] Figure 4 is a schematic diagram of the structure of a computer device in an embodiment of the present invention;

[0034] Figure 5 is another schematic diagram of the structure of a computer device in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0035] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0036] The task processing method based on model parameter adjustment provided by the embodiments of the present invention can be applied in, for example Figure 1In the application environment, the client communicates with the server through the network. The server can collect an auxiliary reference dataset through the client, extract the original weight matrix, decompose it into multiple singular value feature quantities and corresponding singular vectors, screen the target singular value feature quantities and corresponding target singular vectors based on gradient information, and construct a bypass low-rank matrix. Freeze the original weight matrix of the model to be optimized, adjust the model parameters by fine-tuning the bypass low-rank matrix, generate an optimized target model, and use the optimized target model to process tasks and generate execution results. The present invention screens singular value feature quantities and corresponding vectors based on gradient information, constructs a low-rank matrix related to task characteristics, reduces the influence of the randomness of low-rank matrix initialization on the optimization process; freezes the original weight matrix and only fine-tunes the bypass low-rank matrix, reduces the computational overhead, and improves the model optimization efficiency and generalization ability. Among them, the client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail below through specific embodiments.

[0037] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of an embodiment of the task processing method based on model parameter adjustment provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.

[0038] As Figure 2 shown, the task processing method based on model parameter adjustment proposed by the present invention includes the following steps:

[0039] S10, collect an auxiliary reference dataset and extract the original weight matrix from the model to be optimized;

[0040] In this embodiment, the collection of the auxiliary reference dataset is to provide data samples with task relevance or domain representativeness to the model to be optimized. The reference dataset can be selected from public data sources, domain-specific datasets, or data generated during the model training process. During the collection process, it is necessary to ensure the diversity and integrity of the data, and at the same time remove abnormal samples to avoid interference with subsequent gradient information calculation.

[0041] The data preparation work is completed through the data collection module. For the medical and health field, samples can be selected from electronic medical record databases and disease diagnosis datasets, and the data is preprocessed by standardization, such as normalizing numerical data or tokenizing and encoding text data. For the financial field, market transaction data, financial statements, or historical market volatility data can be collected, and at the same time, outliers or invalid records in the data are cleaned through denoising algorithms.

[0042] The original weight matrix is the core parameter structure generated after the pre-training of the model, containing the general features learned during the pre-training process. These weights need to be extracted from the model to be optimized and used as the basis for subsequent low-rank decomposition.

[0043] Using the model parameter extraction module, the core weight layer parameters in the pre-trained model are obtained through the interfaces of deep learning frameworks (such as PyTorch or TensorFlow). For example, in the diagnostic tasks in the medical and health field, the intermediate weight matrix used by the model for text or image feature extraction is extracted; in the time series prediction in the financial field, the weight layer parameters used by the model for trend modeling are extracted.

[0044] In addition, the auxiliary benchmark dataset can be directly associated with the target task. When collecting data, focus on selecting data samples with task-specific features. For example, in the medical and health field for diabetes diagnosis tasks, relevant case data can be collected; in the financial field for bond risk assessment tasks, market volatility data and rating records of specific bonds can be collected.

[0045] Example illustration: If the target task is to optimize disease diagnosis based on a large language model. When collecting the auxiliary benchmark dataset, text data containing diagnostic descriptions can be extracted from the electronic medical records of different medical institutions. For extracting the original weight matrix, the weight matrix of the text embedding layer of the pre-trained model can be selected. Through these data and weights, the model can gradually adapt to the characteristics of different disease types, thereby improving the diagnostic accuracy. If the target task is to optimize market risk prediction. The auxiliary benchmark dataset can collect historical trading data from multiple markets, covering the market volatility characteristics of different time periods. When extracting the original weight matrix, the weight matrix of the time series convolutional layer used for trend extraction in the pre-trained model can be selected. Through these data and weights, the model can more effectively capture market risk changes and improve prediction stability.

[0046] By collecting the auxiliary benchmark dataset from domain-representative data and extracting the original weight matrix of the model to be optimized, it provides a key basis for subsequent singular value decomposition and the construction of task-related low-rank matrices. The introduction of the auxiliary benchmark dataset enhances the flexibility of the model's adaptation to the task and avoids the performance fluctuation problems caused by random initialization. Extracting the weight matrix ensures the accuracy of the starting point of model optimization and the coherence of operations, significantly improving the efficiency and effectiveness of parameter adjustment.

[0047] S20, decompose the original weight matrix into multiple singular value feature quantities and the corresponding singular vectors for each singular value feature quantity;

[0048] In this embodiment, before decomposing the original weight matrix, it is necessary to ensure that the numerical range and dimensional structure of the matrix meet the requirements of singular value decomposition (SVD), including eliminating abnormal data and optimizing the matrix shape. Common preprocessing operations include normalization, removing non-numerical terms, and adjusting the sparsity of the matrix. Use numerical processing tools (such as NumPy or Pandas) to repair extreme values and non-numerical terms in the matrix. For sparse matrices, through filling or cropping operations, the matrix can be made suitable for SVD calculation. The specific implementation can use the electronic medical record matrix in the medical field to ensure the consistency of the numerical range of diagnostic information through normalization; in the financial field, intercept the time series weight matrix to optimize the dimensions.

[0049] Singular value decomposition decomposes a matrix into the product of three matrices, which respectively represent the left singular vector matrix, the diagonal singular value matrix, and the right singular vector matrix. Through SVD, the original weight matrix can be extracted as singular value features and singular vectors. Use deep learning frameworks or matrix operation tools (such as PyTorch, TensorFlow, or SciPy) to perform SVD decomposition, and decompose the original weight matrix into a singular value feature matrix and the corresponding singular vector matrix. In the medical field, the embedding matrix of diagnostic text can be decomposed to extract core diagnostic features; in the financial field, the market transaction matrix can be decomposed to extract the main direction of trading fluctuations.

[0050] The singular value features are the values in the diagonal matrix after SVD decomposition, which represent the important characteristics of the original weight matrix. The purpose of extracting singular value features is to retain the core information in the matrix. Extract the singular value features from the decomposed diagonal matrix and arrange them in descending order for subsequent screening. In the medical field, the singular values in the diagnostic matrix can be extracted to determine the important features of disease diagnosis; in the financial field, the singular values in the market transaction matrix can be extracted to identify the main features of market fluctuations.

[0051] The singular vectors are composed of the left singular vector matrix and the right singular vector matrix in the singular value decomposition, which respectively represent the row space and column space characteristics of the original matrix. Extracting singular vectors is used to characterize the directionality of the matrix structure and content. Combining the singular value features, extract the corresponding singular vectors and store the singular values and singular vectors in association. In the medical field, the singular vectors embedded in the diagnostic text can be used for disease classification tasks; in the financial field, the singular vectors of the transaction matrix can be used for the characteristic analysis of market segments.

[0052] Example illustration: In a disease diagnosis task, the weight matrix in the pre-trained model (such as the embedding layer weight) is used as the original matrix, and singular value decomposition is utilized to extract diagnosis-related features. Through the screening of singular value feature quantities, key diagnostic indicators can be identified, such as the main symptoms or influencing factors of the disease. In a market prediction task, transaction behavior data is transformed into matrix form, and singular value decomposition is performed on its weight matrix. The extracted singular value feature quantities can reflect the core patterns of market fluctuations, while the singular vectors are used to identify the trading characteristics of different industry sectors.

[0053] By performing singular value decomposition on the original weight matrix, the matrix is transformed into a low-rank representation form that is easy to optimize. The extraction of singular value feature quantities retains the core information of the matrix, while the singular vectors provide a structured description of the matrix's directionality. The overall process significantly improves the efficiency of low-rank matrix construction and lays a task-related feature foundation for subsequent model optimization.

[0054] S30, based on the auxiliary benchmark dataset, determine the gradient information of the multiple singular value feature quantities;

[0055] In this embodiment, the auxiliary benchmark dataset is a representative dataset used to interact with the model to be optimized to extract gradient information related to model parameters. The benchmark data is input into the model to be optimized item by item or in batches. Through forward propagation, the model's feature response to the input data is obtained. Load the auxiliary benchmark dataset using a deep learning framework (such as PyTorch or TensorFlow) and input it into the model to be optimized item by item or in small batches. The forward propagation stage of the model generates feature responses or intermediate activation values. For example, in the field of healthcare, when inputting patient symptom data, the possible disease categories are output; in the financial field, when inputting market indicators, the risk score is output.

[0056] The loss function is used to quantify the prediction error of the model for the benchmark data and drive the calculation of the gradient. The loss function should be able to reflect the characteristics of the benchmark data and the requirements of the target task, such as the cross-entropy loss for classification tasks or the mean squared error for regression tasks. Define the loss function according to the nature of the benchmark data. For example:

[0057] In the field of healthcare, the benchmark data may contain disease diagnosis labels, and the cross-entropy loss function is used to measure the diagnostic accuracy;

[0058] In the financial field, the benchmark data may contain market trading trends, and the mean squared error is used to measure the deviation between the predicted value and the actual value.

[0059] Using automatic differentiation techniques, based on model parameters and loss functions, the gradient information of the original weight matrix is calculated by backpropagation layer by layer. The automatic differentiation module can accurately and efficiently calculate the gradients of the model, providing a basis for subsequent singular value gradient extraction. Call the automatic differentiation function of the deep learning framework (such as autograd in PyTorch or GradientTape in TensorFlow) to perform backpropagation on the loss function and generate the gradients of the model parameters (including the original weight matrix). For example, in the medical field, calculate the gradients of diagnostic tasks to identify key disease features; in the financial field, calculate the gradients of trading prediction tasks to reflect the impact of market changes on model weights.

[0060] Map from the gradient information of the original weight matrix to singular value feature quantities, and quantify their contribution to the loss by calculating the change trend of each singular value feature quantity in the gradient. Use the weight matrix and gradient information after singular value decomposition to calculate the gradient changes of each singular value feature quantity. For example, in the medical field, extract the gradients of singular value feature quantities to determine whether certain symptoms have a significant impact on the disease classification model; in the financial field, analyze the gradient changes of singular value feature quantities to evaluate the core factors of market fluctuations.

[0061] Example illustration: In the medical and health field, the benchmark data includes the symptom descriptions of patients and the corresponding disease labels. Input the data into the disease diagnosis model, calculate the gradient information of the model by defining the cross-entropy loss function, and extract the gradient changes of singular value feature quantities to identify the key features that most affect disease classification. In the financial field, the benchmark data is the historical trading data of the stock market and the corresponding risk levels. Input the data into the market risk prediction model, calculate the gradient information of the original weight matrix by defining the mean squared error loss function, and extract the gradient changes of singular value feature quantities to identify the trading features that are most important for market volatility prediction.

[0062] By utilizing the interaction between the auxiliary benchmark dataset and the model to be optimized, combined with automatic differentiation techniques, accurately calculate the gradient information of singular value feature quantities. Extracting the gradient information effectively quantifies the task relevance of singular value feature quantities, providing a scientific basis for subsequent screening of target singular value feature quantities, and significantly improving the adaptability and optimization efficiency of the model.

[0063] S40, determine the target singular value feature quantities whose gradient information intensity exceeds the preset threshold and the target singular vectors corresponding to the target singular value feature quantities;

[0064] In this embodiment, the intensity of the gradient information refers to the magnitude of the gradient value, which is an important indicator for measuring the contribution degree of the singular value feature quantity to the loss function. By analyzing the gradient values of each singular value feature quantity, it can be determined whether it plays a significant role in the optimization process. Extract the numerical magnitude from the gradient information of the singular value feature quantity calculated in the previous step, and calculate its mean value, maximum value or other statistical indicators as the measurement standard of the intensity. For example:

[0065] In the field of medical and health, analyze the gradient intensity related to key disease characteristics;

[0066] In the financial field, analyze the gradient intensity related to market volatility characteristics.

[0067] The preset threshold is used to filter out the singular value feature quantities with relatively small gradient intensity, ensuring that the target singular value feature quantities for constructing the low-rank matrix subsequently have a high task relevance. The preset threshold can be determined by empirical values, data distribution analysis or hyperparameter tuning.

[0068] Set the threshold according to the statistical distribution of the gradient information. For example, take the 90% quantile of the gradient value as the preset threshold, or the threshold can be dynamically adjusted through task-related verification. For example:

[0069] In the field of medical and health, an appropriate threshold can be set according to the complexity of the key disease;

[0070] In the financial field, an appropriate gradient intensity threshold can be set according to the fluctuation range of the market volatility data.

[0071] According to the comparison result between the gradient intensity and the preset threshold, screen out the singular value feature quantities whose gradient intensity exceeds the threshold. These feature quantities are defined as target singular value feature quantities, which are the basis for constructing the low-rank matrix subsequently. Through logical judgment or matrix operations, record the singular value feature quantities whose gradient intensity exceeds the threshold and their corresponding indexes. For example, in the field of medical and health, screen out the important feature quantities related to disease classification; in the financial field, screen out the important transaction feature quantities related to market risk prediction.

[0072] Each target singular value feature quantity corresponds to a singular vector. Extracting these corresponding singular vectors is to ensure that the initialization of the low-rank matrix has a directionality and can better adapt to the optimization process.

[0073] Using the indexes of the selected target singular value feature quantities, extract the corresponding vectors from the singular vector matrix obtained by the previous decomposition and store them in association with the singular value feature quantities. For example:

[0074] In the field of medical and health, extract the vectors used to characterize specific disease characteristics;

[0075] In the financial field, extract the vectors used to capture market trading patterns.

[0076] By analyzing the intensity of gradient information and screening the target singular value feature quantities and corresponding singular vectors, the task relevance in the process of constructing the low-rank matrix is significantly improved. By setting a preset threshold, the singular value feature quantities with less influence on the optimization process are filtered out, thereby reducing the interference of redundant information and enhancing the efficiency of constructing the low-rank matrix and the adaptability of the model.

[0077] S50, construct a bypass low-rank matrix based on the target singular value feature quantities and the target singular vectors;

[0078] In this embodiment, the target singular value feature quantities and the target singular vectors are the core outputs after the screening operation, and they need to be integrated into the basic representation of the low-rank matrix. The singular value feature quantities reflect the importance weights of the matrix, and the singular vectors provide directional information. The combination of the two constitutes the core structure of the low-rank matrix.

[0079] According to the screened target singular value feature quantities and the corresponding target singular vectors, use the formula:

[0080]

[0081] Generate the low-rank matrix W main . Among them, U [:,:r] is the left singular vector of the first r columns of the original weight matrix, providing the row direction characteristics of the matrix; Σ [:r] is the diagonal singular value matrix, containing the first r singular values with the highest task relevance; is the transpose of the first r right singular vectors of the original weight matrix, providing the column direction characteristics of the matrix.

[0082] The generated low-rank matrix W main is the matrix representation with the strongest task relevance and participates in the subsequent fine-tuning process. In the field of medical and health, this matrix can be used in disease classification tasks and associated with key symptoms or diagnostic indicators; in the financial field, this matrix can reflect the characteristics associated with high volatility risks in market transactions.

[0083] The target singular value feature quantities can be used as weights to weight the column vectors of the feature matrix. The purpose of weighting is to enhance the importance of the singular value feature quantities and make the low-rank matrix more accurately reflect the task relevance. After normalizing the target singular value feature quantities, they are used as weights to multiply with the column vectors of the target singular vectors one by one to generate the weighted target feature matrix. For example:

[0084] In the field of medical and health, adjust the contribution of the corresponding vectors according to the singular value weights of different diseases;

[0085] In the financial field, adjust the influence of the trading feature vectors according to the singular value weights under different market conditions.

[0086] To avoid numerical instability or overfitting in the feature matrix, regularization constraints such as L2 regularization can be introduced to limit the range of parameters in the matrix.

[0087] By adding a regularization term, the numerical range of the target feature matrix is restricted. The regularization constraint can be added through the loss function in the optimization framework or completed through numerical operations on the matrix. For example:

[0088] In the field of healthcare, add regularization to the disease diagnosis matrix to ensure that the low-rank matrix adapts to different diagnostic conditions;

[0089] In the financial field, add regularization to the market prediction matrix to enhance the robustness of the matrix to different market fluctuations.

[0090] The rank of the low-rank matrix determines its complexity. By trimming the rank, the representation ability of the matrix can be compressed, the computational overhead can be reduced, and the main features can be retained.

[0091] Use matrix decomposition techniques to trim the rank of the feature matrix and retain the features with the highest rank values. This can be achieved through SVD decomposition or truncating the singular value matrix. For example:

[0092] In the field of healthcare, by trimming the rank of the disease diagnosis matrix, retain the features that contribute more to the main diagnosis tasks;

[0093] In the financial field, by trimming the rank of the market prediction matrix, retain the features that contribute more to the main market fluctuations.

[0094] After completing the above operations, generate the final bypass low-rank matrix as an optimized representation structure related to the task, providing input for the subsequent fine-tuning of the model. Store the weighted, regularized, and trimmed feature matrix as the bypass low-rank matrix in the model structure to replace part of the weights or assist in the optimization process.

[0095] By integrating the target singular value feature quantity and the target singular vector, and constructing the bypass low-rank matrix through weighted, regularization constraint, and rank trimming operations, the task adaptability and optimization efficiency of the low-rank matrix are significantly improved. The generated bypass low-rank matrix not only reduces the adjustment range of model parameters but also enhances the robustness of the model to complex domain tasks.

[0096] S60, freeze the original weight matrix of the model to be optimized, and based on the bypass low-rank matrix, adjust the parameters of the model to be optimized to obtain the optimized target model;

[0097] In this embodiment, freezing the original weight matrix of the model to be optimized is to maintain the stability of the general characteristics of the model during the subsequent fine-tuning process. The specific operation is to define the residual matrix Wresidual to identify the partial weights to be frozen, and this residual matrix is generated from the part of the original weight matrix that is not selected as the bypass low-rank matrix.

[0098] Using the formula:

[0099]

[0100] generate a residual matrix, where:

[0101] U [:,r+1:] is the left singular vector starting from the (r + 1)-th column; Σ [r+1:] is the diagonal singular value starting from the (r + 1)-th singular value; is the transpose of the right singular vector starting from the (r + 1)-th column.

[0102] Set the freezing flag of W residual through the deep learning framework (such as requires_grad = False in PyTorch or trainable = False in TensorFlow) to ensure that the residual matrix remains unchanged during subsequent optimization. In the field of healthcare, this operation can ensure the stability of the model for basic diagnostic features (such as general symptom embeddings); in the financial field, it can retain the ability to capture general market volatility patterns.

[0103] The bypass low-rank matrix W main is the core part of model fine-tuning and needs to be loaded into the model to be optimized and initialized to an adjustable state. The purpose of initialization is to provide a task-related weight starting point for parameter adjustment.

[0104] Insert W main directly into the model to be optimized. The initialization operation can use random initialization or be constructed from the selected target singular value eigenquantities and vectors. For example: in the field of healthcare, the bypass low-rank matrix can be initialized as the weights of specific disease features; in the financial field, the bypass low-rank matrix can be initialized as the weights of specific market behaviors.

[0105] By inputting data related to the target task, generate the prediction results of the model, and calculate the loss value using the loss function defined for the target task. The loss value quantifies the deviation between the model's prediction results and the target expectations. Input the target task input data (such as patient data in healthcare or market transaction data in the financial field) into the optimized model, and perform forward propagation to calculate the prediction results. Define the task-related loss function (such as cross-entropy loss for classification tasks or mean squared error for regression tasks) and calculate the prediction error of the current model.

[0106] The automatic differentiation module is used to calculate the gradient of the loss function to the bypass low-rank matrix, and the parameters of the bypass low-rank matrix are adjusted according to the gradient. This process optimizes the adaptability of the model through multiple iterations.

[0107] Freeze W residual The parameters make the gradient update only act on W main The parameters of the bypass low-rank matrix are iteratively adjusted based on the calculated gradients through the framework’s optimizer (such as Adam or SGD) until convergence to the expected performance.

[0108] Example description: In the disease classification task, freezing the parameters of the general embedding layer (such as common symptom embedding) retains the model's adaptability to common features, and by optimizing the bypass low-rank matrix, the model's classification ability for specific diseases is improved. After preprocessing the patient data, input the model, such as medical images for screening specific diseases, and bypass the low-rank matrix through gradient updates to enhance the model's recognition of specific image patterns.

[0109] In the market risk prediction task, the basic prediction weights are frozen to maintain adaptability to the overall market trend, and the model's predictive ability for specific risk patterns is enhanced by optimizing the bypass low-rank matrix.

[0110] By freezing the original weight matrix and optimizing the bypass low-rank matrix, the general characteristics of the model are retained, while efficient fine-tuning for tasks is achieved, significantly improving the performance and generalization ability of the model in complex domain tasks. By limiting the parameter adjustment range, the computational overhead is reduced, and the overfitting problem caused by excessive modification of the original model is avoided.

[0111] S70, processing the target task based on the optimized target model to generate an execution result of the target task.

[0112] In this embodiment, before processing the target task, the input data needs to be standardized and preprocessed to remove redundant features, unify the data format, and enhance the model's adaptability to different data distributions.

[0113] According to the data characteristics of the target task, select the appropriate standardization method, such as normalization, mean variance scaling or category encoding. In the field of healthcare, patient data (such as medical history, medical images) can be normalized to reduce sampling differences in different hospital equipment. In the financial field, market transaction data (such as stock prices, trading volume) can be normalized in time series to remove the impact of short-term fluctuations.

[0114] The standardized preprocessed data is input into the optimized target model, and the prediction results are generated through the forward propagation of the model.

[0115] Call the inference interface of the optimized target model and load the standardized input data into the input layer of the model. In the field of healthcare, input the processed image data into the classification model to generate disease prediction results. In the financial field, input the processed time series data into the risk prediction model to generate future market volatility prediction values.

[0116] According to the requirements of the target task, perform post-processing operations on the prediction results generated by the model, such as threshold screening, class mapping, or result combination, to obtain the final available output.

[0117] Set the post-processing logic according to the task requirements. For example: in the field of healthcare, perform confidence screening on the prediction results of the disease classification model and output classification labels with high confidence; in the financial field, perform threshold analysis on the results of the risk prediction model to generate buy or sell signals.

[0118] Standardized preprocessing enhances the robustness and adaptability of the model to input data, ensuring that data from different sources or formats can be efficiently processed by the model. Post-processing operations optimize the prediction results according to task requirements, improving the usability and accuracy of the final output results.

[0119] The present invention relates to the fields of artificial intelligence technology, healthcare, and fintech, and discloses a task processing method based on model parameter adjustment, including: collecting an auxiliary benchmark data set and extracting the original weight matrix, decomposing it into multiple singular value feature quantities and corresponding singular vectors, screening the target singular value feature quantities and corresponding target singular vectors based on gradient information, and constructing a bypass low-rank matrix. Freeze the original weight matrix of the model to be optimized, adjust the model parameters by fine-tuning the bypass low-rank matrix, generate an optimized target model, and use the optimized target model to process tasks to generate execution results. The present invention screens singular value feature quantities and corresponding vectors based on gradient information, constructs a low-rank matrix related to task characteristics, reduces the influence of the randomness of low-rank matrix initialization on the optimization process; freezes the original weight matrix and only fine-tunes the bypass low-rank matrix, reducing the computational overhead and improving the model optimization efficiency and generalization ability.

[0120] In one embodiment, the above S20 includes:

[0121] S201, perform dimensionality screening on the rows and columns of the original weight matrix to obtain an optimized weight matrix;

[0122] S202, perform singular value decomposition operation on the optimized weight matrix to obtain multiple singular value feature quantities;

[0123] S203, perform normalization processing on each singular value feature quantity to generate a singular vector corresponding to each singular value feature quantity.

[0124] In this embodiment, dimension screening refers to screening out rows and columns with high task relevance in the original weight matrix according to preset rules to reduce irrelevant information and optimize the matrix size. The basis for screening can be the distribution characteristics of the data or task relevance. For example, task-irrelevant features can be screened out through statistical analysis.

[0125] Row dimension screening: Calculate the statistical characteristics (such as mean, variance, or sparsity) for each row in the original weight matrix. If the mean of a certain row is close to zero, the variance is lower than the set threshold, or the sparsity is too high (the proportion of non-zero elements is too low), it is considered to have low relevance to the target task and is deleted.

[0126] Column dimension screening: Conduct task relevance analysis on each column in the original weight matrix. For example, calculate the correlation coefficient or significance value between it and the output of the target task. If the relevance of a certain column is lower than the preset threshold, its contribution is considered limited and it is deleted. Matrix operation tools (such as NumPy or Pandas) can be used to screen the matrix row by row and column by column. Optimize the matrix after compression storage to avoid redundant memory occupation.

[0127] Singular value decomposition decomposes the weight matrix into a set of specific vectors and eigenvalues to capture the main characteristics of the matrix. The singular value feature quantity reflects the importance of the matrix in different directions. Through this decomposition operation, key features related to the task can be extracted.

[0128] Input the optimized weight matrix and use the singular value decomposition method to decompose it into task-related principal components and secondary components. The decomposition steps can be completed through existing computing tools. For example, use the SVD function provided in the linear algebra library (such as scipy.linalg.svd in Python or the built-in SVD function in MATLAB). When selecting the singular value feature quantity, select the part with higher importance according to the task requirements. The specific criteria can be based on the absolute value size of the singular value or its cumulative contribution rate.

[0129] Singular value extraction: Arrange the singular values in descending order and truncate the part where the cumulative contribution rate reaches a certain proportion for subsequent optimization.

[0130] Parallelized computing: For the decomposition of large-scale matrices, a distributed computing framework (such as Spark or Dask) can be adopted to improve the computing efficiency.

[0131] Normalization processing adjusts the range of the singular value feature quantity to make its characteristics more concentrated, so that subsequent operations are more stable. The normalized singular values are used to generate the corresponding singular vectors, thereby forming matrix characteristics with strong task relevance.

[0132] Perform range scaling on each singular value feature quantity so that it falls within the interval [0, 1] or other preset ranges, avoiding gradient instability caused by excessive differences in numerical ranges. Based on the normalized singular values, combined with the optimized matrix, generate the corresponding singular vectors to describe the characteristics of the matrix in a certain direction. During the normalization process, truncate small values to prevent floating-point precision issues.

[0133] The normalize function in NumPy or manual implementation of linear scaling can be used. The normalization results and singular vectors should be stored in a compressed format to ensure fast invocation in subsequent steps.

[0134] In this embodiment, by screening, decomposing, and normalizing the original weight matrix, features with strong task relevance are extracted, irrelevant and redundant information is removed, and the computational complexity is significantly reduced. While retaining the core features, the stability and accuracy of subsequent optimization are enhanced, laying a solid foundation for the efficient processing of tasks.

[0135] In one embodiment, the above S201 includes:

[0136] S2011, analyze the mean, variance, or non-zero element ratio of each row in the original weight matrix;

[0137] S2012, delete the rows in the original weight matrix with a mean of zero, a variance lower than the preset variance threshold, or a non-zero element ratio lower than the preset ratio value to obtain a preliminarily optimized weight matrix;

[0138] S2013, analyze the variance of each column in the original weight matrix or the correlation between the original weight matrix and the target task;

[0139] S2014, delete the columns in the preliminarily optimized weight matrix with a variance lower than the preset threshold or a correlation with the target task lower than the preset correlation threshold to obtain the finally optimized weight matrix.

[0140] In this embodiment, analyzing the statistical characteristics of each row in the original weight matrix is to evaluate the importance of the rows and perform subsequent screening. The mean, variance, and non-zero element ratio are common indicators for measuring the information content and importance of rows.

[0141] Calculate the mean: Statistically calculate the mean of all elements in each row to determine whether the row contains valid information. If the mean is zero, it means that the data in this row makes no contribution to the task and can be deleted.

[0142] Calculate the variance: Statistically calculate the variance of each row to evaluate the fluctuation of the data distribution. The smaller the variance, the more stable or redundant the row data, and it may have less impact on the task.

[0143] Calculate the proportion of non-zero elements: Count the proportion of non-zero elements in each row to determine sparsity. Rows with excessive sparsity usually have low information content and can be deleted.

[0144] Using a matrix operation library (such as NumPy or Pandas), call functions to directly calculate the row mean, variance, and non-zero proportion.

[0145] Deleting low-contribution or irrelevant rows based on the analysis results is a key step in optimizing the matrix size. Screening rules:

[0146] If the row mean is zero, delete it directly;

[0147] If the row variance is below a preset threshold, it indicates that the data lacks volatility, so delete it;

[0148] If the proportion of non-zero elements in a row is below a preset value, it means the row is sparse, so delete it.

[0149] Mark the rows that meet the conditions through boolean screening for unified deletion. Perform row deletion operations in the matrix operation tool by combining the conditional screening function (such as boolean indexing in NumPy or the drop method in Pandas).

[0150] The analysis of columns is to evaluate the importance of columns in the target task and select the columns that contribute more to the task.

[0151] Analysis of variance: Count the variance of each column to evaluate the degree of data volatility. Columns with low variance may not contribute significantly to the target task.

[0152] Correlation analysis: Calculate the correlation between each column and the target task, such as using the Pearson correlation coefficient or information gain. Columns with low correlation indicate that they have little impact on the task objective and can be deleted.

[0153] Use statistical analysis tools (such as Scipy or Statsmodels) to calculate the variance and correlation and generate analysis results.

[0154] Delete columns with no contribution or low contribution to further optimize the weight matrix and ensure that the remaining columns are the core part related to the task.

[0155] Screening rules:

[0156] If the column variance is below a preset threshold, it indicates insufficient volatility, so delete it;

[0157] If the correlation between the column and the target task is below a preset threshold, it indicates insufficient contribution to the task, so delete it.

[0158] Mark the columns to be deleted through boolean screening to avoid accidental deletion. Combine the matrix operation library (such as NumPy or Pandas) to perform conditional screening and deletion operations on the columns.

[0159] Example illustration: In a disease diagnosis task, each row of the disease feature matrix (such as symptom records of different patients) is analyzed, symptom data lacking volatility or having no diagnostic significance is deleted, and further feature indicators with a relatively high correlation with the disease diagnosis result are screened out to obtain an optimized feature matrix. In a market prediction task, each column of the trading data matrix (such as price fluctuations of different stocks) is analyzed, stock data with low trading volume or insufficient volatility is deleted, and trading features closely related to market fluctuations are screened out to obtain an optimized market prediction matrix.

[0160] In this embodiment, through row and column screening of the original weight matrix, redundant information and characteristics irrelevant to the task are removed, significantly reducing the computational overhead and memory usage. The optimized weight matrix not only has a smaller size but also retains the core characteristics related to the task, providing an efficient basis for subsequent singular value decomposition and model optimization.

[0161] In one embodiment, the above S30 includes:

[0162] S301, input the auxiliary reference data set item by item into the model to be optimized, and obtain the feature response matrix output by the model to be optimized;

[0163] S302, based on the auxiliary reference data set and the feature response matrix, determine the loss function related to the target task;

[0164] S303, based on the loss function, perform gradient analysis on the original weight matrix through an automatic differentiation module to generate gradient information for each singular value feature quantity.

[0165] In this embodiment, the purpose of inputting the auxiliary reference data set is to activate different weights of the model to be optimized and extract its characteristic responses to the data. Processing the data set item by item ensures that the responses of the model to each sample are separately recorded and mapped into the feature space. The model to be optimized converts the input samples into feature responses in the intermediate layer through the forward propagation process, generating a matrix, where each row represents the feature activation value of a sample and the columns represent the feature dimensions of the model.

[0166] The forward propagation of the model calculates the activation values of each layer of neurons through the input samples. In a deep learning framework, functions such as model.predict in TensorFlow or model(input) in PyTorch can be used to complete the item-by-item input of the data set. The feature response matrix is obtained by reading the output of the intermediate layer or the final layer of the model and integrating the activation values into a matrix structure.

[0167] Efficiency can be optimized by batch loading data. For example, dividing the data into small batches for processing can reduce memory occupancy while increasing the computing speed. During implementation, it should be ensured that the feature response matrix is arranged in the sample order for convenient subsequent operations.

[0168] The role of the loss function is to measure the deviation between the predicted values of the model and the expected values of the target task, thereby guiding the optimization of the model. In this step, using the target labels of the auxiliary benchmark dataset and the output results of the feature response matrix, the loss function is defined through differential analysis. For different task types, the definition methods of the loss function may vary.

[0169] For classification tasks, the cross-entropy loss function can be used to minimize the probability distribution difference between the labels and the feature responses of the model. For regression tasks, the mean squared error loss function can be adopted to measure the numerical deviation between the predicted values and the true values. During implementation, use the built-in loss function module of the framework, such as torch.nn.CrossEntropyLoss or torch.nn.MSELoss in PyTorch.

[0170] The column characteristics in the feature response matrix need to be aligned with the task labels to ensure that the loss function can accurately quantify the prediction error. Custom loss functions can also adjust the weights according to task requirements. For example, increasing the influence weight of important features can optimize specific tasks more precisely.

[0171] Based on the defined loss function, the automatic differentiation module calculates the gradient values of each parameter in the model through the backpropagation process. These gradient values represent the sensitivity of the model parameters to the loss, that is, the impact of the change of model parameters on the final loss. Through gradient analysis, the importance of singular value features can be identified.

[0172] The automatic differentiation module will record the forward propagation computational graph of the model and calculate the gradient of each parameter during backpropagation. By freezing the parameters of other parts in the model, only the gradient of the singular value features of the original weight matrix is calculated. The magnitude and direction of the gradient can quantify the contribution degree of the singular value features to the task objective.

[0173] During implementation, use the automatic differentiation function of the deep learning framework (such as torch.autograd.grad in PyTorch or GradientTape in TensorFlow) to efficiently calculate the gradient. To ensure the stability of gradient analysis, gradient clipping can be introduced during the calculation process to prevent unstable problems caused by too large or too small gradient values.

[0174] The generated gradient information should be stored as a gradient table, which records the identifiers of each singular value feature quantity and their corresponding gradient values. This gradient information will provide a scientific basis for screening singular value feature quantities with higher task relevance in subsequent steps.

[0175] In this embodiment, by inputting the auxiliary reference data set item by item to generate a feature response matrix, the intermediate characteristic performance of the model is extracted, and the response of the model to the task data is accurately recorded. Combining the feature response matrix with the loss function definition of the target task can effectively quantify the performance error of the model. The introduction of the automatic differentiation module greatly improves the efficiency and accuracy of gradient calculation. By gradient analysis, the contribution degree of singular value feature quantities to the task is determined, providing a scientific basis for subsequent optimization.

[0176] In one embodiment, the above S50 includes:

[0177] S501, divide the target singular vectors according to a preset grouping method and splice them in column-first order to form a feature matrix;

[0178] S502, based on the target singular value feature quantities, perform weighted processing on the column vectors of the feature matrix, and introduce regularization constraints during the weighted processing to generate a weighted feature matrix;

[0179] S503, perform a rank pruning operation on the weighted feature matrix, and retain the core feature information of the feature matrix by limiting the maximum rank value to compress the representation dimension of the feature matrix;

[0180] S504, use the pruned and compressed feature matrix as the initial representation of the bypass low-rank matrix.

[0181] In this embodiment, the division of the target singular vectors classifies the vectors through a preset grouping rule, such as grouping based on task relevance or the numerical range of the vectors. This grouping method can reduce the task-irrelevant information in the matrix and improve the pertinence of the feature matrix. After the grouping is completed, the singular vectors are spliced into a matrix in column-first order. Such an arrangement can ensure the unified numerical expression of the matrix in the column direction and is convenient for subsequent calculation and processing.

[0182] In the specific implementation process, the target singular vectors can be sorted according to their numerical sizes in a specific dimension, and then sliced according to the grouping rule. Each group of vectors after slicing is spliced column by column into a new feature matrix. After the splicing is completed, it is necessary to check the dimension and numerical range of the feature matrix to ensure that it meets the input requirements of subsequent processing steps.

[0183] In the weighting process, the target singular value feature quantity is used as a weight parameter to adjust the importance of column vectors in the feature matrix. The weight assignment is usually based on the magnitude of the singular values. The larger the singular value, the stronger its relevance to the task, so a higher weight is assigned. The role of the regularization constraint is to prevent instability in matrix calculations caused by weights being too large or too small.

[0184] The weighting process can be completed by directly multiplying each column vector by the corresponding singular value. The regularization constraint can introduce L2 or L1 regularization methods to limit the range of weight values. After the weighting operation, it is necessary to check whether there are outliers in the values of the weighted feature matrix, such as cases where the weights deviate too much or are close to zero. If anomalies occur, they can be adjusted by threshold clipping or reassigning weights.

[0185] Rank clipping is a process of performing low-rank approximation on the feature matrix, aiming to reduce the complexity of the matrix while retaining its most important task-related information. The setting of the maximum rank value can be based on task requirements, for example, determining a reasonable clipping threshold according to the cumulative contribution rate or matrix performance.

[0186] In actual operation, the weighted feature matrix can be subjected to singular value decomposition (SVD), and then the first several singular values and the corresponding singular vectors are selected to reconstruct the matrix. The clipped feature matrix not only has a smaller dimension but also retains the most valuable feature representation for the target task. After clipping, it is necessary to verify the new matrix to ensure that its numerical distribution is reasonable and meets the requirements of the task.

[0187] The clipped and compressed feature matrix is used as the initial state of the bypass low-rank matrix, providing a parameter basis with strong task relevance for subsequent model optimization. In this process, the format and numerical range of the feature matrix need to be consistent with the input layer format of the model to be optimized.

[0188] Specific operations include loading the compressed feature matrix to a specified position in the model, such as as the weight parameter of the bypass channel or as the weight of an independent sub-network during fine-tuning. After loading, the parameters of the bypass low-rank matrix are marked as adjustable, providing a basis for subsequent task fine-tuning.

[0189] This embodiment can effectively reduce the redundant information in the original matrix and enhance its adaptability to specific tasks by constructing a bypass low-rank matrix. The weighting and clipping operations of the feature matrix not only optimize the numerical stability of the matrix but also improve the computational efficiency. The finally constructed bypass low-rank matrix provides high-quality parameter initialization for subsequent model optimization, significantly improving the task processing ability of the model.

[0190] In one embodiment, the above S60 includes:

[0191] S601, freeze the gradient calculation of all parameters of the original weight matrix in the model to be optimized;

[0192] S602, load the bypass low-rank matrix into the model to be optimized, and initialize the parameters of the bypass low-rank matrix to an adjustable state;

[0193] S603, determine the loss value of the model to be optimized based on the input data of the target task;

[0194] S604, generate the gradient of the bypass low-rank matrix based on the loss value through an automatic differentiation module;

[0195] S605, iteratively update the adjustable parameters of the bypass low-rank matrix according to the gradient to generate an optimized target model.

[0196] In this embodiment, freezing the gradient calculation of the original weight matrix is to maintain the stability of the basic part of the model while focusing on the optimization and adjustment of the bypass low-rank matrix. This operation ensures that during the fine-tuning process, only the parameters of the task-related bypass low-rank matrix are adjusted, without destroying the general capabilities of the pre-trained model.

[0197] In the deep learning framework, mark the original weight matrix of the model to be optimized as untrainable. For example, in PyTorch, the parameters can be frozen by setting requires_grad = False. During the forward and backward propagation of the model, the frozen weight matrix will not be updated, thereby reducing the computational overhead and improving the optimization efficiency.

[0198] The loading and initialization of the bypass low-rank matrix are key steps in the optimization process. The loading operation inserts the constructed bypass low-rank matrix into the specified position of the model, and the initialization sets its trainable parameters for subsequent gradient updates.

[0199] The bypass low-rank matrix can be inserted into the model as an independent weight module. For example, bypass channels can be added to the middle layer or output layer of the model. Initializing to an adjustable state means that the parameters of the bypass low-rank matrix are allowed to participate in gradient calculation and optimization. For example, in TensorFlow, this can be achieved by setting trainable = True. After loading, the matrix needs to be verified to ensure that its dimensions and numerical range are consistent with the existing parameters of the model.

[0200] The input data of the target task is used to activate the model to be optimized and generate prediction results. By comparing the prediction results with the expected values of the target task, the loss value of the model is calculated. The loss value is an important indicator to measure the performance of the model and is also the basis for updating parameters during the optimization process.

[0201] After the input data is standardized, it is input into the model one by one to generate prediction results. Select an appropriate loss function according to the type of the target task. For example, the cross-entropy loss function is used for classification tasks, and the mean squared error loss function is used for regression tasks. The calculation of the loss value is completed by comparing the output of the model with the target label. For example, in PyTorch, loss = criterion(output, target).

[0202] The automatic differentiation module is the core tool for gradient calculation. By recording the computational graph of the forward propagation, the automatic differentiation module can calculate the partial derivatives of the loss value with respect to the parameters of the bypass low-rank matrix during the backward propagation. These gradient values are used to optimize the bypass low-rank matrix to make it more suitable for the target task.

[0203] Record the computational graph during the forward propagation to ensure that the parameters of the bypass low-rank matrix can participate in the gradient calculation during the backward propagation. Calculate the gradient through the automatic differentiation function of the deep learning framework. For example, in PyTorch, use loss.backward() to generate the gradient. After the gradient calculation is completed, clip or regularize the gradient values to avoid the problems of gradient explosion or disappearance.

[0204] Use the calculated gradient to update the parameters of the bypass low-rank matrix, gradually reducing the loss value. The optimized target model is a model that adapts to the characteristics of the target task while keeping the original weight matrix stable.

[0205] Use an optimizer (such as SGD or Adam) to update the parameters of the bypass low-rank matrix. In each iteration, adjust the parameter values according to the gradient to make the model gradually converge to the optimal state. During the iteration process, monitor the change trend of the loss value and terminate the optimization according to the preset stopping conditions (such as the loss value no longer decreases significantly or reaches the maximum number of iterations). After the update is completed, verify the performance of the optimized target model to ensure that its effect on the target task meets the requirements.

[0206] This embodiment avoids the destruction of the general characteristics of the model by freezing the original weight matrix, and at the same time optimizes the bypass low-rank matrix to enhance the adaptability of the model to the target task. The introduction of the loss value provides a clear optimization goal for the gradient calculation, and the automatic differentiation module significantly improves the gradient calculation efficiency. Finally, the optimized target model generated by iterative update has higher task processing ability and generalization performance.

[0207] In one embodiment, the above S70 includes:

[0208] S701, perform standardized preprocessing on the input data of the target task to generate preprocessed input data;

[0209] S702. Input the preprocessed input data into the optimized target model, and generate a prediction result based on the input data according to the optimized target model;

[0210] S703. Perform a post-processing operation on the prediction result to generate an execution result of the target task.

[0211] In this embodiment, the input data of the target task may come from multiple heterogeneous data sources and have different formats, distributions, or noise characteristics. Through a series of rules, the standardized preprocessing unifies the input data into the format and distribution required by the optimized target model, ensuring that the data can be efficiently processed by the model. The standardized preprocessing includes the following specific operations: cleaning the input data to remove obvious outliers; supplementing missing values to avoid interruption of the calculation process; normalizing continuous numerical data and scaling it to a fixed interval range; encoding and transforming discrete categorical data, for example, using one-hot encoding to represent categorical features as numerical values.

[0212] In the implementation process, the batch processing method can be used to divide the input data into several batches for preprocessing to reduce memory consumption and improve processing efficiency. For each batch of data, it is necessary to ensure that its feature dimensions are exactly the same as the model input requirements. If the features of the input data change dynamically, a verification module can be added to check the data for consistency.

[0213] The preprocessed data is loaded into the optimized target model, and a prediction result is generated through the forward propagation process of the model. In this process, the model will perform feature extraction and pattern recognition on the input data according to its optimized weights and structure, and finally output a predicted value. This step requires ensuring that the format of the input data exactly matches the requirements of the input layer of the optimized target model.

[0214] The specific steps to implement this technical feature include encapsulating the preprocessed input data as an input tensor of the model, calling the inference function of the deep learning framework (such as model(input) in PyTorch or model.predict in TensorFlow) for forward propagation calculation, and storing the generated prediction results in the specified data structure in the sample order.

[0215] When performing model prediction, the batch inference mode can be selected to improve processing efficiency. For example, divide the input data into small batches and input them into the model successively, and the prediction results of each batch are merged into a complete output result matrix in order.

[0216] Post - processing operations aim to adjust the prediction results generated by the model into the final output form required for the target task, so as to improve the interpretability of the results and the task adaptability. Post - processing operations usually include the following forms: sorting and filtering the prediction results according to confidence levels and retaining high - confidence results; mapping the numerical labels of the prediction results to convert them into descriptions of specific categories or actual meanings; performing logical operations or rule combinations on the prediction results to adapt to the requirements of multi - task scenarios; optimizing the format of the results and outputting them in the specific format required by the target task, such as reports, visualization charts or interface return values.

[0217] During the implementation process, it is necessary to flexibly adjust the post - processing rules according to the requirements of the target task. For example, in a classification task, a confidence threshold can be set to screen the predicted categories; in a multi - task scenario, multiple prediction results can be combined through logical operations. The final execution results need to be verified to ensure that they meet the business requirements of the target task.

[0218] Example illustration: In the disease diagnosis task in the field of medical and health, the input data are the physical examination indicators of patients, including blood indicators, imaging data and medical history information. Standardized pre - processing will process these data separately. For example, normalizing the blood indicators, converting the imaging data into pixel tensors, and encoding the medical history information using a natural language processing model. The pre - processed data are input into the optimized target model, and the model combines the multi - modal information of the patient for comprehensive diagnosis, generating a disease risk score or a list of potential diseases. In the post - processing stage, the diagnosis model will screen out high - risk diseases according to the disease risk score and map the numerical score to the disease name and risk level. For example, mapping the risk value of 0.85 predicted by the model to "Diabetes, high risk". The final result will generate a comprehensive report listing the diagnosis conclusion, main risk factors and recommended follow - up examination items.

[0219] In the market prediction task in the financial field, the input data are the historical prices, trading volumes of stocks and related macro - economic indicators. The standardized pre - processing steps will normalize the price and trading volume data and perform categorical encoding on the macro - economic indicators. For example, encoding the economic policy category as a numerical feature. The pre - processed data are input into the optimized target model, and the model uses historical patterns and feature extraction capabilities to predict future stock price trends. In the post - processing stage, the model will convert the predicted price change range into specific operation suggestions, such as "Increase by 5%, it is recommended to buy". For the prediction results of multiple stocks, investment portfolio suggestions can be generated through logical rules combination. For example, screening out stocks with an increase rate greater than 3% and generating a ranking list. The final execution result will be presented in the form of a visualization report, including stock names, predicted increase rates and investment suggestions.

[0220] Through standardized preprocessing in this embodiment, the input data is unified and optimized, eliminating the fluctuations in model performance caused by differences in data distribution. The optimized target model accurately generates prediction results through forward propagation, and the post-processing operation converts the results into an intuitive business output form, significantly improving the adaptability and usability of the model in actual tasks.

[0221] In one embodiment, a task processing device based on model parameter adjustment is provided, and the task processing device based on model parameter adjustment corresponds one-to-one with the task processing method based on model parameter adjustment in the above embodiment. Refer to Figure 3 , Figure 3 is a schematic diagram of the functional modules of a preferred embodiment of the task processing device based on model parameter adjustment of the present invention. Data processing module 10, matrix decomposition module 20, gradient analysis module 30, feature screening module 40, low-rank matrix construction module 50, model optimization module 60, and task processing module 70. The detailed description of each functional module is as follows:

[0222] The data processing module 10 is used to collect an auxiliary reference data set and extract the original weight matrix from the model to be optimized;

[0223] The matrix decomposition module 20 is used to decompose the original weight matrix into a plurality of singular value feature quantities and the corresponding singular vectors for each singular value feature quantity;

[0224] The gradient analysis module 30 is used to determine the gradient information of the plurality of singular value feature quantities based on the auxiliary reference data set;

[0225] The feature screening module 40 is used to determine the target singular value feature quantity whose gradient information intensity exceeds a preset threshold and the corresponding target singular vector of the target singular value feature quantity;

[0226] The low-rank matrix construction module 50 is used to construct a bypass low-rank matrix based on the target singular value feature quantity and the target singular vector;

[0227] The model optimization module 60 is used to freeze the original weight matrix of the model to be optimized and perform parameter adjustment on the model to be optimized based on the bypass low-rank matrix to obtain an optimized target model;

[0228] The task processing module 70 is used to process the target task based on the optimized target model to generate the execution result of the target task.

[0229] In one embodiment, the matrix decomposition module 20 is specifically used for:

[0230] Perform dimensionality screening on the rows and columns of the original weight matrix to obtain an optimized weight matrix;

[0231] Perform a singular value decomposition operation on the optimized weight matrix to obtain multiple singular value feature quantities;

[0232] Perform a normalization process on each singular value feature quantity to generate a singular vector corresponding to each singular value feature quantity.

[0233] In one embodiment, the matrix decomposition module 20 is specifically configured to:

[0234] Analyze the mean, variance, or non - zero element ratio of each row in the original weight matrix;

[0235] Delete the rows in the original weight matrix where the mean is zero, the variance is lower than a preset variance threshold, or the non - zero element ratio is lower than a preset ratio value to obtain a preliminarily optimized weight matrix;

[0236] Analyze the variance of each column in the original weight matrix or the correlation between the original weight matrix and the target task;

[0237] Delete the columns in the preliminarily optimized weight matrix where the variance is lower than a preset threshold or the correlation with the target task is lower than a preset correlation threshold to obtain the finally optimized weight matrix.

[0238] In one embodiment, the gradient analysis module 30 is specifically configured to:

[0239] Input the auxiliary benchmark dataset into the model to be optimized item by item to obtain the feature response matrix output by the model to be optimized;

[0240] Based on the auxiliary benchmark dataset and the feature response matrix, determine the loss function related to the target task;

[0241] Based on the loss function, perform gradient analysis on the original weight matrix through the automatic differentiation module to generate the gradient information of each singular value feature quantity.

[0242] In one embodiment, the low - rank matrix construction module 50 is specifically configured to:

[0243] Divide the target singular vectors according to a preset grouping method and splice them in column - first order to form a feature matrix;

[0244] Based on the target singular value feature quantities, perform weighted processing on the column vectors of the feature matrix and introduce regularization constraints during the weighted processing to generate a weighted feature matrix;

[0245] Perform a rank pruning operation on the weighted feature matrix and retain the core feature information of the feature matrix by limiting the maximum rank value to compress the representation dimension of the feature matrix;

[0246] Use the cropped and compressed feature matrix as the initial representation of the bypass low-rank matrix.

[0247] In one embodiment, the model optimization module 60 is specifically configured to:

[0248] Freeze the gradient calculation of all parameters of the original weight matrix in the model to be optimized;

[0249] Load the bypass low-rank matrix into the model to be optimized and initialize the parameters of the bypass low-rank matrix to an adjustable state;

[0250] Determine the loss value of the model to be optimized based on the input data of the target task;

[0251] Based on the loss value, generate the gradient of the bypass low-rank matrix through the automatic differentiation module;

[0252] Iteratively update the adjustable parameters of the bypass low-rank matrix according to the gradient to generate an optimized target model.

[0253] In one embodiment, the task processing module 70 is specifically configured to:

[0254] Perform standardized preprocessing on the input data of the target task to generate preprocessed input data;

[0255] Input the preprocessed input data into the optimized target model and generate a prediction result based on the input data according to the optimized target model;

[0256] Perform post-processing operations on the prediction result to generate the execution result of the target task.

[0257] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 4 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage media. The network interface of the computer device is used to communicate with an external client through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a task processing method based on model parameter adjustment.

[0258] In one embodiment, a computer device is provided. The computer device can be a client, and its internal structure diagram can be as Figure 5As shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the user side of a task processing method based on model parameter adjustment

[0259] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are realized:

[0260] Collect an auxiliary reference data set and extract the original weight matrix from the model to be optimized;

[0261] Decompose the original weight matrix into a plurality of singular value features and corresponding singular vectors for each singular value feature;

[0262] Based on the auxiliary reference data set, determine the gradient information of the plurality of singular value features;

[0263] Determine the target singular value feature whose gradient information intensity exceeds a preset threshold and the corresponding target singular vector of the target singular value feature;

[0264] Based on the target singular value feature and the target singular vector, construct a bypass low-rank matrix;

[0265] Freeze the original weight matrix of the model to be optimized, and adjust the parameters of the model to be optimized based on the bypass low-rank matrix to obtain an optimized target model;

[0266] Process the target task based on the optimized target model to generate the execution result of the target task.

[0267] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the following steps are realized:

[0268] Collect an auxiliary reference data set and extract the original weight matrix from the model to be optimized;

[0269] Decompose the original weight matrix into a plurality of singular value features and corresponding singular vectors for each singular value feature;

[0270] Based on the auxiliary reference data set, determine the gradient information of the multiple singular value feature quantities;

[0271] Determine the target singular value feature quantities whose gradient information intensity exceeds a preset threshold and the target singular vectors corresponding to the target singular value feature quantities;

[0272] Based on the target singular value feature quantities and the target singular vectors, construct a bypass low-rank matrix;

[0273] Freeze the original weight matrix of the model to be optimized, and adjust the parameters of the model to be optimized based on the bypass low-rank matrix to obtain an optimized target model;

[0274] Process the target task based on the optimized target model to generate an execution result of the target task.

[0275] It should be noted that for the functions or steps that can be realized by the above computer-readable storage medium or computer device, reference can be made to the relevant descriptions on the server side and the user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0276] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0277] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0278] It should be noted that if there are software tools or components of other companies in the embodiments of this application, they are only used for example introduction and do not represent actual use. The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A task processing method based on model parameter adjustment, characterized in that: The following steps are involved: Collect auxiliary benchmark datasets and extract the original weight matrix from the model to be optimized; Decomposing the original weight matrix into a plurality of singular value eigenvalues ​​and a singular vector corresponding to each singular value eigenvalue; Based on the auxiliary benchmark data set, determining gradient information of the plurality of singular value feature quantities; Determine a target singular value feature quantity whose intensity of gradient information exceeds a preset threshold and a target singular vector corresponding to the target singular value feature quantity; Based on the target singular value feature quantity and the target singular vector, construct a bypass low-rank matrix; Freeze the original weight matrix of the model to be optimized, and adjust the parameters of the model to be optimized based on the bypass low-rank matrix to obtain an optimized target model; The target task is processed based on the optimized target model to generate an execution result of the target task.

2. The task processing method based on model parameter adjustment according to claim 1, characterized in that: Decomposing the original weight matrix into a plurality of singular value eigenvalues ​​and a singular vector corresponding to each singular value eigenvalue, including: Performing dimension screening on the rows and columns of the original weight matrix to obtain an optimized weight matrix; Performing a singular value decomposition operation on the optimized weight matrix to obtain a plurality of singular value feature quantities; Normalization is performed on each singular value feature quantity to generate a singular vector corresponding to each singular value feature quantity.

3. The task processing method based on model parameter adjustment according to claim 2, characterized in that: The rows and columns of the original weight matrix are dimensionally screened to obtain an optimized weight matrix, including: Analyze the mean, variance or non-zero element ratio of each row in the original weight matrix; In the original weight matrix, delete the rows whose mean is zero, whose variance is lower than a preset variance threshold, or whose non-zero element ratio is lower than a preset ratio value, to obtain a preliminarily optimized weight matrix; Analyzing the variance of each column in the original weight matrix or the correlation between the original weight matrix and the target task; In the initially optimized weight matrix, columns whose variance is lower than a preset threshold or columns whose correlation with the target task is lower than a preset correlation threshold are deleted to obtain a final optimized weight matrix.

4. The task processing method based on model parameter adjustment according to claim 1, characterized in that: Determining gradient information of the plurality of singular value feature quantities based on the auxiliary benchmark data set includes: Inputting the auxiliary benchmark data set into the model to be optimized one by one to obtain a characteristic response matrix output by the model to be optimized; Determining a loss function associated with a target task based on the auxiliary benchmark dataset and the feature response matrix; Based on the loss function, the original weight matrix is ​​subjected to gradient analysis through an automatic differentiation module to generate gradient information of each singular value feature.

5. The task processing method based on model parameter adjustment according to claim 1, characterized in that: Based on the target singular value feature quantity and the target singular vector, a bypass low-rank matrix is ​​constructed, including: Dividing the target singular vectors according to a preset grouping method, and splicing them in column priority order to form a feature matrix; Based on the target singular value feature quantity, weighting is performed on the column vectors of the feature matrix, and regularization constraints are introduced in the weighting process to generate a weighted feature matrix; Performing a rank clipping operation on the weighted feature matrix, and retaining core feature information of the feature matrix by limiting a maximum rank value to compress the representation dimension of the feature matrix; The pruned and compressed feature matrix is ​​used as the initial representation of the bypass low-rank matrix.

6. The task processing method based on model parameter adjustment according to claim 1, characterized in that: Freezing the original weight matrix of the model to be optimized, and adjusting the parameters of the model to be optimized based on the bypass low-rank matrix to obtain an optimized target model, including: Freeze the gradient calculation of all parameters of the original weight matrix in the model to be optimized; Loading the bypass low-rank matrix into the model to be optimized, and initializing the parameters of the bypass low-rank matrix to an adjustable state; Determine the loss value of the model to be optimized based on the input data of the target task; Based on the loss value, generating a gradient of the bypass low-rank matrix through an automatic differentiation module; The adjustable parameters of the bypass low-rank matrix are iteratively updated according to the gradient to generate an optimized target model.

7. The task processing method based on model parameter adjustment as claimed in claim 1, characterized in that: Processing the target task based on the optimized target model to generate an execution result of the target task includes: Perform standardized preprocessing on the input data of the target task to generate preprocessed input data; Inputting the preprocessed input data into the optimized target model, and generating a prediction result based on the input data according to the optimized target model; A post-processing operation is performed on the prediction result to generate an execution result of the target task.

8. A task processing device based on model parameter adjustment, characterized in that: The task processing device based on model parameter adjustment includes: A data processing module, used to collect auxiliary benchmark data sets and extract the original weight matrix from the model to be optimized; A matrix decomposition module, used for decomposing the original weight matrix into a plurality of singular value eigenvalues ​​and a singular vector corresponding to each singular value eigenvalue; A gradient analysis module, used to determine the gradient information of the plurality of singular value feature quantities based on the auxiliary benchmark data set; A feature screening module, used to determine a target singular value feature quantity whose intensity of gradient information exceeds a preset threshold and a target singular vector corresponding to the target singular value feature quantity; A low-rank matrix construction module, used to construct a bypass low-rank matrix based on the target singular value feature quantity and the target singular vector; A model optimization module, used for freezing the original weight matrix of the model to be optimized, and adjusting the parameters of the model to be optimized based on the bypass low-rank matrix to obtain an optimized target model; The task processing module is used to process the target task based on the optimized target model and generate the execution result of the target task.

9. A computer device, characterized in that: The computer device includes a memory, a processor, and a task processing program based on model parameter adjustment stored in the memory and executable on the processor. When the task processing program based on model parameter adjustment is executed by the processor, the steps of the task processing method based on model parameter adjustment are implemented as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The storage medium stores a task processing program based on model parameter adjustment, and when the task processing program based on model parameter adjustment is executed by the processor, the steps of the task processing method based on model parameter adjustment according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Large model construction method and device, equipment and medium

    CN121168556A

  • Model processing method and device, equipment, storage medium and program product

    CN121235019A

  • Model processing methods, apparatus, equipment, storage media, and program products

    CN121235019B