A hierarchical parameter adjustment method and system for a multi-task joint large model
By fine-tuning and merging the hierarchical parameters of the large model, a multi-task joint large model is generated, which solves the problem of large model's resource consumption in multi-task processing, and achieves efficient multi-task processing and performance improvement.
Patent Information
- Application Number
- CN202411331520.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-24
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2044-09-24
AI Technical Summary
Existing large models consume high computing resources and have high time costs during multitasking, making it difficult to achieve efficient multitasking in a limited resource environment.
The hierarchical parameters of the large model are iteratively adjusted by Lora fine-tuning, and a multi-task joint model is generated. By layering the parameters of other layers, the specific layer is adjusted, the Lora parameters of different task types are merged, the multi-task joint model is generated, and the performance evaluation is performed.
The number of resources required for parameter adjustment is reduced, the efficiency of parameter adjustment is improved, and the model can handle multiple tasks at the same time, meet the requirements of multi-tasks, and improve the overall performance and iteration cycle of the model.
Smart Images

Figure CN119128111B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of model parameter adjustment, and particularly to a method and system for hierarchical parameter adjustment of a multi-task joint large model. Background Art
[0002] With the rapid development of artificial intelligence technology, especially the increasing maturity of large model technology, its application scenarios in the real world are constantly expanding and deepening. In multiple fields such as intelligent customer service, cross-language translation, and text classification, large models need to simultaneously handle multiple task challenges such as the accuracy of question answering, the effective classification of information, and efficient and accurate translation. In the face of these complex requirements, how to enable a single large model to have the ability to handle multiple tasks when the computing resource environment is limited and multiple large models cannot be deployed simultaneously has become an urgent problem to be solved. This not only requires the model to have high flexibility and scalability, but also requires a fine parameter adjustment strategy to enable the model to adaptively balance the performance requirements between tasks, so as to maximize the use of limited computing resources and achieve collaborative optimization between tasks.
[0003] However, due to the huge number of parameters in large models, their parameters require huge amounts of resources, which brings great pressure to cost control. Therefore, it is particularly important to design an efficient parameter adjustment method. This not only concerns effectively reducing the computing resources and time costs required in the training process, but is also the key to improving the overall performance of the model and accelerating the model iteration cycle. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention provides a method and system for hierarchical parameter adjustment of a multi-task joint large model to solve the problems that the existing model parameter adjustment methods consume a large amount of computing resources and have a high time cost.
[0005] The technical solution adopted by the present invention is as follows:
[0006] In the first aspect, the present invention provides a method for hierarchical parameter adjustment of a multi-task joint large model, including:
[0007] Iteratively adjust the hierarchical parameters of the large model by means of Lora fine-tuning according to the task types executed by the large model and the parameter adjustment data set, and obtain the Lora parameters corresponding to each task type; the parameter adjustment data set includes a question data set, a question-and-answer data set, and a translation data set; the task types include classification tasks, question-and-answer tasks, and machine translation tasks;
[0008] Merge the original large model parameters with the Lora fine-tuning parameters of multiple task types to generate a multi-task joint large model;
[0009] Evaluate the model performance of the multi-task joint large model according to the preset task indicators, and deploy the multi-task joint large model whose model performance evaluation results meet the multi-task requirements to the actual application environment.
[0010] Further, the method also includes the process of obtaining the parameter adjustment data set:
[0011] Collect a variety of model parameter adjustment materials and store them in the object database.
[0012] Perform data preprocessing operations on a variety of model parameter adjustment materials in the object database to obtain preprocessed materials.
[0013] Segment the preprocessed materials to obtain multiple segments of preprocessed materials.
[0014] Input each segment of the preprocessed materials into the preset large model respectively to generate multiple first question texts related to model parameter adjustment, and set the labels of the first question texts as professional questions; extract some questions from the first open-source data set as the second question texts, and set the labels of the second question texts as non-professional questions; combine the first question texts, the second question texts and the corresponding labels of the texts to obtain a question data set.
[0015] Input each segment of the preprocessed materials into the preset large model respectively to generate multiple first question-and-answer pairs related to model parameter adjustment as professional question-and-answer pairs; extract some question-and-answer pairs from the first open-source data set as non-professional question-and-answer pairs, and combine the professional question-and-answer pairs with the non-professional question-and-answer pairs to obtain a question-and-answer data set.
[0016] Input each segment of the preprocessed materials into the preset large model respectively, extract multiple sentences and translate the multiple sentences into sentences in the preset language to obtain professional translation data; extract some translation data from the second open-source data set as non-professional translation data, and combine the professional translation data with the non-professional translation data to obtain a translation data set.
[0017] Further, according to the task type executed by the large model and the parameter adjustment data set, adopt the Lora fine-tuning method to iteratively adjust the hierarchical parameters of the large model respectively to obtain the Lora parameters corresponding to each task type, including:
[0018] Judge the task type executed by the large model. If the task type is a classification task, perform hierarchical parameter Lora fine-tuning on the large model based on the question data to obtain classification Lora parameters; among them, the large model includes an embedding layer, a position encoding layer, an attention layer and a feed-forward layer.
[0019] If the task type is a question-and-answer task, perform hierarchical parameter Lora fine-tuning on the large model based on the question-and-answer data set to obtain question-and-answer Lora parameters.
[0020] If the task type is a machine translation task, then the large model is fine-tuned with hierarchical parameters Lora based on the translation dataset to obtain the translation Lora parameters.
[0021] Furthermore, the Lora fine-tuning process of the hierarchical parameters includes:
[0022] First, freeze other layers of the large model and perform parameter Lora fine-tuning on the embedding layer. Then, freeze other layers of the large model except the position encoding layer and perform parameter Lora fine-tuning on the position encoding layer to assign a position encoding to the position of each element in the input sequence. Freeze other layers of the large model except the attention layer and perform parameter Lora fine-tuning on the attention layer to calculate the similarity between each element in the input sequence, assign a weight to each element, and capture the dependencies and importance between different elements according to the weights of the elements. Next, freeze other layers of the large model except the feed-forward layer and perform parameter Lora fine-tuning on the feed-forward layer. Map the input sequence to a higher-dimensional space than the original dimension through a linear transformation, apply the ReLU activation function, and map the data back to the original dimension through a linear transformation. Finally, freeze other layers of the large model except the output layer and perform parameter Lora fine-tuning on the output layer to output the final result generated by the model.
[0023] Furthermore, the parameter Lora fine-tuning process is specifically as follows: Fix the weight parameters of the original large model, define two low-rank matrices A and B as new parameters to participate in the operation, and sum the results of the two links as the output of this layer in the large model. When updating the hierarchical parameters, calculate the loss function according to the difference between the model's answer and the correct answer, and update the low-rank matrices A and B through the backpropagation algorithm to minimize the loss function. Iteratively update the hierarchical parameters until the performance of the large model reaches the predetermined standard or the training reaches the preset number of iterations.
[0024] Furthermore, the merging of the original large model parameters and the Lora fine-tuning parameters of multiple task types to generate a multi-task joint large model includes:
[0025] Define the original large model parameters as , after completing the parameter hierarchical Lora adjustment for different tasks respectively, the lora parameter of the -th layer of the large model corresponding to the task is , satisfying:
[0026] ;
[0027] Among them, and are two low-rank matrices;
[0028] Merge the original large model parameters with the LoRA fine-tuning parameters of multiple task types to generate a multi-task joint large model that can handle multiple tasks simultaneously. After parameter merging, the parameters of the large model are , and the calculation method is:
[0029] ;
[0030] That is, the LoRA parameters of different tasks are merged into the parameters of the base model according to a preset ratio; among them, is the total number of tasks; is the proportion of the LoRA parameters corresponding to each task in the adjustment of the entire large model parameters. The larger it is, the more important the corresponding task is; are the original large model parameters.
[0031] Furthermore, the model performance evaluation of the multi-task joint large model according to the preset task indicators includes:
[0032] For classification tasks, use accuracy, precision, recall, and F1 score to evaluate the performance of the large model; for question-and-answer tasks, use the C-Eval evaluation suite to evaluate the performance of the model; for machine translation tasks, use the BLEU score to evaluate the performance of the model.
[0033] In a second aspect, the present invention provides a multi-task joint large model hierarchical parameter adjustment system, including:
[0034] A hierarchical parameter adjustment module for iteratively adjusting the hierarchical parameters of the large model respectively by using LoRA fine-tuning according to the task types executed by the large model and the parameter adjustment data set to obtain the LoRA parameters corresponding to each task type; the parameter adjustment data set includes a question data set, a question-and-answer data set, and a translation data set; the task types include classification tasks, question-and-answer tasks, and machine translation tasks;
[0035] A parameter merging module for merging the original large model parameters with the LoRA fine-tuning parameters of multiple task types to generate a multi-task joint large model;
[0036] A model performance evaluation module for evaluating the model performance of the multi-task joint large model according to the preset task indicators and deploying the multi-task joint large model whose model performance evaluation results meet the multi-task requirements to the actual application environment.
[0037] Furthermore, the system further includes a data set preparation module, and the data set preparation module specifically includes:
[0038] A data collection unit for collecting various model parameter adjustment materials and storing the various model parameter adjustment materials in an object database;
[0039] A data preprocessing unit performs data preprocessing operations on various model parameter adjustment data in an object database to obtain preprocessed data;
[0040] A data segmentation unit is used to segment the preprocessed data to obtain multiple segments of preprocessed data;
[0041] A problem dataset combination unit is used to input each segment of the preprocessed data into a preset large model respectively to generate multiple first problem texts related to model parameter adjustment, and set the labels of the first problem texts as professional problems; extract some problems from the first open-source dataset as second problem texts, and set the labels of the second problem texts as non-professional problems; combine the first problem texts, the second problem texts and the corresponding labels of the texts to obtain a problem dataset;
[0042] A question-and-answer dataset combination unit inputs each segment of the preprocessed data into a preset large model respectively to generate multiple first question-and-answer pairs related to model parameter adjustment as professional question-and-answer pairs; extracts some question-and-answer pairs from the first open-source dataset as non-professional question-and-answer pairs, and combines the professional question-and-answer pairs with the non-professional question-and-answer pairs to obtain a question-and-answer dataset;
[0043] A translation dataset combination unit is used to input each segment of the preprocessed data into a preset large model respectively, extract multiple sentences and translate the multiple sentences into sentences in a preset language to obtain professional translation data; extracts some translation data from the second open-source dataset as non-professional translation data, and combines the professional translation data with the non-professional translation data to obtain a translation dataset.
[0044] In summary, the beneficial effects of the present invention are as follows:
[0045] A multi-task joint large model hierarchical parameter adjustment method provided by the present invention collects relevant data through multiple channels; preprocesses the collected data and converts it into a unified text format; generates multi-task data such as problems and labels, question-and-answer pairs, and translations for the collected data to support subsequent model parameter adjustment; for different tasks such as classification, question-and-answer, and machine translation, performs hierarchical parameter adjustment respectively, freezes the parameters of other layers in turn, and adjusts specific layers, thereby reducing the resource quantity required for parameter adjustment and improving the parameter adjustment efficiency; after completing the large model parameter adjustment for different tasks respectively, obtains different LoRA parameters and combines them, so that the large model can process multiple tasks simultaneously; performs performance evaluation on the large model obtained after parameter combination to ensure that it meets the requirements of multiple tasks. Description of the Drawings
[0046] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for use in the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings, and all are within the protection scope of the present invention.
[0047] Figure 1 It is a flowchart of a multi-task joint large model hierarchical parameter adjustment method of the present invention;
[0048] Figure 2 It is a flowchart of hierarchical parameter adjustment of the present invention;
[0049] Figure 3 It is an architecture diagram of the functional modules of a multi-task joint large model hierarchical parameter adjustment system of the present invention. Specific Embodiments
[0050] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. If there is no conflict, the various features in the present invention and its embodiments can be combined with each other, and all are within the protection scope of the present invention.
[0051] Embodiment 1:
[0052] Please refer to Figure 1 , Figure 1 It is a flowchart of a multi-task joint large model hierarchical parameter adjustment method in Embodiment 1 of the present invention. Referring to Figure 1 shown, the method mainly includes:
[0053] According to the task types executed by the large model and the parameter adjustment data set, the hierarchical parameters of the large model are iteratively adjusted respectively in the way of Lora fine-tuning to obtain the Lora parameters corresponding to each task type; the parameter adjustment data set includes a question data set, a question and answer data set, and a translation data set; the task types include classification tasks, question and answer tasks, and machine translation tasks;
[0054] Merge the original large model parameters with the Lora fine-tuning parameters of multiple task types to generate a multi-task joint large model;
[0055] Evaluate the model performance of the multi-task joint large model according to the preset task indicators, and deploy the multi-task joint large model whose model performance evaluation results meet the multi-task requirements to the actual application environment.
[0056] Further, as Figure 1 shown, the method of the embodiment of the present invention further includes the process of obtaining a parameter adjustment data set (i.e., the data preparation process), and this process specifically includes:
[0057] Collect a variety of model parameter adjustment materials and store them in the object database. Specifically, professional books, reports, web pages, etc. can be collected through libraries, the Internet, etc. and stored in the object database to support subsequent data set generation.
[0058] Perform data preprocessing operations on the various model parameter adjustment materials in the object database to obtain preprocessed materials. Among them, the main operations of data preprocessing include text verification and correction, cleaning unnecessary characters, deleting duplicate content, standardizing and formatting time information, etc. By writing automated scripts and using existing text processing tools and libraries (such as NLTK, spaCy, etc.) to assist in completing the preprocessing work.
[0059] The embodiment of the present invention performs preprocessing operations such as cleaning and standardizing the obtained materials to ensure data quality and improve the accuracy of subsequent model training.
[0060] Segment the preprocessed materials to obtain multiple segments of preprocessed materials. Specifically, the preprocessed materials can be segmented into segments of 200 words each to obtain multiple segments of preprocessed materials.
[0061] Input each segment of the preprocessed materials into a preset large model respectively to generate multiple first question texts related to model parameter adjustment, and set the labels of the first question texts as professional questions; then extract some questions from the first open-source data set as second question texts, and set the labels of the second question texts as non-professional questions; finally, combine the first question texts, the second question texts and the corresponding labels of the texts to obtain a question data set.
[0062] Input each segment of the preprocessed materials into a preset large model respectively to generate multiple first question-and-answer pairs related to model parameter adjustment as professional question-and-answer pairs; extract some question-and-answer pairs from the first open-source data set as non-professional question-and-answer pairs, and combine the professional question-and-answer pairs with the non-professional question-and-answer pairs to obtain a question-and-answer data set.
[0063] Among them, the first open-source data set is the OMGEval open-source data set. The preset large model can adopt large models such as qwen and llama.
[0064] Input each piece of preprocessed data into a preset large model respectively, extract multiple sentences and translate the multiple sentences into sentences in a preset language to obtain professional translation data; extract some translation data from the second open-source dataset as non-professional translation data, and combine the professional translation data with the non-professional translation data to obtain a translation dataset. Among them, the preset language can be set by oneself, for example, other languages except Chinese.
[0065] Further, in the embodiments of the present invention, the dataset is adjusted according to the task type and parameters executed by the large model, and the hierarchical parameters of the large model are iteratively adjusted respectively in the way of LoRA fine-tuning to obtain LoRA parameters corresponding to each task type, including:
[0066] Judge the task type executed by the large model. If the task type is a classification task, perform LoRA fine-tuning on the hierarchical parameters of the large model based on the problem data to obtain classification LoRA parameters; among them, the large model includes an embedding layer, a position encoding layer, an attention layer and a feed-forward layer;
[0067] If the task type is a question-and-answer task, perform LoRA fine-tuning on the hierarchical parameters of the large model based on the question-and-answer dataset to obtain question-and-answer LoRA parameters;
[0068] If the task type is a machine translation task, perform LoRA fine-tuning on the hierarchical parameters of the large model based on the translation dataset to obtain translation LoRA parameters.
[0069] Specifically, the LoRA fine-tuning process of the hierarchical parameters includes:
[0070] First, freeze other layers of the large model and perform LoRA fine-tuning on the parameters of the embedding layer. Then, freeze other layers of the large model except the position encoding layer and perform LoRA fine-tuning on the parameters of the position encoding layer to assign a position encoding to the position of each element in the input sequence. Freeze other layers of the large model except the attention layer and perform LoRA fine-tuning on the parameters of the attention layer to calculate the similarity between each element in the input sequence, assign a weight to each element, and capture the dependence and importance between different elements according to the weight of the element. Then, freeze other layers of the large model except the feed-forward layer and perform LoRA fine-tuning on the parameters of the feed-forward layer. Map the input sequence to a higher-dimensional space than the original dimension through a linear transformation, apply the ReLU activation function, and map the data back to the original dimension through a linear transformation. Finally, freeze other layers of the large model except the output layer and perform LoRA fine-tuning on the parameters of the output layer to output the final result generated by the model.
[0071] In the embodiments of the present invention, for different tasks such as classification, question answering, and machine translation, hierarchical parameter adjustment is performed respectively, that is, the parameters of other layers are frozen in turn, and specific layers are adjusted, so as to reduce the amount of resources required for parameter adjustment and improve the efficiency of parameter adjustment. When adjusting the parameters of specific layers, the LoRA fine-tuning method is adopted to further reduce the consumption of computing resources and speed up the completion of parameter adjustment. For the classification task, the question dataset is used for parameter adjustment; for the question answering task, the question answering dataset is used for parameter adjustment; for the machine translation task, the translation dataset is used for parameter adjustment.
[0072] Refer to Figure 2 As shown, the main process of the hierarchical parameter adjustment method in the embodiments of the present invention is as follows: First, freeze other layers and adjust the parameters of the embedding layer to improve the representation ability of the model. Then, freeze other layers and adjust the parameters of the position encoding layer to improve the sensitivity of the model to the element position information. Next, freeze other layers and adjust the parameters of the attention layer to capture the dependency relationships and importance between different elements. Then, freeze other layers and adjust the parameters of the feed-forward layer to process the mutual relationships between words, and enable the model to learn more complex patterns and relationships by introducing non-linear activation functions. Finally, freeze other layers and adjust the parameters of the output layer to output the final result generated by the model. Differentiated learning strategies are implemented for the parameters of different layers. For example, for the bottom-layer shared parameters, a more conservative learning rate is adopted to maintain a stable feature extraction ability; while for the task-specific parameters, a more aggressive learning strategy is adopted to quickly adapt to the requirements of their respective tasks.
[0073] Furthermore, the embedding layer converts each element (such as a word) in the input sequence into a vector representation of a fixed dimension, and these vectors are subsequently used for subsequent calculations and processing. Vector embedding maps each discrete element to a real-valued vector with a fixed dimension. This vector represents the position of the element in an abstract space, enabling the model to process and understand the text data that could not be directly subjected to mathematical operations, thereby capturing the semantic information of the vocabulary.
[0074] Furthermore, in many natural language processing tasks, the order between words is crucial. The position encoding layer assigns a unique vector to each position in the sequence, helping the model distinguish the information of different positions, making up for the shortcoming that the attention layer cannot capture the order information of the elements in the sequence, and thus better understanding the structure and content of the text.
[0075] Furthermore, the attention layer captures the dependency relationships between different elements. By calculating the similarity between each element in the input sequence, a weight is assigned to each element, thereby determining which parts are more important for the current task and capturing information at different levels.
[0076] Furthermore, the feed-forward layer consists of two linear transformations, with a ReLU activation function between the two transformations. First, the input is mapped to a higher-dimensional space (usually referred to as "expansion") through a linear transformation, then the ReLU activation function is applied, and finally, the data is mapped back to the original dimension through another linear transformation. The purpose of this "expansion" and "contraction" design is to handle the relationships between words, and by introducing a non-linear activation function, the model can capture more complex features and improve the model's expressive power.
[0077] Furthermore, the output layer is usually composed of a linear layer, which is the last layer of the model structure. The output layer outputs the final result generated by the model, such as the probability of each word, the classification result, etc. By finely tuning the parameters of this layer, it is ensured that the model can generate accurate and relevant outputs.
[0078] Furthermore, the specific process of parameter LoRA fine-tuning is as follows: Fix the weights of the original large model, define two low-rank matrices A and B as additional parameters to participate in the operation, and sum the results of the two paths as the output of this layer in the large model; when updating the hierarchical parameters, calculate the loss function based on the difference between the model's answer and the correct answer, and update the low-rank matrices A and B through the backpropagation algorithm to minimize the loss function; iteratively update the hierarchical parameters until the performance of the large model reaches the predetermined standard or the training reaches the preset number of iterations.
[0079] Furthermore, the process of merging the original large model parameters with the LoRA fine-tuning parameters of multiple task types to generate a multi-task joint large model includes:
[0080] Define the original large model parameters as , after completing the parameter hierarchical LoRA adjustment for different tasks respectively, the LoRA parameters of the -th layer of the large model corresponding to task are , satisfying:
[0081] ;
[0082] where and are two low-rank matrices;
[0083] Merge the original large model parameters with the LoRA fine-tuning parameters of multiple task types to generate a multi-task joint large model that can handle multiple tasks simultaneously. The parameters of the large model after parameter merging are , and the calculation method is:
[0084] ;
[0085] Merge the LoRA parameters of different tasks into the parameters of the base model according to a preset ratio; where is the total number of tasks; is the proportion of the LoRA parameters corresponding to each task in the adjustment of the parameters of the entire large model. The larger it is, the more important the corresponding task is; is the parameter of the original large model.
[0086] Furthermore, in the embodiments of the present invention, the model performance of the multi-task joint large model is evaluated according to preset task indicators, including:
[0087] For classification tasks, accuracy, precision, recall, and F1 score are used to evaluate the performance of the large model; for question-answering tasks, the C-Eval evaluation suite is used to evaluate the performance of the model; for machine translation tasks, the BLEU score is used to evaluate the performance of the model.
[0088] In the embodiments of the present invention, the performance of the large model obtained after parameter merging is evaluated to ensure that it meets the requirements of multi-tasks. The model that meets the requirements is deployed to the actual application environment for users to call for downstream tasks.
[0089] For classification tasks, indicators such as accuracy, precision, recall, and F1 score are used to evaluate the model performance; for question-answering tasks, C-Eval is used to evaluate the performance of the large model; for machine translation tasks, the BLEU (Bilingual Evaluation Understudy) score is used for evaluation.
[0090] Furthermore, for the classification performance indicators, accuracy refers to the proportion of the number of samples correctly predicted by the model to the total number of samples; precision refers to the proportion of the samples predicted by the model as a certain type that are actually of that type; recall focuses on the proportion of all samples of a certain type that are accurately identified; the F1 score is the harmonic mean of recall and precision.
[0091] Furthermore, the C-Eval is a comprehensive Chinese model evaluation suite, containing 13,948 multiple-choice questions, covering 52 different disciplines and four difficulty levels, used to evaluate the Chinese understanding ability of the large model.
[0092] Furthermore, the BLEU score is an indicator used to evaluate the quality of machine translation, measuring the closeness of the machine translation to a set of high-quality human translations, with its value between 0 and 1, where 1 means that the result of the machine translation is exactly the same as the result of the human translation. The BLEU score is calculated by comparing the output of the machine translation with one or more reference translations, mainly measuring the accuracy, relevance, and fluency of the translation.
[0093] In the embodiments of the present invention, the multi-task joint large model hierarchical parameter adjustment method has the following characteristics:
[0094] 1. The parameter merging technology is adopted to realize multi-task joint parameter adjustment, enabling the large model to simultaneously learn the feature representations of multiple related tasks, promoting the large model's understanding of a wider range of knowledge, and thus improving the large model's ability to execute multiple tasks.
[0095] 2. The hierarchical parameter adjustment method is adopted to reduce the resource requirements for large model parameter adjustment, allowing for the implementation of differentiated learning strategies for parameters at different levels, improving the efficiency and effect of parameter adjustment, enhancing the overall performance of the model, and accelerating the model iteration cycle.
[0096] Embodiment 2: Refer to Figure 3 As shown, the embodiments of the present invention provide a multi-task joint large model hierarchical parameter adjustment system, which includes:
[0097] The hierarchical parameter adjustment module is used to iteratively adjust the hierarchical parameters of the large model by means of Lora fine-tuning according to the task types executed by the large model and the parameter adjustment data set, and obtain the Lora parameters corresponding to each task type; the parameter adjustment data set includes a question data set, a question and answer data set, and a translation data set; the task types include classification tasks, question and answer tasks, and machine translation tasks;
[0098] The parameter merging module is used to merge the original large model parameters with the Lora fine-tuning parameters of multiple task types to generate a multi-task joint large model;
[0099] The model performance evaluation module is used to evaluate the model performance of the multi-task joint large model according to the preset task metrics, and deploy the multi-task joint large model whose model performance evaluation results meet the multi-task requirements to the actual application environment.
[0100] Furthermore, as shown in Figure 3 The system of the embodiments of the present invention further includes a data set preparation module, and the data set preparation module specifically includes:
[0101] The data collection unit is used to collect various model parameter adjustment materials and store the various model parameter adjustment materials in the object database;
[0102] The data preprocessing unit performs data preprocessing operations on the various model parameter adjustment materials in the object database to obtain preprocessed materials;
[0103] The data segmentation unit is used to segment the preprocessed materials to obtain multiple segments of preprocessed materials;
[0104] The question dataset combination unit is used to input each piece of preprocessed data into a preset large model respectively, generate multiple first question texts related to model parameter adjustment, and set the labels of the first question texts as professional questions; extract some questions from the first open-source dataset as second question texts, and set the labels of the second question texts as non-professional questions; combine the first question texts, the second question texts and the corresponding labels of the texts to obtain a question dataset;
[0105] The question-and-answer dataset combination unit inputs each piece of preprocessed data into a preset large model respectively, generates multiple first question-and-answer pairs related to model parameter adjustment as professional question-and-answer pairs; extracts some question-and-answer pairs from the first open-source dataset as non-professional question-and-answer pairs, and combines the professional question-and-answer pairs with the non-professional question-and-answer pairs to obtain a question-and-answer dataset;
[0106] The translation dataset combination unit is used to input each piece of preprocessed data into a preset large model respectively, extract multiple sentences and translate the multiple sentences into sentences in a preset language to obtain professional translation data; extract some translation data from the second open-source dataset as non-professional translation data, and combine the professional translation data with the non-professional translation data to obtain a translation dataset.
[0107] The system according to the embodiments of the present invention collects relevant data through multiple channels; preprocesses the collected data and converts it into a unified text format; generates multi-task data such as questions and labels, question-and-answer pairs, translations, etc. for the collected data to support subsequent model parameter adjustment; for different tasks such as classification, question-and-answer, and machine translation, performs hierarchical parameter adjustment respectively, freezes the parameters of other layers in turn, and adjusts specific layers, so as to reduce the resources required for parameter adjustment and improve the parameter adjustment efficiency; after completing the parameter adjustment of the large model for different tasks respectively, obtains different LoRA parameters and combines them, so that the large model can process multiple tasks simultaneously; performs performance evaluation on the large model obtained after parameter combination to ensure that it meets the requirements of multiple tasks.
[0108] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A hierarchical parameter adjustment method for a multi-task joint large model, characterized in that, Including: Adjust the dataset according to the task type and parameters of the large model, and use the method of LoRA fine-tuning to iteratively adjust the hierarchical parameters of the large model respectively to obtain the LoRA parameters corresponding to each task type, including: Judge the task type executed by the large model. If the task type is a classification task, perform LoRA fine-tuning on the hierarchical parameters of the large model based on the question dataset to obtain classification LoRA parameters; among them, the large model includes an embedding layer, a position encoding layer, an attention layer, and a feed-forward layer; if the task type is a question-and-answer task, perform LoRA fine-tuning on the hierarchical parameters of the large model based on the question-and-answer dataset to obtain question-and-answer LoRA parameters; if the task type is a machine translation task, perform LoRA fine-tuning on the hierarchical parameters of the large model based on the translation dataset to obtain translation LoRA parameters; the parameter adjustment dataset includes a question dataset, a question-and-answer dataset, and a translation dataset; the task types include classification tasks, question-and-answer tasks, and machine translation tasks; The process of the hierarchical parameter LoRA fine-tuning includes: First, freeze other layers except the embedding layer of the large model, and perform parameter LoRA fine-tuning on the embedding layer. Then, freeze other layers except the position encoding layer of the large model, and perform parameter LoRA fine-tuning on the position encoding layer to assign a position encoding to the position of each element in the input sequence. Freeze other layers except the attention layer of the large model, and perform parameter LoRA fine-tuning on the attention layer to calculate the similarity between each element in the input sequence, assign a weight to each element, and capture the dependencies and importance between different elements according to the weights of the elements. Then, freeze other layers except the feed-forward layer of the large model, and perform parameter LoRA fine-tuning on the feed-forward layer. Map the input sequence to a higher-dimensional space than the original dimension through a linear transformation, apply the ReLU activation function, and map the data back to the original dimension through a linear transformation. Finally, freeze other layers except the output layer of the large model, and perform parameter LoRA fine-tuning on the output layer to output the final result generated by the model; Merge the original large model parameters with the LoRA fine-tuning parameters of multiple task types to generate a multi-task joint large model; Evaluate the model performance of the multi-task joint large model according to the preset task metrics, and deploy the multi-task joint large model whose model performance evaluation results meet the multi-task requirements to the actual application environment.
2. The multi-task joint large model hierarchical parameter adjustment method according to claim 1, wherein It also includes the acquisition process of the parameter adjustment dataset: Collect various model parameter adjustment materials and store them in the object database; Perform data preprocessing operations on various model parameter adjustment materials in the object database to obtain preprocessed materials; Segment the preprocessed materials to obtain multiple segments of preprocessed materials; Input each segment of the preprocessed materials into the preset large model respectively to generate multiple first question texts related to model parameter adjustment, and set the labels of the first question texts as professional questions; extract some questions from the first open-source dataset as the second question texts, and set the labels of the second question texts as non-professional questions; Combine the first question texts, the second question texts, and the corresponding labels of the texts to obtain a question dataset; Input each piece of preprocessed data into a preset large model respectively to generate multiple first question-answer pairs related to model parameter adjustment as professional question-answer pairs; extract some question-answer pairs from the first open-source dataset as non-professional question-answer pairs, and combine the professional question-answer pairs with the non-professional question-answer pairs to obtain a question-answer dataset; Input each piece of preprocessed data into a preset large model respectively, extract multiple sentences and translate the multiple sentences into sentences in a preset language to obtain professional translation data; Extract some translation data from the second open-source dataset as non-professional translation data, and combine the professional translation data with the non-professional translation data to obtain a translation dataset.
3. The multi-task joint large model hierarchical parameter adjustment method according to claim 1, wherein The specific process of fine-tuning the parameter Lora is as follows: fix the weight parameters of the original large model, define two low-rank matrices A and B as additional parameters to participate in the operation, and sum the results of the two links as the output of this layer in the large model; when updating the hierarchical parameters, calculate the loss function according to the difference between the answer of the model and the correct answer, and update the low-rank matrices A and B through the backpropagation algorithm to minimize the loss function; iteratively update the hierarchical parameters until the performance of the large model reaches a predetermined standard or the training reaches a preset number of iteration rounds.
4. The method for hierarchical parameter adjustment of the multi-task joint large model according to claim 1, wherein The combination of the original large model parameters and the Lora fine-tuning parameters of multiple task types to generate a multi-task joint large model includes: Define the original large model parameters as , after separately completing the parameter hierarchical LoRA adjustment for different tasks, the LoRA parameters of the -th layer of the large model corresponding to the task are , satisfying: ; Among them, and are two low-rank matrices; Merge the original large model parameters with the LoRA fine-tuning parameters of multiple task types to generate a multi-task joint large model that can handle multiple tasks simultaneously. The parameters of the large model after parameter merging are , and the calculation method is: ; Merge the LoRA parameters of different tasks into the parameters of the base model according to a preset ratio; among them, is the total number of tasks; is the proportion of the LoRA parameters corresponding to each task in the adjustment of the parameters of the entire large model, The larger it is, the more important the corresponding task is; are the parameters of the original large model.
5. The multi-task joint large model hierarchical parameter adjustment method according to claim 1, characterized in that The evaluation of the model performance of the multi-task joint large model according to preset task metrics includes: For the classification task, use accuracy, precision, recall rate, and F1 score to evaluate the performance of the large model; for the question-answer task, use the C-Eval evaluation suite to evaluate the performance of the model; for the machine translation task, use the BLEU score to evaluate the performance of the model.
6. A hierarchical parameter adjustment system for a multi-task joint large model, characterized in that, Includes: A hierarchical parameter adjustment module, which is used to iteratively adjust the hierarchical parameters of the large model respectively in the form of Lora fine-tuning according to the task type executed by the large model and the parameter adjustment dataset to obtain the Lora parameters corresponding to each task type, including: judging the task type executed by the large model, if the task type is a classification task, then perform Lora fine-tuning on the hierarchical parameters of the large model based on the question dataset to obtain classification Lora parameters; among them, the large model includes an embedding layer, a position encoding layer, an attention layer, and a feed-forward layer; if the task type is a question-answer task, then perform Lora fine-tuning on the hierarchical parameters of the large model based on the question-answer dataset to obtain question-answer Lora parameters; if the task type is a machine translation task, then perform Lora fine-tuning on the hierarchical parameters of the large model based on the translation dataset to obtain translation Lora parameters; the parameter adjustment dataset includes a question dataset, a question-answer dataset, and a translation dataset; the task types include classification tasks, question-answer tasks, and machine translation tasks; The process of hierarchical parameter LoRA fine-tuning includes: First, freeze the layers other than the embedding layer of the large model, and perform parameter LoRA fine-tuning on the embedding layer. Then, freeze the layers other than the position encoding layer in the large model, and perform parameter LoRA fine-tuning on the position encoding layer to assign a position encoding to the position of each element in the input sequence. Freeze the layers other than the attention layer in the large model, and perform parameter LoRA fine-tuning on the attention layer to calculate the similarity between each element in the input sequence, assign a weight to each element, and capture the dependencies and importance between different elements according to the weights of the elements. Then, freeze the layers other than the feed-forward layer in the large model, and perform parameter LoRA fine-tuning on the feed-forward layer to map the input sequence to a higher-dimensional space than the original dimension through a linear transformation, apply the ReLU activation function, and map the data back to the original dimension through a linear transformation. Finally, freeze the layers other than the output layer in the large model, and perform parameter LoRA fine-tuning on the output layer to output the final result generated by the model. A parameter merging module for merging the original large model parameters with the LoRA fine-tuning parameters of multiple task types to generate a multi-task joint large model. A model performance evaluation module for evaluating the model performance of the multi-task joint large model according to preset task metrics, and deploying the multi-task joint large model whose model performance evaluation results meet the multi-task requirements to the actual application environment.
7. The multi-task joint large model hierarchical parameter adjustment system according to claim 6, wherein It also includes a dataset preparation module, and the dataset preparation module specifically includes: A data collection unit for collecting various model parameter adjustment materials and storing the various model parameter adjustment materials in an object database. A data preprocessing unit for performing data preprocessing operations on the various model parameter adjustment materials in the object database to obtain preprocessed materials. A data segmentation unit for segmenting the preprocessed materials to obtain multiple segments of preprocessed materials. A problem dataset combination unit for inputting each segment of the preprocessed materials into a preset large model respectively to generate multiple first problem texts related to model parameter adjustment, setting the labels of the first problem texts as professional questions; extracting some problems from the first open-source dataset as second problem texts, and setting the labels of the second problem texts as non-professional questions; combining the first problem texts, the second problem texts and the corresponding labels of the texts to obtain a problem dataset. A Q&A dataset combination unit for inputting each segment of the preprocessed materials into a preset large model respectively to generate multiple first Q&A pairs related to model parameter adjustment as professional Q&A pairs; extracting some Q&A pairs from the first open-source dataset as non-professional Q&A pairs, and combining the professional Q&A pairs with the non-professional Q&A pairs to obtain a Q&A dataset. A translation dataset combination unit for inputting each segment of the preprocessed materials into a preset large model respectively, extracting multiple sentences and translating the multiple sentences into sentences in a preset language to obtain professional translation data; extracting some translation data from the second open-source dataset as non-professional translation data, and combining the professional translation data with the non-professional translation data to obtain a translation dataset.
Citation Information
Patent Citations
Multi-task large model fine tuning method based on adapters and low-rank adaptation
CN116822611A