Data expansion and training method and hardware for generating large model
The expert data set is expanded through the associative data generation and optimization processing model, combined with general question and answer, improved training of the domain model for the data, and updated parameters using accompanying model and domain training model, solving the problem of general capability degradation in vertical domain large model training, and realizing the evolution rather than degradation of the model.
Patent Information
- Application Number
- CN202411971497.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-06
AI Technical Summary
In the training process of vertical domain big models, how to fine-tune the training data of the domain big models by optimizing the training data of the domain big models without damaging general capabilities to achieve the evolution rather than degeneration of the model.
The expert data set is expanded by using the associative data generation and optimization processing model to generate the optimized training data set, and the data is improved for training the domain model in combination with the screened general question and answer. The accompanying model and domain training model are used to update parameters based on KL divergence loss and cross entropy loss.
By designing high-quality field training samples and reasonable data mixing ratios, we ensure the quality and diversity of the training data set, shorten the model training time, avoid model degradation, and improve the professional performance of the field model and the retention of general capabilities.
Smart Images

Figure CN119940408A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of calculation, inference or counting, and in particular to a method and hardware for generating large model data expansion and training. Background Art
[0002] Artificial intelligence big models refer to large-parameter models trained using large-scale data and powerful computing power. These models are usually highly versatile and generalizable, and can be applied to natural language processing, image recognition, speech recognition and other fields.
[0003] A vertical domain big model refers to a large language model that is trained and optimized based on the domain knowledge of a specific domain or industry, with a general big model as the basic model. The vertical domain big model is more focused on the knowledge and skills in a specific field, and has higher domain professionalism and practicality. However, due to the requirements for accuracy and the frequent iteration of the knowledge base, the main problem currently faced in the training and application of vertical domain big models is how to balance the relationship between general capabilities and domain-specific capabilities. Although the big model performs well in general tasks, its application in vertical sub-domains is often unsatisfactory due to the lack of targeted domain data support. In order to improve the performance of vertical fields, it is usually necessary to collect domain-specific data for training, but such data sets are usually difficult to obtain, and during the training process, too much domain data may cause the general capabilities of the model to degrade.
[0004] The current training strategy is generally based on a general model, collecting a certain amount of domain sample data, mixing it with a general data set, and training it in a certain proportion. This method often uses a cross entropy loss function to update parameters, but this method often has a significant impact on the general ability of the model and leads to degradation.
[0005] Therefore, how to fine-tune the training data of large domain models by optimizing them without compromising general capabilities to achieve model evolution rather than degradation has become a key issue that the industry needs to solve urgently. Summary of the invention
[0006] In order to solve the above technical problems, the present invention provides a method and hardware for generating large model data expansion and training.
[0007] The technical solution adopted by the present invention is a method for generating large model data expansion and training. The method expands the expert data set based on the associative data generation and optimization processing model to obtain an optimized training data set, and uses the optimized training data set and the screened general question and answer data set to improve the training of the domain model to obtain an optimized domain model.
[0008] Preferably, the associative data generation and optimization processing model includes a composite LLM-MLLM-LLM model;
[0009] The domain expert labeled data in the expert dataset is input into the initial LLM to obtain the initial domain associative question and answer pair data, the initial domain associative question and answer pair data is input into the MLLM to obtain the corresponding model intrinsic capability question and answer pair data, the initial domain associative question and answer pair data and the model intrinsic capability question and answer pair data are jointly input into the terminal LLM to obtain the optimized training dataset.
[0010] Preferably, the relevant fields of each piece of data in the expert data set are used as background information to form prompt words for the large model.
[0011] Preferably, the terminal LLM selects the optimal answers to the initial domain-associative question-answer pair data and the model-intrinsic capability question-answer pair data formed for the same question as the training sample pair data in the optimized training data set.
[0012] Preferably, the optimized training data set and the screened general question-answer pair data set are distinguished by data type markers.
[0013] Preferably, the improved training includes an adjoint model and a domain training model, and the parameters of the adjoint model remain unchanged during the training process; the overall loss of the domain training model is obtained based on the KL divergence loss and the cross entropy loss, which is used to update the parameter gradient of the training model.
[0014] Preferably, the general question and answer pair data in the screened general question and answer pair data set are simultaneously input into the adjoint model and the domain training model for forward calculation, and the difference between the probability distributions of the adjoint model and the domain training model on the general knowledge data is calculated based on the KL divergence to obtain the KL divergence loss.
[0015] Preferably, the training data in the optimized training data set is input into the domain training model, and the cross entropy loss is calculated based on the domain knowledge data labels.
[0016] A computer-readable storage medium stores a program for generating large model data expansion and training, which implements the method for generating large model data expansion and training when executed by a processor.
[0017] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the method for generating large model data expansion and training is implemented.
[0018] The present invention provides a method and hardware for generating large model data expansion and training. The method expands the expert data set based on the associative data generation and optimization processing model to obtain an optimized training data set, and uses the optimized training data set and the screened general question and answer data set to improve the training of the domain model to obtain the optimized domain model; the execution of the hardware is realized based on the method.
[0019] The beneficial effects of the present invention are:
[0020] (1) We designed and constructed high-quality domain training samples, adopted effective general data screening strategies and reasonable data mixing ratios, and ensured that the training data set can reduce the amount of sample data as much as possible while ensuring quality, diversity, and generation efficiency, thereby shortening the model training time;
[0021] (2) Improve the training method and use the companion model to ensure that the data distribution during the training process is controllable, so that vertical domain knowledge can be effectively learned without causing significant negative impact on general knowledge. Ultimately, the goal of avoiding degradation of large domain models during the evolution process can be achieved, and the accuracy of domain model training can be improved.
[0022] Through such a strategy, we can not only improve the specialized performance of large domain models, but also retain their general capabilities to the maximum extent, ensuring their wide applicability and high efficiency in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 is a flow chart of the present invention;
[0024] Figure 2 A schematic diagram of obtaining an optimized training data set for the present invention;
[0025] Figure 3 A schematic diagram of the improved training of the present invention;
[0026] Figure 4 It is a schematic diagram of the implementation of the corresponding improved training in the present invention. DETAILED DESCRIPTION
[0027] The present invention is further described in detail below in conjunction with embodiments, but the protection scope of the present invention is not limited thereto.
[0028] The present invention relates to a method for generating large model data expansion and training. The method expands an expert data set based on an associative data generation and optimization processing model to obtain an optimized training data set, and uses the optimized training data set and a screened general question-answer pair data set to perform improved training on a domain model to obtain an optimized domain model.
[0029] In the present invention, firstly, an optimized training data set is constructed based on the associative data generation and optimization processing model, and the difference between the large model's own knowledge and the associative stories is used to enhance and select the data set, and the difference data is fully utilized for learning, and the optimized training data set is used for training to clarify the training direction of the model; this increases the amount of training data and the dimension of the present invention, but the model remains non-degraded.
[0030] The present invention mainly includes the construction of an optimized training data set and the execution of improved training using the same.
[0031] (1) Construction of optimized training dataset
[0032] The associative data generation and optimization processing model includes a composite LLM-MLLM-LLM model;
[0033] The domain expert labeled data in the expert dataset is input into the initial LLM to obtain the initial domain associative question and answer pair data, the initial domain associative question and answer pair data is input into the MLLM to obtain the corresponding model intrinsic capability question and answer pair data, the initial domain associative question and answer pair data and the model intrinsic capability question and answer pair data are jointly input into the terminal LLM to obtain the optimized training dataset.
[0034] The relevant fields of each data in the expert dataset are used as background information to form prompt words for the large model.
[0035] The terminal LLM screens the optimal answers to the initial domain-associative question-answer pair data and the model's intrinsic ability question-answer pair data formed for the same question as the training sample pair data in the optimized training data set.
[0036] The optimized training dataset and the filtered general question-answer dataset are distinguished by data type marks.
[0037] In the present invention, the expert dataset refers to a structured dataset annotated by domain experts, and the data is optimized using a composite LLM-MLLM-LLM model, where LLM is a large language model and MLLM is a multimodal large language model.
[0038] In the specific implementation, the data annotated by domain experts is used as the background information of the prompt word (prompt) and as the input of LLM to form the initial domain associative question-answer pair data. Then, with the help of MLLM, the questions and image data in the initial domain associative question-answer pair data generated in the previous step are directly input into MLLM to form the answer of the intrinsic knowledge of the current model, thereby forming the model intrinsic ability question-answer pair data corresponding to the same question. Then, for the same question, two different answers are combined and input into the terminal LLM, so that the terminal LLM can screen out better question-answer pair data as the domain optimized training sample question-answer pair data to build an optimized training data set.
[0039] For general-domain question-answer data sets, a certain data screening strategy is adopted. The screening criteria include but are not limited to stratified sampling of image resolution, question diversity, length of answer content, translation from English to Chinese, etc., to screen out a high-quality general question-answer data subset as the screened general question-answer data set.
[0040] The optimized training data set and the screened general question-answering data set are mixed and labeled using data type labeling to form the final domain model training sample data set.
[0041] The following is further explained in conjunction with the embodiments. It should be particularly noted that the embodiments of the current patent are based on the perspective of cultural relics and cultural heritage, so the embodiments are mainly based on the identification of stone statues. The image represented by the stone statue is only the sum of the characteristics of the stone statue. The actual application scenarios of the present invention are not limited to this. The identified object can be any other image and / or text with an expert data set.
[0042] The selected LLMs are all privately deployed Qwen2-72B-Instruct models. The initial LLM combines the expert-annotated data set and uses the relevant fields of each data as background information to form the prompt of the large model, so that the large model can generate domain-associated question-answer data;
[0043] The output results of the initial domain association question-answering data are as follows:
[0044] {
[0045] "question": "The image depicted in this statue has a prominent armor feature on its clothing and a cross pattern on its chest. This kind of attire and decoration is usually associated with which period of statue style?",
[0046] "answer": "This style of attire and decoration is usually associated with statues from a certain period.",
[0047] "explanation": "The statues of a certain period were influenced by northern culture and wars. The images often wore armor, symbolizing the military role they protected. In addition, the cross pattern on the chest is a decoration on the armor, which belongs to the style of a specific period and is representative."
[0048] }
[0049] For the generation of question-answer data of the intrinsic capability of the model, the present invention uses the privately deployed Qwen2-VL-72B-Instruct, and combines the question generated in the previous step with the real picture information, and inputs it into the MLLM to form the corresponding question-answer data of the intrinsic capability of the model;
[0050] The output results of the model intrinsic capability question-answer pair data corresponding to the MLLM output are as follows:
[0051] The prompt word at this time is
[0052] prompt = """
[0053] The figure depicted in this statue has a distinctive armor-like garment and a cross on the chest. This kind of attire and decoration is usually associated with which period of statue style?
[0054] """
[0055] The result is
[0056] answer = """
[0057] The image depicted in this statue has a distinctive armored costume and a cross on the chest. This kind of attire and decoration is usually associated with the style of statues of a certain period. The artistic style of a certain period is known for its fine carvings and intricate decorations, which often show the solemnity and majesty of the statues.
[0058] """
[0059] For domain-optimized question-answer pair data, the present invention uses the privately deployed Qwen2-72B-Instruct model and designs prompts to allow the terminal LLM to screen out the two different answers formed for the same question in the first two steps, and determine which answer is better, ultimately forming domain-optimized training sample pair data.
[0060] Finally, the mixed optimized training data set and the screened general question-answer data set are distinguished by data type markers for training of large domain models.
[0061] (2) Execute improved training
[0062] The improved training includes an adjoint model and a domain training model. During the training process, the parameters of the adjoint model remain unchanged. The overall loss of the domain training model is obtained based on the KL divergence loss and the cross entropy loss, which is used to update the parameter gradient of the training model.
[0063] In the present invention, LLaVA-OneVision-7B is selected as the base model, and the loss function of the model is redesigned in the training of the domain large model around the training goal of "evolution without degradation" of the domain large model; during the training process, the parameters of the accompanying model are frozen as a whole, and the parameters of the training model need to be fully updated; Figure 4 As shown in FIG. 1 , for sample data in the same batch during the training process, the data in the same batch are divided into two categories of data: general knowledge and domain knowledge based on data type labels, which correspond to the screened general question-answer pair dataset and the optimized training dataset, respectively.
[0064] The total loss is expressed as
[0065]
[0066] Among them, kl_loss is the KL divergence loss corresponding to the general knowledge training data, ce_loss is the cross entropy loss corresponding to the domain knowledge training data, and alpha is the adjustment parameter, which is 0.5 here.
[0067] The general question-answer pair data in the filtered general question-answer pair dataset are simultaneously input into the adjoint model and the domain training model for forward calculation to generate the output logits1 of the adjoint model and the output logits2 of the training model. The difference between the probability distributions of the adjoint model and the domain training model on the general knowledge data is calculated based on the KL divergence to obtain the KL divergence loss kl_loss.
[0068]
[0069] Among them, P and Q correspond to the target probability distribution and the predicted probability distribution respectively.
[0070]
[0071]
[0072] The KL divergence loss is
[0073]
[0074] P represents the probability distribution of the general knowledge training data formed by the adjoint model, and Q represents the probability distribution of the training model on the general knowledge data. The distribution Q predicted by the training model is used to approximate the target distribution P of the adjoint model, so that the model ability of the training model on the general knowledge training data will not be degraded.
[0075] The training data in the optimized training dataset is input into the domain training model, and the cross entropy loss is calculated based on the domain knowledge data labels.
[0076] The domain knowledge data in the optimized training data set is only input into the training model to form the model output logits3, and the cross-entropy loss ce_loss is calculated in combination with the domain knowledge data label labels;
[0077] For multi-classification problems, assuming there are C categories, the true label of the sample It is represented by one-hot encoding. Indicates the true label The value on the class (0 or 1), the probability distribution predicted by the model is ,in Indicates that the sample belongs to The predicted probability of the class;
[0078] The mathematical definition of cross entropy loss is .
[0079] Assume that the output logits3 vector of the model is , and the true label labels vector is , then based on the above formula, the cross entropy loss of this paper is as follows:
[0080]
[0081] in:
[0082] is the number of categories (e.g., vocabulary size);
[0083] Is the first The indicator variable of the class, if the sample belongs to this class, then ,otherwise, ;
[0084] is the model prediction probability after softmax conversion, indicating that the model predicts that the sample belongs to The probability of the class.
[0085] Through the above model training method, the distribution of general data of the training model will not experience violent fluctuations, thereby ensuring that the general ability of the training model will not be significantly degraded, thereby ensuring the retention of the general knowledge ability of the model as much as possible; at the same time, for domain data, the training model will also perform targeted learning optimization, so that the domain model completes domain knowledge learning during the training process while the general ability does not degrade (the domain model evolves but does not degenerate).
[0086] The present invention also relates to a computer-readable storage medium, on which a program for generating large model data expansion and training is stored. When the program is executed by a processor, the method for generating large model data expansion and training is implemented.
[0087] The present invention also relates to a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the method for generating large model data expansion and training is implemented.
[0088] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0089] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0090] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0091] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0092] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0093] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A method for generating large model data expansion and training, characterized in that: The method expands the expert data set based on the associative data generation and optimization processing model to obtain an optimized training data set, and uses the optimized training data set and the screened general question-answer pair data set to improve the training of the domain model to obtain the optimized domain model.
2. The method for generating large model data expansion and training according to claim 1, characterized in that: The associative data generation and optimization processing model includes a composite LLM-MLLM-LLM model; The domain expert labeled data in the expert dataset is input into the initial LLM to obtain the initial domain associative question and answer pair data, the initial domain associative question and answer pair data is input into the MLLM to obtain the corresponding model intrinsic capability question and answer pair data, the initial domain associative question and answer pair data and the model intrinsic capability question and answer pair data are jointly input into the terminal LLM to obtain the optimized training dataset.
3. The method for generating large model data expansion and training according to claim 2, characterized in that: The relevant fields of each data in the expert dataset are used as background information to form prompt words for the large model.
4. The method for generating large model data expansion and training according to claim 2, characterized in that: The terminal LLM screens the optimal answers to the initial domain-associative question-answer pair data and the model's intrinsic ability question-answer pair data formed for the same question as the training sample pair data in the optimized training data set.
5. The method for generating large model data expansion and training according to claim 1, characterized in that: The optimized training dataset and the filtered general question-answer dataset are distinguished by data type marks.
6. The method for generating large model data expansion and training according to claim 1, characterized in that: The improved training includes an adjoint model and a domain training model. During the training process, the parameters of the adjoint model remain unchanged. The overall loss of the domain training model is obtained based on the KL divergence loss and the cross entropy loss, which is used to update the parameter gradient of the training model.
7. A method for generating large model data expansion and training according to claim 6, characterized in that: The general question and answer pair data in the screened general question and answer pair dataset are simultaneously input into the adjoint model and the domain training model for forward calculation. The difference between the probability distributions of the adjoint model and the domain training model on the general knowledge data is calculated based on the KL divergence to obtain the KL divergence loss.
8. The method for generating large model data expansion and training according to claim 6, characterized in that: The training data in the optimized training dataset is input into the domain training model, and the cross entropy loss is calculated based on the domain knowledge data labels.
9. A computer-readable storage medium, characterized in that: A program for generating large model data expansion and training is stored thereon, and when the program is executed by a processor, the method for generating large model data expansion and training as described in one of claims 1 to 8 is implemented.
10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, it implements the method for generating large model data expansion and training as described in one of claims 1 to 8.