Multi-modal large model optimization method for vertical domain
Through automated data construction and quality evaluation, combined with alignment training and iterative optimization, the problem of poor task performance of vertical large models in professional fields is solved, and efficient and low-cost model development and continuous optimization are achieved.
Patent Information
- Application Number
- CN202510484949.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing vertical field large models have poor performance when handling professional field tasks, difficulty in obtaining and building data, inefficient optimization efficiency, complex multimodal processing, and difficult alignment training, resulting in long model development cycle, high cost and difficult to quickly iterate and optimize.
Through automated data construction, quality evaluation, alignment training and iterative optimization, a multimodal large model optimization method for vertical domain is established, and the big model is guided to generate question-and-answer pairs using prompt word templates, setting evaluation rules to filter high-quality data, building an alignment training set, and optimizing the model through iterative training.
It significantly reduces the cost and cycle of vertical domain model development, improves model performance, realizes the automated generation of high-quality training data, simplifies the alignment training process, enhances multimodal processing capabilities, and supports the continuous iterative optimization of the model.
Smart Images

Figure CN120012945A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of large model technology, and in particular to a multi-modal large model optimization method for vertical domains. Background Art
[0002] In recent years, large language models (LLMs) represented by GPT, GLM, and LLaMA and multimodal large models (MLLMs) represented by GPT-4V, LLaVA, and Qwen-VL have made breakthrough progress in the field of general artificial intelligence. These models have demonstrated strong knowledge reserves and task processing capabilities, and can complete complex tasks including natural language understanding, text generation, code writing, and multi-round dialogues. On this basis, the development of large models in vertical fields has gradually become a research hotspot, such as legal large models and financial large models.
[0003] However, building high-performance large vertical domain models faces the following key challenges: 1. Insufficient domain adaptability: Existing large-scale models are usually trained on a wide range of datasets. Although these datasets cover multiple languages and modalities, they often lack in-depth mining and understanding of specific vertical domain knowledge. This results in the model's performance being unsatisfactory when dealing with professional domain tasks, such as traffic situation analysis, legal analysis, etc.
[0004] 2. Difficulty in data acquisition and construction: Professional data in vertical fields are usually scarce and scattered, making them difficult to obtain. In order to train these large models, a large amount of labeled data is required. The collection, cleaning, and labeling of this data is not only costly, but also the manually constructed questions are difficult to cover all the complex scenarios in the field. In addition, the quality and diversity of the data directly affect the performance of the model, which limits the generalization ability and application scope of the model.
[0005] 3. Low optimization efficiency: Existing large vertical model optimization methods often rely on a lot of manual participation, from data collection and annotation to model evaluation and iteration, each link requires in-depth participation of professionals. This not only leads to a long model development cycle and high cost, but also limits the possibility of rapid iteration and continuous optimization of the model.
[0006] 4. Complex multimodal processing: Compared with simple text models, multimodal large models need to process non-text data such as images and audio, and achieve cross-modal understanding and generation. This makes the training and optimization of vertical multimodal large models more complex, and puts higher requirements on data quality and training methods.
[0007] 5. Alignment training is difficult: Alignment training is required to ensure that the content generated by the large vertical model meets human preferences and industry standards. Traditional alignment training methods require a large amount of manually annotated alignment data, which is particularly difficult and costly in professional fields.
[0008] In the existing technology, the industry usually constructs vertical domain training data by manually collecting and labeling, and then uses supervised fine-tuning and reinforcement learning methods with manual feedback to optimize the model. This method is not only time-consuming and labor-intensive, but also difficult to continuously iterate and optimize, which ultimately affects the implementation effect of the vertical domain multimodal large model. Summary of the invention
[0009] In order to address the shortcomings of the prior art, the present invention aims to establish a complete vertical multimodal large model optimization method through automated data construction, quality assessment, alignment training and iterative optimization, which not only greatly reduces the development cost and cycle of the vertical model, but also significantly improves the model performance, laying a technical foundation for the widespread application of vertical multimodal large models.
[0010] In order to achieve the above-mentioned invention object, the technical solution provided by the present invention includes: The multi-modal large model optimization method for vertical domain includes the following steps: S1. Select a target multimodal large model suitable for the target vertical domain; obtain knowledge data of the target vertical domain to construct a first data set; S2. Constructing a prompt word template to guide the first large model to generate a set of question-answer pairs based on the first data set and the task features of the target vertical domain; S3. Setting evaluation rules to guide the second largest model to perform quality evaluation on the question-answer pairs in the question-answer pair set, and screening the question-answer pairs with quality evaluation higher than a preset threshold to construct a target question-answer pair set; S4. taking the target question-answer pair set as an alignment positive sample, and guiding the third largest model to generate an alignment negative sample based on the target question-answer pair set, and combining the alignment positive sample and the alignment negative sample into an alignment training set; S5. Use the target question-answer pair set to complete the supervised fine-tuning training of the target large model, and use the alignment training set to complete the alignment training of the target large model; S6. Repeat steps S2-S5 to complete the iterative training of the target large model.
[0011] Preferably, step S2 comprises: S21. Constructing a first prompt word template for guiding the first large model to extract a first question-answer pair set from the knowledge of the first data set; S22. Based on the task characteristics of the target vertical domain, construct a second prompt word template to guide the first large model to generate a query question set according to the task characteristics, and generate a second question-answer pair set according to the query question set.
[0012] Preferably, the method for generating aligned negative samples in step S4 includes: extracting questions from the question-answer pair set as a first question set, and inputting the first question set into a third largest model to generate aligned negative samples.
[0013] Preferably, step S3 further includes: adding the target question-answer pair set to the first data set, repeating steps S2-S3, and iteratively optimizing the target question-answer pair set.
[0014] Preferably, the method for acquiring knowledge data in step S1 includes: Obtain keywords of the target vertical domain, retrieve and obtain public data on the Internet based on the keywords, construct a third prompt word template to guide the fourth model to generate a third question-answer pair set based on the public data on the Internet, and add it to the first data set as knowledge data.
[0015] Preferably, the first large model, the second large model and the third large model are any one of the multimodal large model and the target large model that have completed training.
[0016] Preferably, the first data set contains at least one of text data, image data, audio data and / or video data; when the first data set contains non-text data, the prompt word template includes instructions for guiding the first large model to analyze and describe the non-text data, and generate question-answer pairs based on the analysis and description.
[0017] Beneficial Effects 1. Significantly reduce data construction costs: The present invention effectively solves the problem of difficulty in obtaining data in professional fields by automatically constructing vertical domain knowledge datasets using existing large models. In particular, the prompt word template is used to guide the first large model to generate a set of question-answer pairs based on the first dataset and task features, which greatly reduces the cost of manual annotation and realizes the automatic generation of high-quality training data.
[0018] 2. Improve data quality: This invention sets evaluation rules to guide the second largest model to evaluate the quality of the generated question-answer pairs. The threshold screening mechanism ensures the high quality of training data and avoids the negative impact of low-quality data on model training. This automated data quality assessment method replaces the traditional manual screening process, which not only improves efficiency but also ensures the objectivity and consistency of data quality.
[0019] 3. Optimize the alignment training process: This paper innovatively proposes a method to use the target question-answer pair set as alignment positive samples, and based on this, guide the third largest model to generate alignment negative samples, and automatically construct an alignment training set. This solves the problem of difficulty in obtaining alignment training data and ensures that the model behavior is consistent with human preferences and industry standards.
[0020] 4. Achieve iterative self-optimization: By repeatedly executing the process of data generation, evaluation and screening, and training optimization, the present invention establishes a self-optimizing closed-loop system. This iterative training mechanism enables the model to continuously learn and improve, gradually improving performance in specific vertical domains, and greatly reducing the need for manual intervention.
[0021] 5. Improve multimodal processing capabilities: This invention specifically designs a specific prompt word template for the processing of multimodal data, which can guide the model to analyze and describe non-text data and generate high-quality question-answer pairs based on this. This significantly enhances the model's ability to process and understand multimodal data, making it perform better in vertical domain applications that include multiple modalities such as images and audio.
[0022] 6. Wide adaptability: The method of the present invention has wide adaptability and can be applied to multiple vertical fields such as law, medical care, finance, education, etc. By simply adjusting the prompt word template and evaluation rules, it can be optimized for the specific needs of different fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 It is a schematic flow chart of a multi-modal large model optimization method for vertical domain provided in a preferred embodiment of the present invention. DETAILED DESCRIPTION
[0024] In order to make the purpose, technical solution and advantages of the present invention clearer, the present invention is further described below in conjunction with the accompanying drawings. In the description of the present invention, it should be understood that the terms "upper", "lower", "front", "back", "left", "right", "top", "bottom", "inside", "outside" and the like indicate directions or positional relationships based on the directions or positional relationships shown in the accompanying drawings, which are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operated in a specific direction, and therefore cannot be understood as limiting the present invention.
[0025] like Figure 1 As shown, this embodiment provides a multi-modal large model optimization method for vertical domain, including the steps of: S1. Select a target multimodal large model suitable for the target vertical domain; obtain knowledge data of the target vertical domain to construct a first data set.
[0026] This step first requires the selection of a multimodal large model suitable for the target vertical domain, which is the basis and starting point of the entire optimization method. Vertical domain refers to a specific vertical professional field, such as law, medical care, finance, education, engineering and other fields with professionalism and special knowledge systems. Each vertical domain has its own unique terminology, knowledge structure and application scenarios.
[0027] Multimodal large models refer to large-scale artificial intelligence models that can simultaneously process and understand multiple data forms such as text, images, audio, video, etc., such as GPT-4V, LLaVA, Qwen-VL, etc. These models have acquired powerful general capabilities through pre-training, but their understanding of specific vertical domains needs further optimization.
[0028] Acquiring knowledge data of the target vertical domain is a key step in constructing the first data set. These knowledge data include, but are not limited to, professional books, academic papers, industry reports, case studies, professional image materials, and related audio and video materials, etc. They contain core professional knowledge and application cases in the vertical domain. Those skilled in the art can know that common methods for constructing the first data set include directly collecting relevant information from professional databases and literature libraries, crawling domain-related data from public websites, inviting domain experts to provide professional knowledge content, extracting text information from image documents through OCR technology, and extracting structured information from multimodal data using computer vision technologies such as target detection.
[0029] In some preferred embodiments, a method for enriching the first data set by using keyword retrieval to obtain public network data is also provided. Specifically, the method for obtaining knowledge data includes: obtaining keywords of the target vertical domain, obtaining public network data based on the keyword retrieval, constructing a third prompt word template to guide the fourth model to generate a third set of question-answer pairs based on the public network data, and adding it to the first data set as knowledge data. Those skilled in the art can implement the above method by crawling. The method provided in this embodiment greatly improves the efficiency of data collection and the quality of data, and lays a solid data foundation for subsequent model optimization.
[0030] When the first data set contains non-text data such as pictures, tables, audio or video, the design of the prompt word template is particularly critical. It needs to accurately guide the first large model to perform professional analysis and description of these complex data, and generate high-quality question-answer pairs based on these analyses. Therefore, in some preferred embodiments, when the first data set contains non-text data, the prompt word template includes instructions for guiding the first large model to analyze and describe the non-text data, and generate question-answer pairs based on the analysis and description. This prompt word template usually includes two key parts: first, it guides the model to identify and parse the visual or auditory elements in the non-text data, such as "Please analyze the main content of this picture and provide a detailed description" or "Please identify the key scenes and objects in this video"; second, it guides the model to generate question-answer pairs related to the vertical domain based on the analysis results, such as "Based on your analysis, please generate 5 questions that professionals may ask and their detailed answers." In practical applications, this template-driven approach has two main implementation paths: one is to provide visual information such as images directly as input data to the multimodal large model, use the model's own cross-modal understanding ability to automatically generate questions that users may ask, and then generate professional and accurate answers based on these questions; the other method is to first extract text information from the image with the help of OCR technology, or use computer vision technology such as target detection to identify key objects and relationships in the image, and then provide these structured visual element information to the multimodal large model to guide it to generate more accurate question-answer pairs. This method of processing non-text data significantly expands the diversity and richness of training data, enabling the optimized vertical multimodal large model to better understand and process complex scenarios in real-world applications, such as medical image analysis, engineering drawing interpretation, traffic scene understanding, etc., thereby greatly improving the application value and practicality of the model in specific vertical domains.
[0031] S2. Construct a prompt word template to guide the first large model to generate a set of question-answer pairs based on the first data set and the task features of the target vertical domain.
[0032] The task characteristics refer to the types, characteristics and requirements of professional tasks that the model needs to perform in a specific vertical domain, such as imaging diagnosis and case interpretation in the medical vertical domain, regulatory interpretation and case analysis in the legal vertical domain, and drawing recognition and fault diagnosis in the engineering vertical domain. These task characteristics determine the professional capabilities and knowledge boundaries that the model needs to have, and directly affect the design of the prompt word template and the generation direction of the question and answer pairs.
[0033] A question-and-answer pair refers to a data unit consisting of a question and its corresponding answer. As the basic data form for model training and fine-tuning, high-quality question-and-answer pairs should have characteristics such as strong vertical professionalism, wide coverage, diverse question expressions, and professional and accurate answers.
[0034] It should be understood that the first large model, the second large model and the third large model used in the present invention may be the same model applied to different stages or for different purposes, or different models may be used for different links of the optimization process, which may be other multimodal large models that have completed training, or directly be the target model used as the optimization target. At this time, the specific configuration can be designed by those skilled in the art according to actual needs and actual conditions on site, and the present invention does not make further requirements.
[0035] The design of the prompt word template may include a brief introduction to the vertical domain background, a clear description of the task characteristics, format requirements for the generation of question and answer pairs, and requirements for professionalism and diversity. Those skilled in the art may know that for existing vertical domain-related materials, such as text materials in articles and books, key knowledge points can be directly extracted through the prompt word template to construct question pairs. Considering that the large model itself has a rich knowledge reserve, in some preferred embodiments, considering this feature, the model is triggered to autonomously generate user query questions, thereby enriching the content of the first data set. Specific methods include: S21. Construct a first prompt word template to guide the first large model to extract a first set of question-answer pairs from the knowledge of the first data set. This bottom-up method of directly extracting question-answer pairs can ensure that the generated question-answer pairs have a solid knowledge foundation and professional accuracy, and retain the professional terms, standard definitions, and domain specifications in the original data set.
[0036] S22. Based on the task characteristics of the target vertical domain, a second prompt word template is constructed to guide the first large model to generate a query question set according to the task characteristics, and generate a second question-answer pair set according to the query question set. This step adopts a top-down approach, making full use of the rich knowledge reserves of the large model itself, integrating the task description into the prompt words of the model, and triggering the model to autonomously generate user query questions. Then, based on the generated query questions, the model will execute the answer generation process to output the corresponding answers.
[0037] This dual-track parallel strategy design has multiple advantages: first, it ensures the diversity and comprehensiveness of the training data, covering both basic knowledge and practical application scenarios; second, the two methods can complement and verify each other, and the knowledge depth provided by S21 is combined with the application breadth provided by S22 to form a more three-dimensional knowledge system; third, this method can flexibly respond to the special needs of different vertical domains. For example, the medical field may require more question-and-answer pairs based on standard definitions, while the engineering field may need solutions based on practical problems; finally, the question-and-answer pairs generated in two different ways can better simulate the diverse query habits of actual users, thereby improving the performance of the model in real application environments.
[0038] S3. Setting evaluation rules to guide the second largest model to perform quality evaluation on the question-answer pairs in the question-answer pair set, and screening the question-answer pairs with quality evaluation higher than a preset threshold to construct a target question-answer pair set.
[0039] Those skilled in the art can know that the evaluation rules are essentially a set of multi-dimensional quality assessment standards, which usually include but are not limited to the following core dimensions: professional accuracy (whether the content of the question and answer pair conforms to the vertical domain professional knowledge and is error-free), relevance (whether the question and answer pair is closely related to the target vertical domain task characteristics), completeness (whether the answer comprehensively answers all aspects of the question), difficulty adaptability (whether the difficulty of the question can effectively train the various levels of capabilities of the model), clarity of expression (whether the question statement is clear and the answer is clear and coherent) and breadth of knowledge coverage (whether the overall question and answer pair set covers the key parts of the vertical domain knowledge system). The specific implementation process of quality evaluation usually adopts a multi-level scoring mechanism. The second largest model will quantitatively score each dimension of each question and answer pair (such as a 1-5 point system or a 1-10 point system), and may set weights for different dimensions, and finally calculate a comprehensive score. The preset threshold is a quality threshold set based on business needs and model performance goals. Only question and answer pairs with a comprehensive score or a key dimension score exceeding the threshold will be screened into the target question and answer pair set. This strict quality control mechanism ensures that the data ultimately used for model training is highly professional and practical, and can effectively improve the professional capabilities and application performance of the target large model in the vertical domain. At the same time, this evaluation mechanism itself is also an optimizable process, and the evaluation dimensions, weights and thresholds can be continuously adjusted according to the feedback on the model training effect, forming a positive cycle of quality evaluation, and continuously improving the quality of question and answer pairs and model performance. It should be noted that the specific setting method of the evaluation rules and preset thresholds can be reasonably designed by those skilled in the art based on the content of the above-mentioned prior art and conventional methods in the field, combined with the characteristics of the target vertical domain, etc. It is not the focus of protection of the present invention, so it will not be repeated.
[0040] In some preferred embodiments, a dynamically evolving data enhancement strategy is also provided, which realizes the self-evolution and continuous optimization of the data set by returning a high-quality target question-answer pair set to the original first data set and then repeating the cycle of data generation and quality evaluation. Specifically, step S3 also includes: adding the target question-answer pair set to the first data set, repeating steps S2-S3, and iteratively optimizing the target question-answer pair set.
[0041] This mechanism builds a positive feedback loop: the high-quality question-answer pairs selected in each round of iteration not only enrich the scale of the original data set, but more importantly, improve its quality density and knowledge depth, so that the next round of question-answer pairs generated based on this can be built on a higher starting point. In this way, the data set presents an exponential quality improvement effect like a snowball. Each round of iteration can not only generate new knowledge points, but also derive more complex and in-depth question-answer pairs based on existing knowledge, forming a multi-level, multi-angle, and multi-granular knowledge structure network. This iterative optimization process is particularly suitable for vertical domain model training, because vertical domain knowledge often has an inherent progressive relationship and hierarchical structure. The question-answer pairs generated by subsequent iterations can capture the logical connection and application transformation between this knowledge, so that the model can not only master single-point knowledge, but also understand the overall framework and application context of the knowledge system, thereby showing stronger adaptability and problem-solving capabilities in practical applications.
[0042] S4. Use the target question-answer pair set as aligned positive samples, and guide the third largest model to generate aligned negative samples based on the target question-answer pair set, and set the aligned positive samples and the aligned negative samples as an aligned training set.
[0043] The alignment refers to the process of adjusting the model behavior to conform to human values, preferences and expectations, ensuring that the model output content is not only accurate but also in line with human ethical and professional standards. This step constructs an efficient and automated alignment data generation mechanism, which enables the model to learn human preferences through a positive and negative sample comparison learning framework. Specifically, the high-quality target question-answer pair set that has been strictly screened in the previous step is directly used as the alignment positive sample, which represents the professional, accurate and valuable answer mode in the vertical domain; while the alignment negative sample is guided by the third largest model (usually a smaller model) to generate suboptimal answers based on the same question, which may have various deficiencies in professionalism, accuracy, completeness or expression. This positive and negative sample construction method cleverly avoids the cost bottleneck of a large number of manual annotations in traditional RLHF (reinforcement learning based on human feedback), and realizes the large-scale automatic generation of alignment data.
[0044] In some preferred embodiments, a method for generating negative samples based on a target question-answer pair set is provided, specifically comprising: extracting questions from the question-answer pair set as a first question set, and inputting them into a third large model to generate aligned negative samples. It is worth noting that the prompt word template of the negative sample may include intentionally guiding the model to generate answers that are not professional enough, not comprehensive enough, not logically clear, or deviated from the focus, thereby covering various possible types of low-quality answers.
[0045] The alignment training set constructed by the above method can fully capture the human preference characteristics for high-quality answers, allowing the model to clearly distinguish between "good answers" and "bad answers" through comparative learning, thereby internalizing human value judgment standards. In practical applications, this automated alignment data construction method can significantly improve the professionalism, reliability and practical value of vertical domain model answers, while greatly reducing the resource consumption of traditional manual annotation, providing an efficient and feasible technical path for the precise alignment of large models in professional fields.
[0046] S5. Use the target question-answer pair set to complete the supervised fine-tuning training of the target large model, and use the alignment training set to complete the alignment training of the target large model.
[0047] Supervised Fine-Tuning (SFT) is a guided learning process that uses a high-quality set of target question-answer pairs as training samples to enhance the model's professional capabilities in a specific vertical domain by minimizing the gap between the model's predictions and the standard answers. At this stage, the model learns how to transform domain knowledge into accurate and professional answers, absorbing domain-specific terminology, logical reasoning methods, and problem-solving frameworks. This process essentially allows the model to master what to do and how to do it, consolidating its vertical domain professional capabilities.
[0048] Alignment Training is a deeper optimization process that uses the positive and negative sample comparison datasets constructed in the previous steps to teach the model to distinguish between high-quality answers and low-quality answers, so that the model can internalize human preference standards. Unlike supervised fine-tuning, alignment training focuses more on the value judgment dimension of "why this is better", so that the model not only has the ability to answer questions correctly, but also can express these answers in the way humans expect - for example, clearer, more comprehensive, organized, with appropriate professional depth, while avoiding simple, vague or wrong response patterns.
[0049] This step embodies the "two-track" training idea: supervised fine-tuning injects professional knowledge and skills into the model, while alignment training shapes the form and style of the model output to ensure that it meets human expectations. This combined optimization strategy ensures that the final trained vertical domain large model not only understands professional knowledge, but also can express this knowledge in the standards and manner of human experts, thereby demonstrating higher practical value and user satisfaction in actual applications. Compared with the traditional method of relying solely on large-scale manual annotation for RLHF, this method of automatically constructing training data and combining SFT with alignment training is more efficient and is particularly suitable for vertical domain model development scenarios with limited resources.
[0050] S6. Repeat steps S2-S5 to complete the iterative training of the target large model. A positive feedback loop is formed through iterative training: the improvement of model capabilities after each round of training will directly affect the quality of data screening in the next round of iteration, thereby optimizing the screening criteria for question-answer pairs and producing higher-quality training corpus; at the same time, the model's deepening understanding of vertical domain knowledge can also help generate more professional data expansion samples and build more discriminative aligned positive and negative sample pairs. It is worth noting that iterative training is not an infinite loop, and usually clear termination conditions are set, such as the performance of the validation set reaches the preset target, the improvement of the evaluation index for multiple consecutive rounds is lower than the threshold, or the predetermined number of iterations is reached. The limit enables the model to gradually approach or even surpass the professional level of human experts in a specific vertical domain through a closed-loop process of multiple rounds of data screening, expansion, alignment and fine-tuning, while maintaining the humanity and understandability of the response. Technicians in this field can conduct a comprehensive evaluation after each round of iteration, including automated testing and manual sampling review, to ensure that the model is steadily improving in the expected direction and to prevent overfitting or capability degradation caused by over-optimization.
[0051] The above shows and describes the basic principles, main features and advantages of the present invention. It should be understood by those skilled in the art that the present invention is not limited to the above embodiments, and the above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention to be protected. The scope of protection of the present invention is defined by the attached claims and their equivalents.
Claims
1. A multi-modal large model optimization method for vertical domain, characterized by: Includes steps: S1. Select a target multimodal large model suitable for the target vertical domain; obtain knowledge data of the target vertical domain to construct a first data set; S2. Constructing a prompt word template to guide the first large model to generate a set of question-answer pairs based on the first data set and the task features of the target vertical domain; S3. Setting evaluation rules to guide the second largest model to perform quality evaluation on the question-answer pairs in the question-answer pair set, and screening the question-answer pairs with quality evaluation higher than a preset threshold to construct a target question-answer pair set; S4. taking the target question-answer pair set as an alignment positive sample, and guiding the third largest model to generate an alignment negative sample based on the target question-answer pair set, and combining the alignment positive sample and the alignment negative sample into an alignment training set; S5. Use the target question-answer pair set to complete the supervised fine-tuning training of the target large model, and use the alignment training set to complete the alignment training of the target large model; S6. Repeat steps S2-S5 to complete the iterative training of the target large model.
2. The vertical domain-oriented multi-modal large model optimization method according to claim 1, characterized in that: Step S2 includes: S21. Constructing a first prompt word template for guiding the first large model to extract a first question-answer pair set from the knowledge of the first data set; S22. Based on the task characteristics of the target vertical domain, construct a second prompt word template to guide the first large model to generate a query question set according to the task characteristics, and generate a second question-answer pair set according to the query question set.
3. The vertical domain-oriented multi-modal large model optimization method according to claim 1, characterized in that: The method for generating aligned negative samples in step S4 includes: extracting questions from the question-answer pair set as a first question set, and inputting the first question set into a third largest model to generate aligned negative samples.
4. The vertical domain-oriented multi-modal large model optimization method according to claim 1, characterized in that: Step S3 also includes: adding the target question-answer pair set to the first data set, repeating steps S2-S3, and iteratively optimizing the target question-answer pair set.
5. The vertical domain-oriented multi-modal large model optimization method according to claim 1, characterized in that: The method for acquiring knowledge data in step S1 includes: Obtain keywords of the target vertical domain, retrieve and obtain public data on the Internet based on the keywords, construct a third prompt word template to guide the fourth model to generate a third question-answer pair set based on the public data on the Internet, and add it to the first data set as knowledge data.
6. The vertical domain-oriented multi-modal large model optimization method according to claim 1, characterized in that: The first large model, the second large model, and the third large model are any one of the multimodal large models and the target large model that have completed training.
7. The vertical domain-oriented multi-modal large model optimization method according to claim 1, characterized in that: The first data set includes at least one of text data, image data, audio data and / or video data; when the first data set includes non-text data, the prompt word template includes instructions for guiding the first large model to analyze and describe the non-text data, and generate question-answer pairs based on the analysis and description.
Citation Information
Patent Citations
Model fine tuning method and device, electronic equipment and storage medium
CN119337832A
Visual language feature fine alignment method for medical multi-mode large model
CN119357443A
Multi-modal model generation method, multi-modal processing method, and device
WO2025031090A1
Cited By
Training method and device of vertical domain question and answer model, equipment and storage medium
CN121117615A