Multi-stage semantic reordering fine tuning method and device applied to vertical field

Through the multi-stage semantic reordering fine-tuning method, combined with large language model, difficult negative sample mining and knowledge distillation technology, the problems of extradomain data interference, negative sample quality and generalization ability of vertical domain reordering models are solved, achieving higher discrimination ability and adaptability, and reducing computing resource consumption.

CN120181097AActive Publication Date: 2025-06-20CHINA DATANG GRP DIGITAL TECH CO LTD

Patent Information

Application Number
CN202510252651.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-20
Estimated Expiration
2045-03-05

AI Technical Summary

Technical Problem

When used in vertical fields, existing reordering models have problems such as large extraterritorial data interference, low negative sample quality, and poor generalization capabilities of the model. It is difficult to effectively capture subtle semantic differences in the field, and the computing resources are consumed and the optimization efficiency is low.

Method used

The multi-stage semantic reordering fine-tuning method is adopted to guide the big model to generate training data by designing a propt template, and multi-stage hard-negative sample mining is used to mine multiple stages with vector models and reordering models. Knowledge distillation technology is introduced, and the reordering model is optimized with a multi-stage training strategy combining full-parameter fine-tuning and LoRA fine-tuning.

Benefits of technology

It significantly improves the discriminant ability and generalization ability of the model in the vertical field, reduces the impact of extraterritorial interference, improves the adaptability and accuracy of the model, and effectively balances the computing efficiency and model performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120181097A_ABST
    Figure CN120181097A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-stage semantic reordering fine tuning method and device applied to the vertical field. The method comprises the following steps: S1, receiving original document input and segmenting document segments; s2, designing a prompt template to guide a large model to generate training data based on a high-quality question-answer pair labeled by a domain expert; s3, performing multi-stage difficult-to-load sample mining by using the vector model and the reordering model; s4, generating a knowledge distillation signal based on the teacher model; and S5, optimizing the reordering model by adopting a multi-stage training strategy combining all-parameter fine tuning and LoRA fine tuning. According to the method, the performance of the reordering model in the vertical field can be effectively improved. Particularly, the two-stage fine tuning strategy not only ensures that the model can fully learn global knowledge, but also optimizes difficult sample scenes in a targeted manner; the knowledge distillation mechanism helps the model to inherit the discrimination ability of a teacher model, and a good effect is achieved while the light weight of the model is kept.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a multi-stage semantic re-ranking fine-tuning method and device applied to vertical fields. Background Art

[0002] With the rapid development of artificial intelligence technology, information retrieval and ranking systems are becoming increasingly widely used in various vertical fields. Traditional retrieval ranking models usually adopt a single-stage approach, that is, directly ranking the retrieval results at one time. However, in highly specialized vertical fields such as medical, legal, and financial, in these tasks, how to effectively rank candidate answers or retrieval results to ensure higher accuracy and relevance is a challenging technical problem.

[0003] Currently, certain progress has been made in the research of re-ranking methods, mainly single-stage fine-tuning methods based on machine learning. The single-stage fine-tuning method directly optimizes the ranking performance by training the model once. However, this method often has certain limitations when facing complex vertical field tasks. Due to the particularity of vertical field data, a single training process cannot fully capture the diverse features within the field, resulting in poor performance of the model during application.

[0004] In the prior art, some patents have proposed model fine-tuning methods based on re-ranking. For example, some methods use contrastive learning or reinforcement learning for multi-stage optimization, but these methods often only focus on optimization at the data level and ignore the improvement of the model's capabilities and their complementarity in different task stages. Therefore, how to effectively combine the objectives and data characteristics of each stage in multi-stage training to improve the re-ranking effect of vertical field tasks remains an urgent problem to be solved. Summary of the Invention

[0005] In order to solve the problems existing in the prior art: existing re-ranking models have problems such as large interference from out-of-domain data, low quality of negative samples, and poor generalization ability when applied in vertical fields. In particular, traditional re-ranking models often use random sampling or simple similarity calculation to construct negative samples, and this method is difficult to obtain challenging negative samples, resulting in easy misjudgment of the model in actual applications. At the same time, existing methods lack in-depth understanding of specific knowledge in vertical fields and are difficult to accurately capture subtle semantic differences within the field. In addition, traditional model training methods consume a large amount of computing resources and have low model optimization efficiency, and are difficult to quickly adapt to the specific requirements of vertical fields.

[0006] The present invention provides a multi-stage semantic re-ranking fine-tuning method and device applied to vertical fields. By optimizing the fine-tuning strategy of each stage and making full use of data and task characteristics, the application performance of the model in vertical fields is improved. The specific solutions are as follows:

[0007] A multi-stage semantic re-ranking fine-tuning method applied to vertical fields, comprising the following steps:

[0008] S1, receiving the original document input and performing segmentation;

[0009] S2, designing a prompt template based on a small amount of high-quality question-and-answer pairs annotated by domain experts to guide the large model to generate training data;

[0010] S3, using a vector model and a re-ranking model to perform multi-stage hard negative sample mining;

[0011] S4, generating a knowledge distillation signal based on a teacher model;

[0012] S5, adopting a multi-stage training strategy combining full-parameter fine-tuning and LoRA fine-tuning to optimize the re-ranking model.

[0013] Preferably, the steps of designing a prompt template to guide the large model to generate training data in step S2 include:

[0014] S21, the domain expert selects representative document fragments from the original document for question-and-answer pair annotation;

[0015] S22, designing a prompt template including domain background information, role setting, answer format requirements, quality control rules, and input-output examples;

[0016] S23, inputting the prompt template and the document fragment into the large model to generate question-and-answer pair data;

[0017] S24, evaluating the information content of the document fragment, and outputting "null" for fragments with insufficient information content.

[0018] Preferably, the steps of using a vector model and a re-ranking model to perform multi-stage hard negative sample mining in step S3 include:

[0019] S31, preparing training data including questions and corresponding positive samples and a candidate sample pool;

[0020] S32, using a semantic embedding model to generate text embedding vectors;

[0021] S33, using FAISS to build an index and performing semantic retrieval to recall the TOP100 candidate samples;

[0022] S34, using a re-ranking model to perform refined ranking on the candidate samples;

[0023] S35, screening negative samples from the ranking results and selecting the TOP K as training negative samples.

[0024] Preferably, the step of screening negative samples from the sorting results in step S35 includes:

[0025] S351, using the re-ranking model trained in the previous stage to mine negative samples again;

[0026] S352, calculating the scores of positive samples and all negative samples;

[0027] S353, when the score difference between the positive sample and the highest negative sample is less than the preset threshold, using this sample as a hard negative sample for the next stage of training.

[0028] Preferably, the step of generating the knowledge distillation signal in step S4 includes:

[0029] S41, selecting a re-ranking model with a model parameter scale greater than 1B as the teacher model;

[0030] S42, scoring the positive and negative samples in the training data;

[0031] S43, adding the scoring result as a knowledge distillation signal to the training data.

[0032] Preferably, step S5 collectively includes the following steps:

[0033] S51, performing full-parameter fine-tuning on the model in the first stage to optimize the weighted sum of the cross-entropy loss and the knowledge distillation loss;

[0034] S52, in the second stage, mining hard negative samples based on the model in the first stage, using the LoRA technique to fine-tune the query, key, value, and dense layers, and at the same time performing full-parameter fine-tuning on the classifier layer.

[0035] Preferably, the knowledge distillation loss is calculated by the following formula:

[0036]

[0037] where N is the batch size, M is the group size, t ij is the output distribution target of the teacher model, and p ij is the output distribution of the student model.

[0038] Preferably, the fine-tuning device based on the above multi-stage semantic re-ranking fine-tuning method includes:

[0039] A data generation module for generating training data based on domain expert annotation data;

[0040] A negative sample mining module for performing multi-stage hard negative sample mining;

[0041] A knowledge distillation module for generating knowledge distillation signals;

[0042] A model training module for performing multi-stage fine-tuning training.

[0043] Preferably, the data generation module includes:

[0044] A document processing unit for segmenting and processing input documents;

[0045] A prompt design unit for designing prompt templates to guide large models;

[0046] A data generation unit for generating training data using large models.

[0047] The present invention also discloses a computer device, including a processor and a memory, where a computer program is stored in the memory, and when the computer program is executed by the processor, the steps of the multi-stage reordering fine-tuning method described in any one of the above are implemented.

[0048] The beneficial effects of the present invention are as follows:

[0049] First, through the construction of high-quality question-and-answer pairs and the hard negative sample mining strategy, the discriminative ability of the model in the vertical domain is significantly improved. Especially the multi-stage iterative negative sample mining mechanism enables the model to gradually learn more fine-grained semantic discrimination ability and effectively reduces the impact of out-of-domain interference.

[0050] Second, the introduction of the knowledge distillation mechanism enables the model to better learn the knowledge representation of the teacher model and significantly improves the generalization ability of the model. Through knowledge transfer at the distribution level, the model not only learns the discriminative ability at the sample level but also masters more abstract domain knowledge representation, enabling it to better handle unseen query scenarios.

[0051] Third, the innovative multi-stage fine-tuning strategy effectively balances computational efficiency and model performance. Compared with traditional full-parameter fine-tuning methods, the solution of the present invention significantly reduces the consumption of computational resources while maintaining similar performance. Especially the LoRA technology adopted in the second stage enables the model to quickly adapt to the specific requirements of the vertical domain and provides an efficient solution for the iterative optimization of the model.

[0052] Fourth, the overall solution has strong scalability and can be flexibly applied to different vertical domain scenarios. By adjusting the prompt template, negative sample mining strategy, and fine-tuning parameters, the solution can quickly adapt to new domain requirements. Practice shows that this method has been successfully applied to multiple professional fields such as medical, legal, and financial, demonstrating good adaptability and stability.

[0053] In addition, the method of the present invention exhibits excellent stability and reliability in practical applications, meeting the stability requirements of the production environment. At the same time, the modular design of this solution makes it easy to maintain and upgrade, providing technical support for sustainable development of the intelligent retrieval system in the vertical field. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0055] Figure 1 It is a flowchart of generating Q&A pair data for training based on a large model in an embodiment of the present invention.

[0056] Figure 2 It is a flowchart of mining hard negative samples based on a vector model and a re-rank model in an embodiment of the present invention.

[0057] Figure 3 It is a flowchart of fine-tuning a multi-stage re-rank model based on knowledge distillation in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0059] As can be seen from the background art, there are the following problems in the process of fine-tuning the re-rank model in the vertical field: First, it is difficult to obtain training data, and the cost of manual annotation is high and the efficiency is low; second, improper selection of negative samples may lead to a decline in model performance, and it is difficult to mine effective hard negative samples; third, there is a lack of an effective knowledge distillation mechanism in the model training process, and it is difficult to fully utilize the knowledge of the teacher model to optimize the performance of the student model.

[0060] To solve the above technical problems, the present invention provides a multi-stage semantic re-ranking fine-tuning method and device for the vertical field. This method innovatively combines advanced technologies such as large language models, hard negative sample mining, and knowledge distillation. First, by designing a specific prompt template to guide the large language model, a large-scale question-and-answer pair training data is automatically constructed based on a small amount of high-quality manually annotated samples, effectively reducing the data acquisition cost. Second, a multi-stage iterative approach is adopted to mine hard negative samples. Through the pre-trained model for preliminary recall and re-ranking, the most valuable negative samples are screened out, and on this basis, multiple rounds of iterative optimization are carried out to continuously improve the quality of negative samples. At the same time, the knowledge distillation technology is introduced, and a teacher model with stronger performance is used to guide the training process of the student model, enabling it to learn richer knowledge representations. Finally, by adopting the LoRA fine-tuning strategy, only the key parameters of the model are updated, which not only ensures the improvement of model performance but also significantly reduces the consumption of computing resources. The experimental results show that the method of the present invention can significantly improve the adaptability and accuracy of the re-ranking model in the vertical field, providing an efficient and feasible solution for solving the text matching problem in the vertical field.

[0061] As Figure 1 shown, a method for generating question-and-answer pair data for training based on a large model provided in this embodiment aims to quickly generate high-quality question-and-answer pairs by combining a small amount of manually annotated data with the capabilities of the large model for subsequent fine-tuning and training of the model. Specifically, the method includes the following steps:

[0062] S1, Receive the input of the original document and segment the document fragments.

[0063] Load the input original document into the system. The document format supports multiple common types such as PDF, TXT, and WORD.

[0064] Use the open-source tool MinerU to parse the document. MinerU is an efficient document parsing tool that supports extracting content from unstructured documents and has good multi-format compatibility. Through this tool, the body text, title information, and table content of the document can be parsed into a structured data format (such as JSON).

[0065] During the parsing process, for PDF documents, text extraction, paragraph separation, etc. can be achieved; for documents containing tables, the content of the tables and related title information can be extracted.

[0066] Segment the parsed document content. The segmentation granularity determines the size and content of each chunk fragment.

[0067] Use the semantic chunk method to segment the text. Semantic chunk is a semantic-based text chunking technique that can ensure the semantic integrity of the segmented segments and avoid affecting the subsequent generation effect due to overly short segments or semantic interruptions.

[0068] When segmenting, each chunk segment needs to contain the following key information: text content (text): the main information part of the chunk, containing specific semantic content; main document title (main_title): the main title information of the document, used to provide global context; table title (table_title, if any): if the chunk segment contains table data, the title information of the table needs to be extracted as context reference.

[0069] The format is as follows:

[0070] {"text": "[text content]", "main_title": "[main document title]", "table_title": "[table title]"}

[0071] S2. Design a prompt template based on high-quality question-answer pairs annotated by domain experts to guide the large model to generate training data.

[0072] S21. Domain experts annotate a small number of high-quality question-answer pairs. Through the professional knowledge and experience of domain experts, provide examples of high-quality question-answer pairs to provide directional guidance and quality standards for generating more question-answer pair data through the large model in the future.

[0073] Domain experts are personnel with professional knowledge in relevant fields. Their responsibility is to extract key information on the basis of fully understanding the document content, select representative and information-rich chunk segments from the original document for manual annotation of question-answer pairs. Prioritize segments containing important facts, statistical data, core business processes, or key technical descriptions to ensure the content of the annotated data is highly relevant and practical.

[0074] S22. Based on the segmented chunk segments and the high-quality data annotated by domain experts, design a set of refined prompt templates to guide the large model to generate high-quality question-answer pair data.

[0075] The design of the prompt template includes the following key contents:

[0076] Domain background information: Provide the background information of a specific domain required for generating question-answer pairs;

[0077] Role setting: Clearly define the role of the large model, such as "RAG question-answer pair generation expert";

[0078] Answer format requirements: Standardize the format of generated Q&A pairs, including the correspondence between questions and answers;

[0079] Quality control rules: Establish rules for information screening, text understanding, and handling special situations;

[0080] Input examples and generation examples: Use a small amount of high-quality manually annotated data as reference examples to provide clear generation directions for large models.

[0081] One of the prompt designs is as follows:

[0082] #Role: RAG Q&A pair generation expert

[0083] ##Profile

[0084] You are a RAG (Retrieval-Augmented Generation) example generator, responsible for generating accurate and detailed questions corresponding to the provided document fragments, focusing on generating high-quality Q&A pairs from the given document fragments. You need to generate a most appropriate question based on the provided text fragment (text), main title (main_title), and table title (table_title if any) information, and use the original text as the answer.

[0085] ##Knowledge

[0086] - Question generation: Questions should be accurate, specific, and able to fully cover the core information in the text fragment

[0087] - Text understanding: Accurately understand the text content to ensure that the generated questions are highly relevant to the text content

[0088] - Data processing: Be able to process text containing HTML tags and data in tabular form

[0089] - Information screening: Be able to judge whether the text fragment contains sufficient information to generate meaningful questions

[0090] ##Rules

[0091] 1. The input data format contains the following fields:

[0092] - text: The text content (required)

[0093] - main_title: The main title of the document (required)

[0094] - table_title: The table title (optional)

[0095] 2. Question generation rules:

[0096] -You need to extract information from the three fields text, main_title, and table_title to generate a question that accurately reflects the key information in the document. The question should be as detailed and specific as possible.

[0097] - Questions should contain necessary contextual information (such as year, company name, etc.)

[0098] - For tabular data, generate questions targeting specific data items

[0099] - For business data, generate questions that reflect the complete business situation

[0100] - The generated questions should be helpful for further knowledge retrieval or text generation tasks, so the questions should lead to detailed answers.

[0101] 3. Handling of special situations:

[0102] - If the text fragment contains very little information (such as only a title), output "null"

[0103] -If the text contains HTML tags, you need to extract the actual text content before generating questions

[0104] 4. Output format:

[0105] -Output is in JSON format, including two fields: question and answer

[0106] - If the text fragment contains very little information (such as only a title), output "null"

[0107] ##Example (specific data content omitted here)

[0108] Input 1:

[0109] {“text”:“[text content]”,“main_title”:“[main title of document]”,“table_title”:“[table title]”}

[0110] Output 1:

[0111] {"question":"[manually annotated question]","answer":"[manually annotated fragment answer]"}

[0112] Input 2:

[0113] {“text”: “[text content]”, “main_title”: “[main title of document]”}

[0114] Output 2:

[0115] {"Question": "[Manually marked question]", "Answer": "[Manually marked fragment answer]"}

[0116] Input 3:

[0117] {"text": "[Text content]", "main_title": "[Main title of the document]"}

[0118] Output 3:

[0119] "null"

[0120] ##Output Format

[0121] {"Question": "[Specific question generated based on the document content]", "Answer": "[Original text fragment]"}

[0122] Or

[0123] "null"

[0124] S23. Input the prompt template and the document fragment into the large model to generate Q&A pair data. That is, input the designed prompt together with the information (text, main_title, table_title) in the chunk fragment into the large model. The large model generates the corresponding Q&A pair data through the understanding and reasoning of the input information.

[0125] S24. Before the large model generates Q&A pairs, it is necessary to evaluate the information volume of the content in the chunk fragment to determine whether it contains sufficient meaningful information.

[0126] For the chunk fragment with sufficient information volume, the large model generates Q&A pair data in the specified format. If the information volume of the chunk fragment is insufficient, such as only containing a title or no specific content, then directly generate "null" as the output result.

[0127] S3. Use the vector model and the re-ranking model to perform multi-stage hard negative sample mining. As Figure 2 shown Figure 2 is a schematic flowchart of a method for hard negative sample mining in an embodiment of the present invention. The method of the present invention includes the following steps:

[0128] S31, Data Preparation: Prepare the training data including questions and corresponding positive samples, as well as the candidate sample pool. That is, prepare the input file input_file and the candidate sample pool candidate_pool. Among them, input_file is the dataset for training, containing the query information of the questions and the pos information of the corresponding positive samples. This file can be converted from the QA pairs in the format of {"question": "[Manually annotated question]", "answer": "[Manually annotated fragment answer]"} generated above; candidate_pool is an optional extended data source, containing a set of possible negative sample candidates, providing basic data support for subsequent mining of hard negative samples.

[0129] The data format required for input_file is

[0130] {"query": str, "pos": List[str]}

[0131] The data format required for candidate_pool is

[0132] {"text": "[Text content]", "main_title": "[Document main title]", "table_title": "[Table title]"}

[0133] S32, Generate text embedding vectors using a semantic embedding model. Specifically:

[0134] Model Loading: Load a semantic embedding model (embedding model), such as the bge-large-zh-v1.5 model, which is used to convert the input text data into a fixed-length vector representation, providing semantic embedding support for subsequent semantic retrieval and ranking tasks.

[0135] Embedding Generation: Process the input data in the data preparation stage to generate the embedding vectors (p_vecs) of the candidate texts and the embedding vectors (q_vecs) of the queries respectively. Among them, the embedding representation of the candidate texts represents the semantic features of all texts in the sample pool, and the query embedding represents the semantic features for each query.

[0136] S33, Use FAISS to build an index and perform semantic retrieval to recall the top 100 candidate samples. Specifically:

[0137] FAISS Index Construction and Search: Use the FAISS (Facebook AI Similarity Search) tool to construct a semantic retrieval index, and use the query embedding vectors (q_vecs) to perform fast semantic search on the embedding vectors (p_vecs) of the candidate sample pool, recalling the TOP100 candidate samples that are semantically closest to each query, providing input data for the subsequent re-ranking step.

[0138] S34. Use the re-ranking model to perform fine-grained ranking on the candidate samples. Specifically:

[0139] Re-ranking means loading the re-ranking model, such as the bge-reranker-large model, to perform further fine-grained ranking on the candidate samples recalled by FAISS. According to the semantic relevance, re-score and rank a specified number of samples in the recalled samples, and select the top-ranked samples from them.

[0140] S35. Filter negative samples from the ranking results and select the TOP K as training negative samples. Specifically:

[0141] Negative sample mining means further filtering negative samples from the re-ranked samples, mainly including the following two sub-steps: 1. Filter negative samples: Remove the positive sample information that is positively relevant to the query from the re-ranked samples, and only retain the negative sample data; 2. Select TOP K negative samples from the filtered negative sample set as the negative samples for training.

[0142] It should be noted here that the present invention involves multi-stage hard negative sample mining. The steps of filtering negative samples from the ranking results include:

[0143] S351. Use the re-ranking model trained in the previous stage to mine negative samples again, and extract a small number of hard negative samples for fine-tuning training. S352. The rule for screening a small number of hard negative samples is to calculate the scores of positive samples and all negative samples at the same time. S353. If the score difference between the positive sample and the highest negative sample is less than a pre-set threshold, then this sample is considered a hard negative sample and is used for training in the next stage.

[0144] Output result: Output the final result of negative sample mining in the form of a file, providing data support for the subsequent model fine-tuning training.

[0145] The final output data format is

[0146] {"query":str,"pos":List[str],"neg":List[str]}

[0147] S4. Generate knowledge distillation signals based on the teacher model.

[0148] After obtaining the positive and negative sample training data in this embodiment, it is also necessary to add distillation signals to the training data, including the following steps: S41, select a re-ranking model with a model parameter scale greater than 1B as the teacher model; S42, score the positive and negative samples in the training data; S43, add the scoring results as knowledge distillation signals to the training data.

[0149] Specifically, input the training data set. The training data set is Figure 2 the data containing question and positive / negative sample information obtained after execution. Select appropriate teacher and student models. For example, bge-reranker-v2-minicpm-layerwise with a larger number of parameters and better performance can be selected as the teacher model, and the lighter bge-reranker-large can be selected as the student model. The teacher model is responsible for generating distillation signals, and it has stronger capabilities and can provide more useful information, while the student model gradually learns the knowledge of the teacher model through contrastive learning and the distillation process. At this stage, the output of the teacher model is used as the learning target of the student model to improve the performance of the student model through the distillation process. The final output data format is

[0150] {"query":str,"pos":List[str],"neg":List[str],"pos_scores":List[int],"neg_scores":List[int]}

[0151] where pos_scores and neg_scores are the scores calculated by the teacher model.

[0152] S5, optimize the re-ranking model using a multi-stage training strategy that combines full-parameter fine-tuning and LoRA fine-tuning. As Figure 3 shown, Figure 3 is a flowchart of the method for fine-tuning a re-ranking model based on knowledge distillation according to an embodiment of the present invention. The method of the present invention includes the following steps:

[0153] Step 1) Construct an input data set. The basic format of the input data is a dictionary structure containing "query" and "positive / negative sample word vector sequences", and the format is as follows:

[0154] {"query":str,"pos":List[str],"neg":List[str],"pos_scores":List[int],"neg_scores":List[int]}

[0155] Among them, query is the query text, pos is the set of positive samples related to the query text, neg is the set of negative samples related to the query text, and pos_scores and neg_scores are the matching scores of positive and negative samples respectively, which are used for knowledge distillation training.

[0156] Step 2) Sample the samples according to the group size. In this step, the input positive and negative samples are grouped and sampled according to the preset group size, and each group contains a query and its corresponding positive and negative sample pairs. Specifically, for each query, the positive sample list and the negative sample list are merged and randomly sampled with group size samples to form a set of training data.

[0157] Step 3) Concatenate (tokenizer) the input query with the positive / negative sample word vector sequences to construct the following structure:

[0158] [CLS] query word vector sequence [SEP] positive / negative sample word vector sequence [SEP], where [CLS] is the classification identifier and [SEP] is the separator identifier.

[0159] Step 4) Re-rank model encoding. Adopt a two-stage fine-tuning strategy. In the first stage, full-parameter fine-tuning is performed on all training data, and in the second stage, LoRA fine-tuning is performed on a small number of selected hard negative samples. Specifically, it includes:

[0160] S51, in the first stage, fine-tune all parameters of the model, and optimize the model by minimizing the weighted sum of the cross-entropy loss and the knowledge distillation loss. The knowledge distillation loss uses the soft labels of the teacher model as the supervision signal to help the student model learn the knowledge of the teacher model.

[0161] S52, in the second stage, based on the model trained in the first stage, further mine negative samples, select a small number of hard negative samples to construct new training data. The specific screening method is described in detail in Figure 2 Step B6. Then use the LoRA fine-tuning technique to adjust some parameters of the model. Only perform LoRA fine-tuning on the query, key, value, and dense layers, and keep full-parameter fine-tuning on the classifier layer, and freeze the remaining parameters to improve the model's discriminative ability for hard negative samples.

[0162] Step 5) Classifier prediction. The re-rank model after two-stage fine-tuning predicts the relevance of the input query-document pair and outputs the relevance score. The classifier performs binary classification based on the high-dimensional feature representation encoded by the model to determine whether the document is relevant to the query.

[0163] Step 6) Calculate the loss and optimize the model. During training, joint optimization is performed by combining the cross-entropy loss and the knowledge distillation loss. Among them, the cross-entropy loss is used to supervise the model to learn correct relevance judgments, and the knowledge distillation loss guides the student model to imitate the prediction distribution of the teacher model, making full use of the knowledge in the teacher model.

[0164] Specifically, the knowledge distillation loss is achieved by minimizing the cross-entropy between the output distributions of the student model and the teacher model after softmax. The mathematical expression is:

[0165]

[0166] where N is the batch size, M is the group size, t ij is the output distribution target of the teacher model, and p ij is the output distribution of the student model. Minimizing this loss function makes the output of the student model as close as possible to that of the teacher model.

[0167] Through the above steps, the method of the present invention can effectively improve the performance of the re-ranking model in the vertical field. In particular, the two-stage fine-tuning strategy not only ensures that the model can fully learn global knowledge but also can specifically optimize difficult sample scenarios; while the knowledge distillation mechanism helps the model inherit the discriminative ability of the teacher model and achieves good results while keeping the model lightweight.

[0168] The present invention significantly improves the discriminative ability of the model in the vertical field through the construction of high-quality question-answer pairs and the difficult negative sample mining strategy. Especially the multi-stage iterative negative sample mining mechanism enables the model to gradually learn more fine-grained semantic discrimination ability and effectively reduces the impact of out-of-domain interference.

[0169] Secondly, the introduction of the knowledge distillation mechanism enables the model to better learn the knowledge representation of the teacher model and significantly improves the generalization ability of the model. Through knowledge transfer at the distribution level, the model not only learns the discriminative ability at the sample level but also masters more abstract domain knowledge representation, enabling it to better handle unseen query scenarios.

[0170] Thirdly, the innovative multi-stage fine-tuning strategy effectively balances computational efficiency and model performance. Compared with the traditional full-parameter fine-tuning method, the solution of the present invention significantly reduces the consumption of computing resources while maintaining similar performance. Especially the LoRA technology adopted in the second stage enables the model to quickly adapt to the specific needs of the vertical field and provides an efficient solution for the iterative optimization of the model.

[0171] Fourth, the overall solution has strong scalability and can be flexibly applied to different vertical domain scenarios. By adjusting the prompt template, negative sample mining strategy, and fine-tuning parameters, the solution can quickly adapt to the new domain requirements. Practice has shown that this method has been successfully applied to multiple professional fields such as medical, legal, and financial, demonstrating good adaptability and stability.

[0172] In addition, the method of the present invention shows excellent stability and reliability in practical applications, meeting the stability requirements of the production environment. At the same time, the modular design of the solution makes it easy to maintain and upgrade, providing technical support for the sustainable development of intelligent retrieval systems in vertical domains.

[0173] The above-described embodiments merely represent several implementation manners of the present invention. Their descriptions are relatively specific and detailed, but should not be construed as limiting the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the appended claims. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multi-stage semantic re-ranking fine-tuning method applied to vertical fields, characterized in that: The following steps are involved: S1, receives the original document input and divides the document fragments; S2, based on high-quality question-answer pairs annotated by domain experts, designs prompt templates to guide large models to generate training data; S3, multi-stage hard negative sample mining using vector model and re-ranking model; S4, generates knowledge distillation signals based on the teacher model; S5, a multi-stage training strategy combining full parameter fine-tuning and LoRA fine-tuning is used to optimize the re-ranking model.

2. The method according to claim 1, characterized in that: The steps of designing a prompt template in step S2 to guide the large model to generate training data include: S21, domain experts select representative document fragments from the original documents to annotate question-answer pairs; S22, design a prompt template that includes domain background information, role settings, answer format requirements, quality control rules, and input and output examples; S23, input the prompt template and document fragment into the big model to generate question-answer pair data; S24, evaluate the information content of the document fragments, and output "null" for the fragments with insufficient information.

3. The method according to claim 1, characterized in that Step S3 uses the vector model and the re-ranking model to perform multi-stage hard negative sample mining, including: S31, prepare training data including questions and corresponding positive samples and a candidate sample pool; S32, generate text embedding vector using semantic embedding model; S33, use FAISS to build an index and perform semantic retrieval to recall the TOP100 candidate samples; S34, use the re-ranking model to fine-tune the ranking of candidate samples; S35, filter negative samples from the sorting results and select TOP K as training negative samples.

4. The method according to claim 3, characterized in that Step S35 of screening negative samples from the sorting results includes: S351, use the re-ranking model trained in the previous stage to mine negative samples again; S352, calculating the scores of the positive sample and all negative samples; S353: When the score difference between the positive sample and the highest negative sample is less than a preset threshold, the sample is used as a hard negative sample for the next stage of training.

5. The method according to claim 1, characterized in that The step of generating the knowledge distillation signal in step S4 includes: S41, select a re-ranking model with a model parameter scale greater than 1B as the teacher model; S42, scoring the positive and negative samples in the training data; S43, adding the scoring results as knowledge distillation signals to the training data.

6. The method according to claim 1, characterized in that Step S5 collectively includes the following steps: S51, in the first stage, all parameters of the model are fine-tuned to optimize the weighted sum of cross entropy loss and knowledge distillation loss; S52, the second stage mines hard negative samples based on the first stage model, uses LoRA technology to fine-tune the query, key, value and dense layers, and fine-tunes all parameters of the classifier layer.

7. The method according to claim 6, characterized in that The knowledge distillation loss is calculated by the following formula: Where N is the batch size, M is the group size, and t ij is the output distribution target of the teacher model, p ij Output distribution for the student model.

8. A fine-tuning device based on the multi-stage semantic reordering fine-tuning method according to claims 1-7, characterized in that: include: Data generation module, used to generate training data based on data annotated by domain experts; Negative sample mining module, used for multi-stage difficult negative sample mining; A knowledge distillation module, used to generate knowledge distillation signals; Model training module, used to perform multi-stage fine-tuning training.

9. The fine-tuning device according to claim 8, characterized in that: The data generation module comprises: A document processing unit, used for segmenting and processing input documents; The prompt design unit is used to design prompt templates for guiding large models; The data generation unit is used to generate training data using the large model.

10. A computer device comprising a processor and a memory, characterized in that: The memory stores a computer program, which, when executed by the processor, implements the steps of the multi-stage reordering fine-tuning method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Semantic retrieval and question and answer processing method and device for long text and electronic equipment

    CN115630136A

  • Generation method and equipment of instruction fine tuning data and storage medium

    CN117763113A

  • Industrial standard data processing method and device based on vertical field large language model

    CN117852497A

  • Power accident event extraction method based on knowledge distillation and preference optimization

    CN118780249A

  • Intelligent question answering method based on vertical domain knowledge

    CN118820444A

Cited By

  • Lightweight target detection knowledge distillation method

    CN121119044A

  • Wind power operation and maintenance knowledge management method, device, equipment and medium

    CN121144534A