Big language model-based double-stage fine tuning training method and device
Through the two-stage fine-tuning training method based on the large language model, the large language model is initially fine-tuned and re-fine-tuned, which solves the problem that the existing technology is difficult to achieve efficient search and accurate correlation judgment in complex queries and massive data processing, and achieves more efficient and accurate information retrieval effects.
Patent Information
- Application Number
- CN202510145586.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-27
AI Technical Summary
In the prior art, when processing complex queries and massive data, it is difficult to achieve efficient search and accurate correlation judgment, especially when understanding the context, keyword-based matching algorithms perform poorly.
A two-stage fine-tuning training method based on a large language model is adopted to obtain the correlation judgment data set and the correlation selection data set, and the large language model is initially fine-tuned and re-fine-tuned, so that it can master the basic correlation judgment ability and correlation judgment ability in complex environments.
It improves the accuracy and robustness of searches, enhances the accuracy of information retrieval, and can adapt to complex application environments and recall the most relevant information.
Smart Images

Figure CN120045749A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of neural networks, and particularly to a two-stage fine-tuning training method and device based on a large language model. Background Art
[0002] In the application scenarios of retrieval question answering and knowledge base construction, information retrieval technology faces huge challenges. With the explosive growth of knowledge base data, traditional retrieval methods are unable to cope when dealing with complex queries and massive data. Large language models (such as GPT-4), due to their powerful natural language processing capabilities, have become potential tools to solve these problems.
[0003] However, there are still some technical obstacles in directly applying large language models for retrieval relevance judgment. Users need to quickly find the content most relevant to their queries from a large amount of content, such as finding specific scenarios or conversations. This not only requires the retrieval system to have efficient search capabilities but also accurate relevance judgment capabilities. Currently, many retrieval systems rely on keyword-based matching algorithms, which work well for simple queries but perform poorly in the face of complex queries or when context understanding is required. In recent years, retrieval-augmented generation (RAG) technology based on large language models has gradually emerged. Through vector language retrieval, it attempts to improve the accuracy and efficiency of retrieval. However, the results retrieved on a large scale have similarities with the semantics of the query but lack sufficient discrimination, making it difficult to recall the most relevant information for refined retrieval requirements. Summary of the Invention
[0004] To solve the above technical problems or at least partially solve the above technical problems, this application provides a two-stage fine-tuning training method and device based on a large language model.
[0005] In a first aspect, this application provides a two-stage fine-tuning training method based on a large language model, the method comprising:
[0006] Obtain a relevance judgment data set and a relevance selection data set, wherein both the relevance judgment data set and the relevance selection data set include triples composed of queries, positive samples, and negative samples, and the complexity of the relevance selection data set is higher than that of the relevance judgment data set;
[0007] Perform preliminary fine-tuning on the pre-trained large language model using the relevance judgment data set, wherein the preliminary fine-tuning is used to enable the large language model to master basic relevance judgment capabilities;
[0008] Perform re-fine-tuning on the preliminarily fine-tuned large language model using the relevance selection data set, wherein the re-fine-tuning is used to enable the large language model to master relevance discrimination capabilities in complex environments;
[0009] Deploy the large language model after fine-tuning to an information retrieval system.
[0010] Optionally, obtaining a relevance judgment dataset includes:
[0011] Obtain a relevance data corpus, where each query in the relevance data corpus corresponds to at least zero positive samples and at least zero negative samples;
[0012] Randomly sample from the relevance data corpus to generate a simple triple containing one query, one positive sample, and one negative sample;
[0013] Fill the simple triple into a preset prompt template, where the prompt template is used to standardize the input format of the simple triple;
[0014] Fill the relationship between the query and the positive sample and the negative sample into a label template, where the label template is used to represent the relevance between the query and the sample;
[0015] Use the filled prompt template as input data and the filled label template as the label of the input data to form the relevance judgment dataset.
[0016] Optionally, obtaining a relevance selection dataset includes:
[0017] Adjust the relevance data corpus, and generate a complex triple containing one query, multiple positive samples, and multiple negative samples by random sampling;
[0018] Fill the complex triple into a preset prompt template, and fill the relationship between the query and multiple positive and negative samples into a label template;
[0019] Use the filled prompt template as input data and the filled label template as the label of the input data to form the relevance selection dataset.
[0020] Optionally, adjusting the relevance data corpus includes:
[0021] Adjust the queries in the relevance data corpus to simple queries, complex queries, or fuzzy queries; or,
[0022] Adjust the queries in the relevance data corpus to long-tail queries or multi-modal queries; or,
[0023] Add context information to the queries and samples in the relevance data corpus, where the context information is used to enable the large language model to understand complex semantic relationships.
[0024] Optionally, the preliminary fine-tuning of the pre-trained large language model using the correlation judgment data set includes:
[0025] Obtain multiple pre-trained candidate large language models, where the number of parameters of the candidate large language model is less than a set threshold;
[0026] Use the correlation judgment data set to perform preliminary fine-tuning on multiple candidate large language models respectively, where semi-precision training is used during the preliminary fine-tuning;
[0027] Select the candidate large language model with the best fine-tuning effect as the large language model after preliminary fine-tuning.
[0028] Optionally, after deploying the fine-tuned large language model to the information retrieval system, the method further includes:
[0029] Input the target query content into the fine-tuned large language model;
[0030] Perform inference using mixed precision to obtain the information query result output by the large language model.
[0031] In a second aspect, the present application provides a two-stage fine-tuning training device based on a large language model, and the device includes:
[0032] An acquisition module for acquiring a correlation judgment data set and a correlation selection data set, where both the correlation judgment data set and the correlation selection data set include triples composed of queries, positive samples, and negative samples, and the complexity of the correlation selection data set is higher than that of the correlation judgment data set;
[0033] A first fine-tuning module for performing preliminary fine-tuning on the pre-trained large language model using the correlation judgment data set, where the preliminary fine-tuning is used to enable the large language model to master basic correlation judgment capabilities;
[0034] A second fine-tuning module for performing re-fine-tuning on the preliminarily fine-tuned large language model using the correlation selection data set, where the re-fine-tuning is used to enable the large language model to master correlation discrimination capabilities in complex environments;
[0035] A deployment module for deploying the fine-tuned large language model to the information retrieval system.
[0036] Optionally, the acquisition module is used for:
[0037] Obtain a correlation data corpus, where each query in the correlation data corpus corresponds to at least zero positive samples and at least zero negative samples;
[0038] Randomly sample from the correlation data corpus to generate a simple triple containing a query, a positive sample, and a negative sample;
[0039] Fill the simple triple into a preset prompt template, where the prompt template is used to standardize the input format of the simple triple;
[0040] Fill the relationship between the query and the positive sample and the negative sample into a label template, where the label template is used to represent the correlation between the query and the sample;
[0041] Use the filled prompt template as input data and the filled label template as the label of the input data to form the correlation judgment data set.
[0042] In a third aspect, an electronic device is provided, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus;
[0043] The memory is used to store computer programs;
[0044] The processor is used to implement the steps of any of the above-mentioned two-stage fine-tuning training methods based on the large language model when executing the program stored on the memory.
[0045] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned two-stage fine-tuning training methods based on the large language model are implemented.
[0046] The above technical solutions provided by the embodiments of the present application have the following advantages compared with the prior art:
[0047] In the initial fine-tuning stage of the method provided by the embodiments of the present application, the large language model is made to learn to judge the correlation between the query and the sample through a simple correlation judgment data set. In the re-fine-tuning stage, the large language model is made to learn to select relevant content in complex scenarios through a complex correlation selection data set. In this way, the trained large language model can adapt to complex application environments, enhancing the accuracy and robustness of retrieval and improving the precision of information retrieval. Description of the Drawings
[0048] The drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments that conform to the present invention and are used together with the specification to explain the principles of the present invention.
[0049] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0050] Figure 1 Schematic diagram of the hardware environment of a two-stage fine-tuning training method based on a large language model provided by an embodiment of the present application;
[0051] Figure 2 Flowchart of a two-stage fine-tuning training method based on a large language model provided by an embodiment of the present application;
[0052] Figure 3 Schematic diagram of the structure of a two-stage fine-tuning training device based on a large language model provided by an embodiment of the present application;
[0053] Figure 4 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts belong to the scope of protection of the present application.
[0055] In the subsequent description, the suffixes such as "module", "component", or "unit" used to represent elements are only for the convenience of description of the present application, and they have no specific meaning in themselves. Therefore, "module" and "component" can be used interchangeably.
[0056] To solve the problem of difficult to accurately recall the most relevant information mentioned in the background art, the present application enables the large language model to have the ability to judge relevance in complex environments by adopting a two-stage fine-tuning strategy, thereby improving the accuracy of recalling relevant information.
[0057] The embodiments of the present application are applicable to various scenarios that require high-precision information retrieval and relevance judgment, including but not limited to: document query, product recommendation, video recommendation, etc.
[0058] Optionally, in the embodiments of the present application, the above two-stage fine-tuning training method based on a large language model can be applied to Figure 1 the hardware environment composed of the terminal 101 and the server 103 as shown inFigure 1 As shown in the figure, the server 103 is used to train the large language model by means of two-stage fine-tuning, and then deploy the fine-tuned large language model into the information retrieval system. Users can perform information retrieval through the information retrieval system on the terminal. The database 105 can be set on the server or independently of the server to provide data storage services for the server 103. The above network includes but is not limited to: wide area network, metropolitan area network or local area network. The terminal 101 includes but is not limited to PC, mobile phone, tablet computer, etc.
[0059] The two-stage fine-tuning training method based on the large language model in the embodiments of the present application can be executed by the server 103.
[0060] Next, in combination with specific implementation manners, a two-stage fine-tuning training method based on the large language model provided by the embodiments of the present application will be described in detail. As Figure 2 shown, the specific steps are as follows:
[0061] Step 201: Obtain a relevance judgment data set and a relevance selection data set. Among them, both the relevance judgment data set and the relevance selection data set include triples composed of queries, positive samples, and negative samples. The complexity of the relevance selection data set is higher than that of the relevance judgment data set.
[0062] Step 202: Use the relevance judgment data set to perform preliminary fine-tuning on the pre-trained large language model. Among them, the preliminary fine-tuning is used to enable the large language model to master basic relevance judgment capabilities.
[0063] Step 203: Use the relevance selection data set to perform secondary fine-tuning on the preliminarily fine-tuned large language model. Among them, the secondary fine-tuning is used to enable the large language model to master the relevance discrimination capabilities in complex environments.
[0064] Step 204: Deploy the fine-tuned large language model into the information retrieval system.
[0065] In step 201, when building an efficient retrieval system, a large amount of relevant data corpus needs to be collected first. These corpora usually come from the queries of users in various information retrieval tasks and their corresponding retrieval contents. The retrieval contents include but are not limited to documents, pictures, videos, commodities, etc. For example, the relevant data corpus includes the query terms input by users in the search engine and the document summaries in the search results. To ensure the quality and format consistency of the data, the system will clean and organize these raw data. This includes steps such as removing duplicates, handling missing values, and standardizing text formats. Through this series of data preprocessing operations, it can be ensured that the data set finally used to train the large language model has high quality and consistency.
[0066] Based on the organized data, the system extracts a relevance judgment dataset and a relevance selection dataset from the organized corpus. Both datasets consist of triples containing queries, positive samples, and negative samples, but they are each applicable to query environments of different complexities. Each dataset is described separately below.
[0067] The triples in the relevance judgment dataset help the large language model understand the basic concept of relevance. The relevance judgment dataset is mainly used for the initial fine-tuning of the large language model to enable it to master the basic relevance judgment ability. Through training with a large number of triples, the large language model can gradually establish an understanding of basic semantic relationships, laying a foundation for subsequent more complex training.
[0068] The relevance selection dataset not only contains queries, positive samples, and negative samples but also adds complexity, enabling the large language model to be trained in a more complex query environment and being applicable to more complex query environments.
[0069] A triple is a structured way to represent a query and its related positive and negative samples, including the query, positive sample, and negative sample. Among them, the query refers to the search term or phrase input by the user to express an information need, the positive sample refers to a material segment highly relevant to the query, and the negative sample refers to a material segment irrelevant to the query.
[0070] Exemplarily, the triple includes the following content:
[0071] Query: "Science fiction novels about time travel".
[0072] Positive sample: "AA" is a science fiction novel about time travel, which tells the story of the protagonist solving a series of historical puzzles through time travel.
[0073] Negative sample: "BB" tells the story of a romantic encounter between two young people in Paris.
[0074] In step 202, in the initial fine-tuning stage, the system uses the relevance judgment dataset to train the pre-trained large language model. The goal of this stage is to enable the large language model to master the basic relevance judgment ability. Specifically, the large language model needs to learn how to distinguish the strong relevance between the query and the positive sample, as well as the weak relevance or irrelevance between the query and the negative sample, and the large language model can gradually establish an understanding of basic semantic relationships.
[0075] In step 203, after the initial fine-tuning is completed, the large language model already has basic relevance judgment capabilities. However, to handle more complex query scenarios, further fine-tuning is required. At this stage, the system uses a relevance selection dataset to further train the large language model after the initial fine-tuning. This dataset is characterized by a high degree of complexity, simulating diverse and complex query requirements in the real world. Through further fine-tuning, the large language model can better understand and process complex relationships.
[0076] In step 204, the system deploys the fine-tuned large language model into the information retrieval system, and then uses the deployed large language model for inference. Since the large language model has undergone strict training during both the initial fine-tuning and the further fine-tuning processes, it can accurately identify content relevant to the query and exclude irrelevant interference items, achieving high-precision retrieval capabilities.
[0077] This application provides a two-stage fine-tuning strategy. In the initial fine-tuning stage, the large language model learns to judge the relevance between queries and samples through a simple relevance judgment dataset. In the further fine-tuning stage, the large language model learns to select relevant content in complex scenarios through a complex relevance selection dataset. In this way, the trained large language model can adapt to complex application environments, enhancing the accuracy and robustness of retrieval and improving the precision of information retrieval.
[0078] As an optional implementation manner, in step 201, obtaining the relevance judgment dataset includes:
[0079] Step S11: Obtain a relevance data corpus, where each query in the relevance data corpus corresponds to at least zero positive samples and at least zero negative samples;
[0080] Step S12: Randomly sample from the relevance data corpus to generate a simple triple containing one query, one positive sample, and one negative sample;
[0081] Step S13: Fill the simple triple into a preset prompt template, where the prompt template is used to standardize the input format of the simple triple;
[0082] Step S14: Fill the relationship between the query, the positive sample, and the negative sample into a label template, where the label template is used to represent the relevance between the query and the sample;
[0083] Step S15: Use the filled prompt template as input data and the filled label template as the label of the input data to form a relevance judgment dataset.
[0084] The system first needs to obtain a large amount of relevant data corpus. Each query in this corpus corresponds to at least zero positive samples and at least zero negative samples, which means that some queries may not have exactly matching positive or negative samples. Then the system randomly samples from the relevant data corpus to generate simple triples containing one query, one positive sample, and one negative sample. This structured data form enables the large language model to learn basic relevance concepts and provides a high-quality data basis for subsequent training. An example of a simple triple is shown below.
[0085] Query: Science fiction novels about time travel.
[0086] Positive Sample: "AA" tells the story of humanity's first step on Mars, and the protagonist solves a series of historical puzzles through time travel.
[0087] Negative Sample: "BB" tells the romantic encounter of two young people in Paris.
[0088] After generating the simple triples, the next step is to fill these triples into a pre-set prompt template to standardize the input format. The role of the prompt template is to unify the format of the input data and ensure that the large language model can process data from different sources consistently. The filled prompt template is as follows.
[0089] Query: Science fiction novels about time travel.
[0090] Positive Sample: "AA" tells the story of humanity's first step on Mars, and the protagonist solves a series of historical puzzles through time travel.
[0091] Negative Sample: "BB" tells the romantic encounter of two young people in Paris.
[0092] To represent the relevance between the query and the samples, it is also necessary to fill the relationships between the query and the positive and negative samples into the label template. The label template is used to clearly label the relevance between the query and the samples, where 1 represents relevant and 0 represents irrelevant. The filled label template is as follows:
[0093] Query: Science fiction movies about space exploration.
[0094] Positive Sample: 1.
[0095] Negative Sample: 0.
[0096] The system uses the filled prompt template as input data and the filled label template as the label of the input data to form a relevance judgment dataset. To ensure the effectiveness and generalization ability of the large language model, the dataset is usually divided into a training set and a validation set. Specifically, 80% of the data in the dataset is used for training, and 20% of the data is used for validation. This division method helps to evaluate the performance of the large language model on unseen data and prevent overfitting problems.
[0097] In this application, using the relevance judgment dataset to perform preliminary fine-tuning on the pre-trained large language model can enable the large language model to master the basic relevance judgment ability. Specifically, through a large number of simple triple trainings, the large language model can gradually establish an understanding of basic semantic relationships, so as to accurately distinguish the strong relevance between the query and the positive sample, as well as the weak relevance or irrelevance between the query and the negative sample.
[0098] As an optional implementation manner, in step 201, obtaining the relevance selection dataset includes:
[0099] Step S21: Adjust the relevance data corpus, and generate complex triples including one query, multiple positive samples, and multiple negative samples through random sampling;
[0100] Step S22: Fill the complex triples into the preset prompt template, and fill the relationships between the query and multiple positive and negative samples into the label template;
[0101] Step S23: Use the filled prompt template as input data and the filled label template as the label of the input data to form a relevance selection dataset.
[0102] When constructing a more complex query environment, it is necessary to adjust the relevance data corpus to generate complex triples including one query, multiple positive samples, and multiple negative samples. This process extracts diverse queries and their corresponding content sets from the sorted data through random sampling. Each query not only corresponds to one positive sample and one negative sample, but may also involve multiple positive and negative samples. Examples of complex triples are as follows.
[0103] Query: Science fiction movies about space exploration with time travel plots.
[0104] Positive sample 1: "AA" tells the story of humanity's first step on Mars, and the protagonist solves many scientific problems through time travel.
[0105] Positive sample 2: "CC" is a story about future human exploration of the universe, and the protagonist discovers new planets through time travel.
[0106] Negative sample 1: "BB" tells the romantic encounter of two young people in Paris.
[0107] Negative sample 2: "DD" records the daily life of students at school.
[0108] This complex data structure enables large language models to be trained in more complex query environments, process diverse inputs, and improve their generalization ability in practical applications. In this way, large language models can learn how to distinguish between multiple relevant and irrelevant contents of a query, thereby enhancing the accuracy of their judgments.
[0109] To ensure the consistency and standardization of input data, the generated complex triples need to be filled into a pre-set prompt template. At the same time, to represent the relationships between a query and multiple positive and negative samples, these relationships also need to be filled into a label template, which is used to clearly label the relevance between a query and a sample. The filled prompt template is used as input data, and the filled label template is used as the label of the input data, constituting a relevance selection dataset. This dataset not only includes simple binary classification tasks (i.e., the relationship between a single query and a single positive and negative sample), but also covers more complex multi-classification tasks (i.e., the relationship between a single query and multiple positive and negative samples). This complex dataset design helps large language models learn to process diverse inputs and improve their generalization ability in practical applications.
[0110] Similarly, the relevance selection dataset will also divide the dataset into a training set and a validation set. Specifically, 80% of the data in the dataset is used for training, and 20% of the data is used for validation.
[0111] In this application, through a large number of complex triple trainings, large language models can better understand and process the differences between multiple positive and negative samples, thereby providing high-precision retrieval results. This not only improves the relevance of search results, but also reduces the occurrence of irrelevant or low-quality results, thereby enhancing the relevance and accuracy of the information retrieval system.
[0112] As an optional implementation manner, in step S21, adjusting the relevance data corpus includes:
[0113] Adjusting the queries in the relevance data corpus to simple queries, complex queries, or fuzzy queries; or,
[0114] Adjusting the queries in the relevance data corpus to long-tail queries or multi-modal queries; or,
[0115] Adding context information to the queries and samples in the relevance data corpus, where the context information is used to enable large language models to understand complex semantic relationships.
[0116] To ensure that constructing a relevant selection dataset is more complex than a relevance judgment dataset, the system needs to introduce more variables and complexities during the dataset construction process, specifically including the following.
[0117] 1. Diverse query types.
[0118] Simple query: The goal of this type of query is very clear and it is easy to match relevant content. For example, "Science fiction novels about time travel".
[0119] Compound query: A query that contains multiple keywords or conditions, usually requiring comprehensive consideration of multiple factors to accurately match relevant content. For example, "Science fiction novels about time travel and with time travel plots".
[0120] Fuzzy query: A query with unclear intent or imprecise expression, which may contain some uncertain words or expressions. This type of query usually requires the large language model to have a certain reasoning ability to infer the actual needs of the user. For example, "Books similar to 'Interstellar Adventure'".
[0121] By adjusting queries to different types (such as simple queries, complex queries, or fuzzy queries), the large language model can better understand and process diverse user needs. For example, for simple queries, the large language model can directly match relevant content; while for complex queries or fuzzy queries, the large language model needs to comprehensively consider multiple factors and information forms to provide more accurate results. This flexibility enables the large language model to perform well in different application scenarios and meet the diverse needs of users.
[0122] 2. Simulating complex queries in real scenarios.
[0123] Long-tail query: A less common but very specific query, usually containing multiple keywords and having a relatively small search volume. For example, "Science fiction novels about time travel and containing elements of quantum mechanics".
[0124] Multimodal query: A query that combines multiple information forms (such as text, images, videos, etc.). This type of query not only relies on text information but may also involve data in other modalities.
[0125] By introducing long-tail queries, the large language model can capture those less common but very specific user needs. The large language model can find content that meets the requirements through detailed matching algorithms instead of simply returning popular results, which not only improves the relevance of search results but also enhances the user experience.
[0126] By supporting multimodal queries, large language models can process and understand different forms of information, such as text, images, videos, etc. Large language models not only need to search for relevant text content but also match corresponding image resources. This ability enables large language models to perform excellently in multimedia application scenarios, such as video recommendation systems, text-image search engines, etc.
[0127] 3. Introduce context information.
[0128] In addition to adjusting the query type, context information can also be added to the query and samples. Context information can help large language models better understand complex semantic relationships. For example, add more background information or context description in the document, "Interstellar Adventure is a work by the famous writer Zhang San. The book not only has exciting time-travel plots but also explores the future of humanity."
[0129] By adding context information to the query and samples, large language models can better understand the user's intentions and preferences. This context-aware ability improves the personalization level of the recommendation system and enhances user satisfaction and stickiness.
[0130] As an optional implementation method, in step 202, using the relevance judgment dataset to perform preliminary fine-tuning on the pre-trained large language model includes:
[0131] Step S31: Obtain multiple pre-trained candidate large language models, where the number of parameters of the candidate large language model is less than the set threshold;
[0132] Step S32: Use the relevance judgment dataset to perform preliminary fine-tuning on multiple candidate large language models respectively, where semi-precision training is used during the preliminary fine-tuning process;
[0133] Step S33: Select the candidate large language model with the best fine-tuning effect as the large language model after preliminary fine-tuning.
[0134] When using the relevance judgment dataset to perform preliminary fine-tuning on the pre-trained large language model, several suitable candidate large language models need to be selected first. To balance training computing resources and inference speed, pre-trained large language models with a smaller number of parameters are usually selected. For example, large language models such as glm-4-9b-chat or llama-3-8b-chat can be selected. These large language models not only have high performance but also can run efficiently with limited computing resources.
[0135] During the initial fine-tuning process, selecting an appropriate loss function is crucial for measuring the difference between the model's predictions and the true values. One of the commonly used loss functions is the Cross-Entropy Loss, which is widely applied in the fine-tuning tasks of large language models. The cross-entropy loss is used to calculate the distance between the predicted probability distribution and the true probability distribution, and can effectively measure the difference between the output of the large language model and the actual labels.
[0136] During the initial fine-tuning process, a reasonable training strategy has an important impact on the convergence speed and final performance of the large language model. Key fine-tuning strategies include parameter update strategies, learning rate adjustment strategies, early stopping strategies, and half-precision training strategies.
[0137] Parameter update strategy: Use optimization algorithms (such as Adam or SGD) to update the model parameters to minimize the loss function. The choice of optimization algorithm directly affects the convergence speed and stability of the model.
[0138] Learning rate adjustment strategy: Setting an appropriate learning rate is the key to controlling the training speed of the model. The initial learning rate is set to 1e-5 and adjusted dynamically according to the performance of the validation set. Using the cosine annealing strategy can reduce the learning rate in the later stage of training, helping the model to converge better.
[0139] Early stopping strategy: To avoid overfitting, an early stopping strategy is adopted. Stop training in advance when the performance of the validation set no longer improves. This method can prevent the model from overfitting the training data, thereby improving its generalization ability on unseen data.
[0140] Half-precision training strategy: Adopt half-precision floating-point numbers (such as FP16) for training. FP16 reduces the video memory occupancy compared to FP32, enabling larger models or larger batches to be trained under the same hardware conditions. In addition, FP16 has a faster calculation speed and can improve the training efficiency without affecting the model accuracy.
[0141] During the fine-tuning process, it is also necessary to monitor the performance metrics of the model (such as accuracy, recall, etc.) to ensure that the model is continuously improved during training. By regularly evaluating the performance of the model on the validation set, potential problems can be detected in a timely manner and adjusted.
[0142] As an alternative implementation, in step 203, the process of using the relevance selection dataset to further fine-tune the pre-fine-tuned large language model is similar to the initial fine-tuning process. In the initial fine-tuning stage, first, a large language model with the best performance needs to be selected. This large language model should perform well in the relevance judgment task and have good generalization ability. Then, the relevance selection dataset is used for further fine-tuning to simulate more complex retrieval scenarios. The fine-tuning strategy also includes parameter update strategy, learning rate adjustment strategy, early stopping strategy, and half-precision training strategy. The difference is that in a complex environment, it may be necessary to adjust the learning rate decay strategy to adapt to longer training requirements.
[0143] In this application, by using the relevance selection dataset for further fine-tuning, the large language model can perform excellently in more complex query environments. For example, for complex queries containing multiple positive and negative samples, the large language model can accurately distinguish relevant content from irrelevant content and provide high-precision results. This enhanced generalization ability improves the large language model's ability to retrieve relevant content.
[0144] As an alternative implementation, after deploying the fine-tuned large language model to the information retrieval system, the method further includes:
[0145] Input the target query content into the fine-tuned large language model;
[0146] Perform inference using mixed precision to obtain the information query result output by the large language model.
[0147] Before deploying the fine-tuned large language model to actual applications, the export and optimization of the large language model need to be carried out first. This process aims to reduce the model size, improve the inference speed, and ensure the effective utilization of computing resources. The optimization of the large language model refers to model quantization, that is, the process of converting model parameters from floating-point numbers (such as FP32) to integers (such as INT8). This not only reduces the storage requirements but also improves the inference speed. For example, an FP32 model that originally occupied a large amount of video memory can be reduced to one-fourth of its original size after INT8 quantization, and the inference speed can be increased by 2 - 4 times.
[0148] The deployment of the large language model can use the vllm framework as the deployment tool. The vllm framework is an efficient inference framework suitable for the deployment of large-scale language models. In addition, this application provides RESTful API interfaces for external systems to call. To ensure the compatibility and security of the interfaces, OpenAI-compatible API calls are also supported.
[0149] After the large language model is deployed, the inference process can be executed. The user inputs the target query content into the large language model, and the large language model performs inference using mixed precision, thereby outputting the information query result. Among them, using mixed precision inference (such as FP16) can further reduce the use of computing resources without affecting the model accuracy. Mixed precision inference combines the advantages of FP16 and FP32, can reduce the computational overhead while maintaining high precision, and can improve efficiency and reduce the inference time.
[0150] Based on the same technical concept, the embodiment of the present application also provides a two-stage fine-tuning training process based on a large language model, including the following contents:
[0151] 1. Data collection and preprocessing.
[0152] Step 1.1: Collect a large amount of relevant data corpus, which contains user queries and their corresponding contents. Each query corresponds to at least zero positive samples and at least zero negative samples.
[0153] Step 1.2: Clean and sort the original data, remove duplicates, handle missing values, standardize the text format, etc., to ensure the quality and consistency of the data.
[0154] 2. Construct a relevance judgment data set.
[0155] Step 2.1: Randomly sample from the relevant data corpus to generate a simple triple containing one query, one positive sample, and one negative sample.
[0156] Step 2.2: Fill the generated simple triple into a preset prompt template to unify the format of the input data.
[0157] Step 2.3: Fill the relationship between the query and the positive and negative samples into the label template to clearly label the relevance between the query and the samples.
[0158] Step 2.4: Use the filled prompt template as the input data and the filled label template as the label of the input data to form a relevance judgment data set.
[0159] 80% of the data in the data set is used for training, and 20% of the data is used for verification.
[0160] 3. Initial fine-tuning stage.
[0161] Step 3.1: Select a pre-trained large language model with a relatively small number of parameters (such as glm-4-9b-chat or llama-3-8b-chat), use the validation set to evaluate the performance of the model, and select the large language model with the best performance.
[0162] Step 3.2: Use the relevance judgment dataset to perform preliminary fine-tuning on the selected large language model.
[0163] 4. Construct a relevance selection dataset.
[0164] Step 4.1: Adjust the relevance data corpus to generate complex triples containing a query, multiple positive samples, and multiple negative samples.
[0165] Step 4.2: Fill the generated complex triples into the preset prompt template to unify the format of the input data.
[0166] Step 4.3: Fill the relationship between the query and multiple positive and negative samples into the label template to clearly label the relevance between the query and the samples.
[0167] Step 4.4: Use the filled prompt template as the input data and the filled label template as the label of the input data to form a relevance selection dataset. 80% of the data in the dataset is used for training, and 20% of the data is used for validation.
[0168] 5. The second fine-tuning stage.
[0169] Step 5.1: Select the large language model with the best validation effect in the preliminary fine-tuning stage.
[0170] Step 5.2: Use the relevance selection dataset to perform second fine-tuning on the selected large language model.
[0171] 6. Export and optimization of the large language model.
[0172] Step 6.1: Use model quantization techniques (such as INT8 quantization) to reduce the model size and improve the inference speed.
[0173] Step 6.2: Adopt mixed-precision inference (such as FP16) to reduce the use of computing resources.
[0174] The technical effects that this application can achieve include the following.
[0175] 1. By constructing a relevance judgment dataset containing queries, positive samples, and negative samples, as well as a more complex relevance selection dataset, high-quality training data is provided to ensure that the large language model can perform effective relevance judgment in diverse scenarios.
[0176] 2. Through two-stage fine-tuning training on the relevance judgment dataset and the relevance selection dataset, the relevance judgment ability of the large language model in complex query environments is improved, enhancing the accuracy and robustness of retrieval.
[0177] 3. By means of model quantization and mixed-precision inference, the inference speed is improved, the use of computing resources is reduced, enabling large language models to operate efficiently on more types of devices.
[0178] 4. Accurate retrieval results: Large language models can provide high-quality retrieval results, meeting the diverse needs of users, and improving the user experience and the overall performance of the system.
[0179] This application can be applied to the field of video and film materials. Among a vast amount of video and film materials, it can quickly locate the segments or scenes most relevant to the user's query, supporting users in efficiently searching for materials.
[0180] Based on the same technical concept, the embodiment of this application also provides a two-stage fine-tuning training device based on a large language model, as Figure 3 shown. This device includes:
[0181] An acquisition module 301, configured to acquire a relevance judgment data set and a relevance selection data set. Among them, both the relevance judgment data set and the relevance selection data set include triples composed of queries, positive samples, and negative samples, and the complexity of the relevance selection data set is higher than that of the relevance judgment data set;
[0182] A first fine-tuning module 302, configured to perform preliminary fine-tuning on the pre-trained large language model using the relevance judgment data set. Among them, the preliminary fine-tuning is used to enable the large language model to master basic relevance judgment capabilities;
[0183] A second fine-tuning module 303, configured to perform re-fine-tuning on the preliminarily fine-tuned large language model using the relevance selection data set. Among them, the re-fine-tuning is used to enable the large language model to master relevance discrimination capabilities in complex environments;
[0184] A deployment module 304, configured to deploy the fine-tuned large language model into an information retrieval system.
[0185] Optionally, the acquisition module 301 is used for:
[0186] Acquire a relevance data corpus, where each query in the relevance data corpus corresponds to at least zero positive samples and at least zero negative samples;
[0187] Randomly sample from the relevance data corpus to generate a simple triple containing one query, one positive sample, and one negative sample;
[0188] Fill the simple triple into a preset prompt template, where the prompt template is used to standardize the input format of the simple triple;
[0189] Fill the relationships between the query and the positive and negative samples into the label template, where the label template is used to represent the correlation between the query and the samples;
[0190] Use the filled prompt template as the input data and the filled label template as the label of the input data to form a correlation judgment data set.
[0191] Optionally, the obtaining module 301 is used for:
[0192] Adjust the correlation data corpus, and generate complex triples containing one query, multiple positive samples, and multiple negative samples through random sampling;
[0193] Fill the complex triples into the preset prompt template, and fill the relationships between the query and multiple positive and negative samples into the label template;
[0194] Use the filled prompt template as the input data and the filled label template as the label of the input data to form a correlation selection data set.
[0195] Optionally, the obtaining module 301 is used for:
[0196] Adjust the queries in the correlation data corpus to simple queries, complex queries, or fuzzy queries; or,
[0197] Adjust the queries in the correlation data corpus to long-tail queries or multi-modal queries; or,
[0198] Add context information to the queries and samples in the correlation data corpus, where the context information is used to enable the large language model to understand complex semantic relationships.
[0199] Optionally, the first fine-tuning module 302 is used for:
[0200] Obtain multiple pre-trained candidate large language models, where the number of parameters of the candidate large language models is less than the set threshold;
[0201] Use the correlation judgment data set to perform preliminary fine-tuning on multiple candidate large language models respectively, where half-precision training is used during the preliminary fine-tuning process;
[0202] Select the candidate large language model with the best fine-tuning effect as the large language model after preliminary fine-tuning.
[0203] Optionally, the device is further used for:
[0204] Input the target query content into the large language model after fine-tuning;
[0205] Perform inference using mixed precision to obtain the information query result output by the large language model.
[0206] Based on the same inventive concept, an embodiment of the present invention further provides an electronic device, such as Figure 4 shown, which includes a processor 401, a communication interface 402, a memory 403, and a communication bus 404. Among them, the processor 401, the communication interface 402, and the memory 403 complete mutual communication through the communication bus 404.
[0207] The memory 403 is used to store a computer program.
[0208] The processor 401, when executing the program stored on the memory 403, implements the above steps.
[0209] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.
[0210] The communication interface is used for communication between the above electronic device and other devices.
[0211] The memory may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0212] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0213] In another embodiment provided by the present invention, a computer-readable storage medium is further provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above methods are implemented.
[0214] In another embodiment provided by the present invention, a computer program product including instructions is further provided. When it runs on a computer, the computer is caused to execute any of the methods in the above embodiments.
[0215] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state disk (SSD)).
[0216] It should be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, article, or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device including the element.
[0217] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.
Claims
1. A two-stage fine-tuning training method based on a large language model, characterized in that: The method comprises: Acquire a relevance judgment data set and a relevance selection data set, wherein the relevance judgment data set and the relevance selection data set both contain triplets consisting of a query, a positive sample, and a negative sample, and the complexity of the relevance selection data set is higher than the complexity of the relevance judgment data set; Using the relevance judgment data set to perform preliminary fine-tuning on the pre-trained large language model, wherein the preliminary fine-tuning is used to enable the large language model to master basic relevance judgment capabilities; The relevance selection data set is used to further fine-tune the large language model after preliminary fine-tuning, wherein the further fine-tuning is used to enable the large language model to master the relevance discrimination ability in a complex environment; Deploy the fine-tuned large language model into the information retrieval system.
2. The method according to claim 1, characterized in that: Obtaining the relevance judgment data set includes: Obtaining a corpus of relevant data, wherein each query in the corpus of relevant data corresponds to at least zero positive samples and at least zero negative samples; Randomly sampling from the relevant data corpus to generate a simple triple containing a query, a positive sample and a negative sample; Filling the simple triple into a preset prompt word template, wherein the prompt word template is used to standardize the input format of the simple triple; Filling the relationship between the query and the positive sample and the negative sample into a label template, wherein the label template is used to represent the correlation between the query and the sample; The filled prompt word template is used as input data, and the filled label template is used as the label of the input data to form the relevance judgment data set.
3. The method according to claim 2, characterized in that Obtaining relevance selection data sets includes: The correlation data corpus is adjusted to generate a complex triple comprising a query, a plurality of positive samples, and a plurality of negative samples by random sampling; Filling the complex triple into a preset prompt word template, and filling the relationship between the query and multiple positive and negative samples into a label template; The filled prompt word template is used as input data, and the filled label template is used as the label of the input data to form the correlation selection data set.
4. The method according to claim 3, characterized in that: Adjusting the correlation data corpus includes: Adjusting the query in the correlation data corpus to a simple query, a complex query or a fuzzy query; or, Adjusting the query in the relevant data corpus to a long-tail query or a multimodal query; or, Context information is added to queries and samples in the correlation data corpus, wherein the context information is used to enable the large language model to understand complex semantic relationships.
5. The method according to claim 1, characterized in that Using the relevance judgment dataset to perform preliminary fine-tuning on the pre-trained large language model includes: Acquire multiple pre-trained large language models to be selected, wherein the parameter amount of the large language models to be selected is less than a set threshold; Using the correlation judgment data set to perform preliminary fine-tuning on the plurality of large language models to be selected, wherein half-precision training is used in the preliminary fine-tuning process; The candidate large language model with the best fine-tuning effect is selected as the large language model that has completed preliminary fine-tuning.
6. The method according to claim 1, characterized in that After deploying the fine-tuned large language model to the information retrieval system, the method further includes: Input the target query content into the fine-tuned large language model; Mixed precision is used for reasoning to obtain the information query result output by the large language model.
7. A two-stage fine-tuning training device based on a large language model, characterized in that: The device comprises: An acquisition module, used to acquire a relevance judgment data set and a relevance selection data set, wherein the relevance judgment data set and the relevance selection data set both contain triplets consisting of a query, a positive sample, and a negative sample, and the complexity of the relevance selection data set is higher than the complexity of the relevance judgment data set; A first fine-tuning module is used to perform preliminary fine-tuning on the pre-trained large language model using the relevance judgment data set, wherein the preliminary fine-tuning is used to enable the large language model to master basic relevance judgment capabilities; A second fine-tuning module is used to use the relevance selection data set to fine-tune the large language model after preliminary fine-tuning, wherein the fine-tuning is used to enable the large language model to master the relevance discrimination ability in a complex environment; The deployment module is used to deploy the fine-tuned large language model to the information retrieval system.
8. The device according to claim 7, characterized in that The acquisition module is used for: Obtaining a corpus of relevant data, wherein each query in the corpus of relevant data corresponds to at least zero positive samples and at least zero negative samples; Randomly sampling from the relevant data corpus to generate a simple triple containing a query, a positive sample and a negative sample; Filling the simple triple into a preset prompt word template, wherein the prompt word template is used to standardize the input format of the simple triple; Filling the relationship between the query and the positive sample and the negative sample into a label template, wherein the label template is used to represent the correlation between the query and the sample; The filled prompt word template is used as input data, and the filled label template is used as the label of the input data to form the relevance judgment data set.
9. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, for implementing the method steps described in any one of claims 1 to 6 when executing a program stored in a memory.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps described in any one of claims 1 to 6 are implemented.