Correlation determination method and related equipment
By supervising and fine-tuning the large language model and learning the correlation conclusions and reasons, the information retrieval difficulties caused by inconsistent meaning of ‘code words’ in vertical categories are solved, and more accurate information retrieval and better generalization capabilities are achieved.
Patent Information
- Application Number
- CN202510168782.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-06
AI Technical Summary
In specific vertical fields, existing information retrieval techniques are difficult to accurately search for information needed by users because the meaning of ‘code words’ in these fields is inconsistent with general understanding.
By supervising and fine-tuning the large language model, it can not only learn the correlation conclusions, but also fully learn the correlation reasons corresponding to the correlation conclusions, thereby achieving better learning effects and generalization ability.
It realizes more accurate information retrieval in specific vertical fields, can effectively learn and understand the meaning of "code words" in different vertical categories, and improves the relevance of search results.
Smart Images

Figure CN120104768A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a correlation determination method and related devices. Background Art
[0002] With the development of the Internet, a large number of vertical fields have emerged, also referred to as vertical categories. Simply put, each of the above vertical categories represents a specific field or classification consisting of producers of a type of content, that type of content, and consumers of that type of content. For example, beauty, fashion, technology, sports, novels, animation, etc. can all be regarded as different vertical categories.
[0003] In particular, there are a large number of "code words" in each vertical neighborhood. In the same vertical field, users usually use these "code words" to communicate. However, since the meanings expressed by these "code words" are usually inconsistent with the general understanding, it is usually difficult for current information retrieval technology to accurately search for the information users need when searching for information in a specific vertical field. Summary of the invention
[0004] In view of this, an embodiment of the present disclosure provides a correlation determination method, which performs supervised fine-tuning on a large language model so that the large language model can not only learn the correlation conclusions but also fully learn the correlation reasons corresponding to the correlation conclusions, thereby achieving better learning effects and ultimately achieving better generalization capabilities.
[0005] The correlation determination method described in the embodiment of the present disclosure may include: obtaining a query statement input by a user; obtaining candidate texts from a to-be-retrieved information library; generating a prompt based on the query statement and the candidate texts; inputting the prompt into a large language model that has been fine-tuned under supervised conditions; and obtaining a correlation conclusion and a correlation reason between the query statement and the candidate texts output by the large language model.
[0006] The relevance determination method described in the embodiment of the present disclosure may further include: constructing a training data set; wherein each training sample in the training data set includes: a query statement, a candidate text, a relevance label, and a relevance reason; generating a model fine-tuning prompt described in natural language for each training sample in the training data set; and performing supervised fine-tuning on the large language model based on the model fine-tuning prompt to obtain the supervised fine-tuned large language model.
[0007] In an embodiment of the present disclosure, the above method further includes: constructing a pre-training data set; wherein the pre-training data set includes: multiple contents in a pre-set vertical field; generating pre-training prompts of natural language description for each content in the pre-training data set; and before performing supervised fine-tuning on the large language model, continuing to pre-train the large language model based on the pre-training prompts.
[0008] In an embodiment of the present disclosure, continuing pre-training the large language model based on the pre-training prompt includes: continuing pre-training the large language model based on the pre-training prompt using a next word prediction technology.
[0009] In an embodiment of the present disclosure, constructing a training data set includes: generating self-enhanced training samples; extracting a corresponding number of self-enhanced training samples and manually annotated training samples according to a preset mixing ratio; and mixing the extracted self-enhanced training samples and manually annotated training samples to obtain the training data set.
[0010] In an embodiment of the present disclosure, generating a self-enhancement training sample includes: obtaining a pre-labeled data sample; wherein the pre-labeled data sample includes: a query statement, a candidate text, and a first relevance label; generating a task prompt described in natural language for the query statement and the candidate text; inputting the task prompt into a large language model that has undergone supervised fine-tuning to obtain an answer to the task prompt; wherein the answer includes a second relevance label and a relevance reason corresponding to the query statement and the candidate text; and in response to determining that the first relevance label is consistent with the second relevance label, generating a self-enhancement training sample based on the query statement, the candidate text, the first relevance label, and the relevance reason.
[0011] In an embodiment of the present disclosure, generating a model fine-tuning prompt with a natural language description for each training sample in the training data set includes: generating a task description in the model fine-tuning prompt based on a preset task description template; generating a search term in the model fine-tuning prompt based on a query statement in the training sample; generating a text in the model fine-tuning prompt based on a candidate text in the training sample; and generating an answer in the model fine-tuning prompt based on a relevance label and a relevance reason in the training sample.
[0012] In an embodiment of the present disclosure, fine-tuning the large language model in a supervised manner based on the model fine-tuning prompt includes: fine-tuning the large language model in a supervised manner based on the model fine-tuning prompt using a next word prediction technique.
[0013] Embodiments of the present disclosure may further include: testing a large language model that has undergone supervised fine-tuning using a test data set; obtaining error cases in which the relevance conclusion output by the large language model is inconsistent with the corresponding relevance label in the test set; analyzing the relevance reasons corresponding to the error cases to determine the error cause of the large language model; collecting supplementary information related to the error cause based on the error cause; generating targeted training samples based on the supplementary information; and performing supplementary training on the large language model based on the targeted training samples.
[0014] Corresponding to the above correlation determination method, an embodiment of the present disclosure further discloses a correlation determination device, including:
[0015] A query statement acquisition module is used to acquire the query statement input by the user;
[0016] A candidate text acquisition module is used to acquire candidate texts from the information database to be searched;
[0017] A query prompt generating module, generating prompts based on the query statement and the candidate text;
[0018] The relevance determination module inputs the prompt into a large language model that has been fine-tuned in a supervised manner, and obtains a relevance conclusion and a relevance reason between the query statement and the candidate text output by the large language model.
[0019] In addition, an embodiment of the present disclosure further provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned correlation determination method when executing the program.
[0020] An embodiment of the present disclosure further provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the above correlation determination method.
[0021] An embodiment of the present disclosure further provides a computer program product, including computer program instructions. When the computer program instructions are executed on a computer, the computer is enabled to execute the above correlation determination method.
[0022] It can be seen from this that the above-mentioned relevance determination method and related devices enable the large language model to learn more difficult semantic understanding problems by simultaneously constraining the conclusion (Label) and the relevance reason (CoT) during the model training process, especially to effectively learn the meaning of "code words" in different vertical categories. Compared with the traditional solution that only optimizes the conclusion, the large language model can not only learn the relevance conclusion but also fully learn the relevance reasons corresponding to the relevance conclusion, so as to achieve better learning results and ultimately achieve better generalization capabilities. The above method solves the problem of using a large language model to learn and model proprietary corpus, and innovatively adds training on relevance reasons, so that the large language model can learn the meaning of the "code words" in a certain vertical category and achieve the effect of learning from one example, so that the relevance between the user's search terms and the retrieved documents can be accurately measured. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the present disclosure or related technologies, the drawings required for use in the embodiments or related technical descriptions are briefly introduced below. Obviously, the drawings described below are only embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0024] Figure 1 The implementation process of the method for supervised fine-tuning of a large language model described in some embodiments of the present disclosure is shown.
[0025] Figure 2 The implementation process of the method for continuing pre-training a large language model described in some embodiments of the present disclosure is shown.
[0026] Figure 3 The implementation process of the method for automatically generating training samples through self-enhancement described in some embodiments of the present disclosure is shown.
[0027] Figure 4 The implementation process of the targeted completion described in some embodiments of the present disclosure is shown.
[0028] Figure 5 The implementation process of the correlation determination method described in some embodiments of the present disclosure is shown.
[0029] Figure 6 The internal structure of the correlation determination device described in some embodiments of the present disclosure is shown.
[0030] Figure 7 A more specific schematic diagram of the hardware structure of an electronic device described in some embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0031] In order to make the objectives, technical solutions and advantages of the present disclosure more clearly understood, the present disclosure is further described in detail below in combination with specific embodiments and with reference to the accompanying drawings.
[0032] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should be understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in the embodiments of the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. "Including" or "comprising" and similar words mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0033] It is understandable that before using the technical solutions of each embodiment of the present disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner, and the user's authorization will be obtained.
[0034] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly remind the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can independently choose whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.
[0035] As an optional but non-limiting implementation, in response to receiving an active request from the user, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0036] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0037] As mentioned above, different vertical neighborhoods have different "code words", and the meanings of these "code words" are usually inconsistent with general knowledge. This makes it difficult for current information retrieval technology to accurately search for the information users need when performing information retrieval in a specific vertical field.
[0038] For example, traditional natural language processing methods commonly used currently will use an encoding (Encoder) model similar to BERT, inputting the query statement (Query) entered by the user and the text to be matched (Doc) into the BERT Encoder model simultaneously to obtain a relevance score. The distribution of this relevance score is usually between 0 and 1. Among them, 0 represents that the two are not relevant, and 1 represents that the semantics are exactly the same. Currently, during the training process of the above BERT Encoder model, manually annotated data, that is, a data group of <Query, Doc, Label> is usually used, where Label represents the annotation of the relevance between the above <Query, Doc>. By learning the Label in the above data group, the above BERT Encoder model can achieve the prediction of the relevance of the input <Query, Doc> data pair. However, in the above traditional solution, the BERT Encoder model can only learn the Label itself, but cannot constrain it to learn the true meaning. For some Queries, the BERT Encoder model can only forcibly remember the relevance conclusion, or attribute the conclusion to other reasons, resulting in poor learning effects. For example, for the data group <Query = be novel, Doc = Which is the most tragic book on Tomato, Label = 0.7> in the novel vertical category, the trained BERT Encoder model will forcibly remember the conclusion that the relevance is 0.7, but it does not learn the connection between "be" representing a bad ending (Bad Ending) in the novel field and "knife" (which refers to a heart-breaking tragic novel in the novel field). Instead, it may attribute the relevance of 0.7 to the relationship between "novel" and "book", resulting in the inability to generalize when facing other data.
[0039] To this end, the embodiments of the present disclosure provide a method for supervised fine-tuning of a large language model, which can optimize a large language model to obtain a relevance determination model for a specific vertical field, so as to achieve precise information retrieval in the specific vertical field.
[0040] Figure 1 Shows the implementation process of the training method for supervised fine-tuning of a large language model according to some embodiments of the present disclosure. As Figure 1 shown, the method for supervised fine-tuning of a large language model according to the embodiments of the present disclosure includes the following steps:
[0041] In step 110, a training data set is constructed.
[0042] In the embodiments of the present disclosure, the training data set may include multiple training samples. Each training sample may include: a query statement, a candidate text, a relevance label, and a relevance reason. That is, each training sample may be represented as<Query,Doc,Label,CoT> A data group in the form of . Query represents the query statement. Doc represents the candidate text. Label represents the relevance label. In some applications, Label can be represented by a score that is usually distributed between 0 and 1, where 0 represents that the two are irrelevant and 1 represents that the semantics are completely consistent. In other applications, Label can also be a textual expression representing relevance, such as irrelevant, weakly relevant, moderately relevant, and strongly relevant. CoT represents the reason for the relevance, that is, the explanation of the Label.
[0043] As an example, the following Table 1 shows an example of a training sample.
[0044]
[0045] Table 1
[0046] In an embodiment of the present disclosure, the training data set can be constructed in the following manner: first, self-enhancement training samples are generated; then, a corresponding number of self-enhancement training samples and manually labeled training samples are extracted according to a preset mixing ratio; finally, the extracted self-enhancement training samples and manually labeled training samples are mixed to obtain the training data set. The method for generating self-enhancement training samples will be described in detail later, which is omitted here for the time being.
[0047] In step 120, a model fine-tuning prompt described in natural language is generated for each training sample in the training data set.
[0048] It can be understood that, usually, the large language model cannot directly understand the above data set, and it is necessary to convert it into a prompt described in natural language. In the embodiment of the present disclosure, the above model fine-tuning prompt usually includes four parts: task description, search terms, text, and answer. Among them, the task description can use a pre-set task description template; the search term can be the Query in the training sample; the text can be the Doc in the training sample; the answer can be composed of the Label and CoT in the training sample. It can be seen that in the above step 120, the task description in the model fine-tuning prompt can be first generated based on the pre-set task description template; then, the search term in the model fine-tuning prompt can be generated based on the Query in the training sample; the text in the model fine-tuning prompt can be generated based on the Doc in the training sample; finally, the answer in the model fine-tuning prompt can be generated based on the Label and CoT in the training sample.
[0049] As an example, Table 2 below shows an example of a prompt.
[0050]
[0051]
[0052] Table 2
[0053] In step 130, based on the above-mentioned model fine-tuned prompt, the large language model is subjected to supervised fine-tuning to obtain a large language model that has undergone supervised fine-tuning.
[0054] In the embodiments of the present disclosure, the above-mentioned supervised fine-tuning of the large language model can be completed by using the training method of next token prediction.
[0055] Specifically, in the embodiments of the present disclosure, the above-mentioned training method of next token prediction means that starting from the second token in the answer part of the above-mentioned fine-tuned prompt, each token is used as the target word in turn, and the large language model is made to estimate the probability of each word in a pre-set vocabulary as the next word according to the previous text. Among them, the training objective is to make the probability of the target word greater than the probability of other words. For example, assuming that the third word "be" in the above answer is the target word, the training objective is to require the large language model to estimate p(be|in, relevant)=1. That is to say, for the above-mentioned training method of next token prediction, the total loss Loss of the model can be expressed as the following expression: Among them, V represents the above-mentioned pre-set vocabulary, that is, the set of all possible tokens in the text; i represents the i-th token in the above vocabulary; p i is the probability estimated by the large language model according to the previous text that the i-th word appears as the next word; y i is the label value corresponding to the i-th word. Specifically, if the i-th word is the target word, then y i =1, otherwise, y i =0. For example, for the above example, if the current target word is "be", then only the y i corresponding to "be" in the vocabulary is 1, and the y i corresponding to other words is 0. It can be seen that the training objective of next token prediction is to minimize the above-mentioned total loss Loss.
[0056] It can be seen that the supervised fine-tuning method for the large language model described in the embodiment of the present disclosure uses the method of simultaneously constraining the conclusion (Label) and the relevance reason (CoT), so that the large language model can learn more difficult semantic understanding problems, especially can effectively learn the meaning of "code words" in different vertical fields. Compared with the traditional solution of only optimizing the conclusion, the large language model can not only learn the relevance conclusion but also fully learn the relevance reason corresponding to the relevance conclusion, so as to achieve better learning effect and finally achieve better generalization ability.
[0057] In order to make the large language model better adapt to downstream tasks in a specific vertical field, before the above-mentioned large language model is fine-tuned in a supervised manner, the embodiment of the present disclosure also discloses a method for further pre-training the large language model. Figure 2 The implementation process of the method for continuing pre-training a large language model described in the embodiment of the present disclosure is shown. Figure 2 As shown, the above-mentioned continued pre-training may include the following multiple steps.
[0058] In step 210, a pre-training dataset is constructed.
[0059] In an embodiment of the present disclosure, the above-mentioned pre-training data set may include: multiple contents in a pre-set vertical field.
[0060] In the embodiments of the present disclosure, since the pre-training data set does not need to be labeled, it can be collected manually or automatically by a machine. For example, the collection of pre-training data sets for different vertical fields can be completed from encyclopedic knowledge, professional forums in various fields, forums, or various sections of news. For example, for the field of novels, novel introductions, novel texts, articles in novel forums, encyclopedic knowledge related to novels, and news related to novels, etc. can be collected as pre-training data sets for the field of novels. The embodiments of the present disclosure do not limit the generation method of the above-mentioned pre-training data set.
[0061] In step 220, a pre-trained prompt described in natural language is generated for each content in the pre-trained data set.
[0062] As mentioned above, the large language model usually cannot directly understand the collected content, so it needs to be converted into prompts described in natural language. In the embodiment of the present disclosure, the above step 220 can usually be constructed into a form of dialogue or question and answer based on each content, which is easy for the large language model to understand. For example, suppose that in the field of novels, one of the collected content is the following encyclopedia entry: "Portable space text: refers to the protagonist or some people in the book have their own space, which can be carried with them, and the area can be upgraded to become larger and more spacious. Generally speaking, the space time of space text is static, and the stored things will not deteriorate or break down. Sometimes there will be fields, houses, springs, small lakes, etc. in it." In this way, in the above step 220, it can be constructed into the following question and answer form: "In the novel, what does portable space text mean? It means that the protagonist or some people in the book have their own space, which can be carried with them. The area can be upgraded to become larger and more spacious. Generally speaking, the space time of space text is static, and the stored things will not deteriorate or break down. Sometimes there will be fields, houses, springs, small lakes, etc. in it."
[0063] Specifically, for content organized in different ways, rewriting templates and rewriting rules can be pre-set. In this way, the rewriting of the collected content into the pre-trained prompts can be automatically completed based on the set rewriting templates and rewriting rules.
[0064] In step 230, the large language model is further pre-trained based on the pre-training prompt.
[0065] In the embodiments of the present disclosure, the above-mentioned continued pre-training of the large language model can also be performed by self-supervised training using the Next Token Prediction training method. The specific training process can refer to the aforementioned supervised fine-tuning process of the large language model, and the specific training process will not be repeated here.
[0066] It can be seen that in the above-mentioned continued pre-training process, by collecting a large amount of corpus in a specific vertical field and using the Next Token Prediction technology to perform self-supervisory training on the large language model, the large language model can be first migrated from the knowledge of the entire field to a specific vertical field. For example, by injecting a large amount of collected novel-related knowledge into the large language model, the traditional large language model can be migrated to the novel field, so that the large language model itself after continued pre-training is consistent with the distribution of the corpus for subsequent supervised fine-tuning, thereby reducing training losses and improving training efficiency.
[0067] For the training data set constructed in the above step 110, in some embodiments, data collection and data labeling can be completed completely manually. However, the method of manually generating training data sets is not only costly, but also the amount of data in a short period of time is relatively limited. In order to better adjust the effect of the large language model, it is usually necessary to continuously polish the quality of the training data. In an embodiment of the present disclosure, the above training data set can be polished by self-enhancement (SelfImprovement) or targeted supplementation (Targetedsupplement). In some applications, the above self-enhancement and targeted supplementation of data can also be repeated alternately.
[0068] For the above self-enhancement method, as mentioned above, the training data set will include two parts of data: manually annotated training samples and training samples automatically generated by self-enhancement. The above two parts of data are mixed together in a pre-set mixing ratio to form the above training data set. In this case, the above method of automatically generating training samples by self-enhancement can refer to Figure 3 .like Figure 3 As shown, generating a self-enhancement training sample may include the following steps:
[0069] In step 310, a pre-labeled data sample is obtained.
[0070] In the embodiment of the present disclosure, the above-mentioned pre-labeled data sample may include: a query sentence, a candidate text, and a first relevance label. That is, each pre-labeled data sample can be represented as<Query,Doc,Label> data set. Query represents the query statement; Doc represents the candidate text; and Label represents the first relevance label. Since the above pre-annotated data samples do not include the aforementioned relevance reason part, the above pre-annotated data samples can be created using an existing training data set for text relevance model training or by manual annotation. For manual annotation, since only Label needs to be annotated, the labor cost is relatively low.
[0071] In step 320, a task prompt described in natural language is generated for the query statement and the candidate text.
[0072] In an embodiment of the present disclosure, in the above step 320, the task description in the task prompt may be first generated based on a preset task description template; then, the search term in the task prompt may be generated based on the Query in the training sample; and the text in the task prompt may be generated based on the Doc in the training sample. It should be noted that the preset task description template used in the above step 320 may be the same as the task description template used in the above step 120, and thus, the generated task prompt may be similar to the prompt shown in Table 2, and the only difference may be that the task prompt generated in this step will not include the answer part shown in Table 2.
[0073] In step 330, the task prompt is input into a large language model that has been fine-tuned in a supervised manner to obtain a response to the task prompt.
[0074] In an embodiment of the present disclosure, the answer to the task prompt output by the large language model that has undergone supervised fine-tuning will include a second relevance label and a relevance reason corresponding to the query statement and the text.
[0075] In step 340, in response to determining that the first relevance label is consistent with the second relevance label, a self-enhancement training sample is generated based on the query sentence, the candidate text, the first relevance label and the relevance reason.
[0076] In an embodiment of the present disclosure, after obtaining the answer to the task prompt output by the supervised fine-tuned large language model, it is possible to preliminarily determine whether the relevance reason output by the supervised fine-tuned large language model is accurate based on the consistency between the first relevance label and the second relevance label. If the first relevance label is consistent with the second relevance label, it can be preliminarily determined that the relevance reason output by the supervised fine-tuned large language model is basically accurate and can be used as a training sample for training other large language models.
[0077] In this way, after obtaining the above-mentioned multiple self-enhanced training samples, they can be mixed with other manually annotated training samples according to a preset mixing ratio to generate the above-mentioned training data set, which can be further used for supervised fine-tuning of other large language models. It should be noted that the above-mentioned mixing ratio can be set according to demand or experience.
[0078] Since the training samples generated by the self-enhancement method are not real knowledge, when using the training data set containing the training samples generated by the self-enhancement method for training, the large language model may have overfitting problems. Therefore, in this case, it is necessary to pay attention to the loss trend of the large language model training. When the loss of the large language model is significantly reduced compared to the previous version, it is necessary to readjust the above mixing ratio, that is, to reduce the ratio of the training samples generated by the self-enhancement method mixed into the training data set, so as to ensure the training effect of the large language model.
[0079] It can be seen from this that a large number of training samples required by the embodiments of the present disclosure can be obtained through self-enhancement based on a large language model that has been fine-tuned in a supervised manner and pre-labeled data samples that only contain relevant labels, thereby reducing the cost of generating training data sets and improving training efficiency.
[0080] The above-mentioned targeted data filling method usually needs to be completed after testing the large language model based on the test data set. Specifically, Figure 4 The implementation process of the targeted filling method described in the embodiment of the present disclosure is shown. Figure 4 As shown, the above-mentioned targeted filling method may include:
[0081] At step 410 , the large language model that has undergone supervised fine-tuning is tested using a test dataset.
[0082] In step 420, an error case in which the relevance conclusion output by the large language model is inconsistent with the corresponding relevance label in the test data set is obtained.
[0083] In step 430, the correlation reasons corresponding to the error cases are analyzed to determine the error causes of the large language model.
[0084] In step 440, supplementary information related to the error cause is collected based on the error cause.
[0085] In step 450, targeted training samples are generated based on the above supplementary information.
[0086] In step 460, supplementary training is performed on the large language model based on the targeted training samples.
[0087] Furthermore, in some embodiments, the generated targeted training samples may be added to the above-mentioned pre-training data set as training data for other large language models.
[0088] For example, in a specific example, the query input into the supervised fine-tuned large language model is "Pokemon cheats", and the doc is "The three great Pokémons, level one gods, level two gods, with systems, and more battles." After the above data set is input into the supervised fine-tuned large language model, the output is "Not relevant, the search term requires content related to Pokémon, and the post mentions the three great Pokémons, level one gods, etc., which are different from Pokémon, so it is not relevant." By analyzing the reasons for the relevance of the above error cases, it can be found that the above supervised fine-tuned large language model is not clear about the relevant concepts of Pokémon intellectual property (IP), especially the concepts related to the three great Pokémons, so it can be supplemented with relevant concepts in a targeted manner. After obtaining supplementary information related to Pokémon and the three great Pokémons, targeted training samples can be generated. The following is an example of a targeted training sample: "What are the three starter Pokémon? The three starter Pokémon is a term in Pokémon. In the games or animations in this series, the protagonist will initially choose one of three Pokémon to start his adventure. The three starter Pokémon usually have three attributes: water / fire / grass, and have good growth potential, and are loved by everyone. The three starter Pokémon of the first generation of Pokémon are Charmander, Squirtle, and Bulbasaur. In the ninth generation of games, the three starter Pokémon are Charmeleon, Creaturedrake, and Creatureleaf." After using the above targeted training samples to supplement the large language model, the large language model can have a clearer understanding of the relevant concepts of the Pokémon IP, so that more accurate relevance judgment can be achieved for the relevant texts of Pokémon.
[0089] It can be seen that through the above-mentioned targeted data completion method, the missing knowledge or concepts in the large language model training process can be quickly discovered based on the relevance reasons of the model output during the test process, and corresponding supplements can be made, thereby greatly improving the efficiency and effect of model training. At the same time, it can be seen from the above process that adding relevance reasons to the training samples and requiring the large language model to not only give relevance conclusions but also relevance reasons during the relevance judgment process can quickly lock in the erroneous intentions of the large language model, thereby greatly improving the ability to migrate the large language model to a specific vertical field, and also greatly improving the training efficiency and effect of the model.
[0090] Furthermore, in order to measure the effect of model correction and prevent the newly added targeted training samples from affecting the effect of the original large language model on other test cases, the embodiment of the present disclosure also proposes a method for measuring the model. Specifically: Assume that the model retrained after correction is theta, and the original model is ref. Compared with the original model ref, the model theta is required to maximize the benefits on the targeted training samples and the relevant test sets (Maximizes therewards), and at the same time, on other test data, the difference between the model theta and the model ref is as small as possible (Prevents the model from changing too drastically). If the targeted training samples cannot meet the above requirements, the above targeted training samples need to be re-produced.
[0091] Corresponding to the above correlation determination method, some embodiments of the present disclosure further disclose a correlation determination method. Figure 5 The implementation process of the correlation determination method described in the embodiment of the present disclosure is shown. Figure 5 As shown, the above method may include the following steps:
[0092] In step 510, a query statement input by a user is obtained;
[0093] In step 520, candidate texts are obtained from the information database to be searched;
[0094] At step 530, a prompt is generated based on the query statement and the candidate text;
[0095] At step 540, the prompt is input into the large language model that has undergone supervised fine-tuning; and
[0096] In step 550, the correlation conclusion and correlation reason between the query sentence output by the large language model and the candidate text are obtained.
[0097] Figure 6 The internal structure of the correlation determination device described in some embodiments of the present disclosure is shown. Figure 6 The above-mentioned correlation determination device may include:
[0098] A query statement acquisition module 610 is used to acquire a query statement input by a user;
[0099] A candidate text acquisition module 620 is used to acquire candidate text from the information library to be searched;
[0100] A query prompt generating module 630, which generates prompts based on the query statement and the candidate text;
[0101] The relevance determination module 640 inputs the prompt into the large language model that has undergone supervised fine-tuning, and obtains the relevance conclusion and relevance reason between the query statement and the candidate text output by the large language model.
[0102] It should be noted that the implementation method of each module in the above device can refer to the implementation method of each step in the above embodiment, and will not be repeated here.
[0103] Similar to the above-mentioned relevance determination method, the above-mentioned relevance determination device enables the large language model to learn more difficult semantic understanding problems by simultaneously constraining the conclusion (Label) and the relevance reason (CoT) during the supervised fine-tuning of the large language model, especially to effectively learn the meaning of "code words" in different vertical categories. Compared with the traditional solution of only optimizing the conclusion, the large language model can not only learn the relevance conclusion but also fully learn the relevance reasons corresponding to the relevance conclusion, so as to achieve better learning effect and ultimately achieve better generalization ability. The above method solves the learning and modeling method of proprietary corpus using a large language model, and innovatively adds the training of relevance reasons, so that the large language model can learn the meaning of "code words" in a certain vertical category and achieve the effect of learning from one example, so as to accurately measure the relevance between the user's search terms and the retrieved documents.
[0104] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present disclosure also provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the correlation determination method described in any of the above embodiments is implemented.
[0105] Figure 7 A more specific hardware structure diagram of an electronic device provided in this embodiment is shown, and the device may include: a processor 2010, a memory 2020, an input / output interface 2030, a communication interface 2040, and a bus 2050. The processor 2010, the memory 2020, the input / output interface 2030, and the communication interface 2040 are connected to each other in communication within the device through the bus 2050.
[0106] The processor 2010 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0107] The memory 2020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 2020 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program codes are stored in the memory 2020 and are called and executed by the processor 2010.
[0108] The input / output interface 2030 is used to connect input / output devices to realize information input and output. The input / output devices can be configured in the device as components, or can be externally connected to the device to provide corresponding functions. The input devices can include microphones, various sensors, etc., and the output devices can include displays, speakers, vibrators, indicator lights, etc.
[0109] The communication interface 2040 is used to connect a communication module (not shown) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired mode (such as USB, network cable, etc.) or a wireless mode (such as mobile network, WIFI, Bluetooth, etc.).
[0110] The bus 2050 includes a path that transmits information between the various components of the device (eg, the processor 2010, the memory 2020, the input / output interface 2030, and the communication interface 2040).
[0111] It should be noted that, although the above device only shows the processor 2010, the memory 2020, the input / output interface 2030, the communication interface 2040, and the bus 2050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include the components necessary for implementing the embodiments of the present specification, and does not necessarily include all the components shown in the figure.
[0112] The electronic device of the above embodiment is used to implement the corresponding correlation determination method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.
[0113] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present disclosure also provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the correlation determination method described in any of the above embodiments.
[0114] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.
[0115] The computer instructions stored in the storage medium of the above embodiments are used to enable the computer to execute the task processing method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0116] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples. Based on the concept of the present disclosure, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present disclosure as described above, which are not provided in detail for the sake of simplicity.
[0117] In addition, to simplify the description and discussion, and in order not to make the embodiments of the present disclosure difficult to understand, the known power / ground connections to the integrated circuit (IC) chips and other components may or may not be shown in the provided figures. In addition, the device can be shown in the form of a block diagram to avoid making the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure will be implemented (that is, these details should be fully within the scope of understanding of those skilled in the art). Where specific details (e.g., circuits) are set forth to describe exemplary embodiments of the present disclosure, it is apparent to those skilled in the art that the embodiments of the present disclosure can be implemented without these specific details or with changes in these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0118] Although the present disclosure has been described in conjunction with specific embodiments of the present disclosure, many replacements, modifications and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.
[0119] The embodiments of the present disclosure are intended to cover all such substitutions, modifications and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present disclosure should be included in the scope of protection of the present disclosure.
Claims
1. A method for determining relevance, comprising: Get the query statement entered by the user; Obtain candidate texts from a to-be-retrieved information database; Generate a prompt based on the query statement and the candidate text; Inputting the prompt into a large language model that has been fine-tuned in a supervised manner; and Obtain a correlation conclusion and a correlation reason between the query statement output by the large language model and the candidate text.
2. The method according to claim 1, further comprising: Constructing a training data set; wherein each training sample in the training data set includes: a query statement, a candidate text, a relevance label, and a relevance reason; Generating a model fine-tuning prompt described in natural language for each training sample in the training data set; The supervised fine-tuning of the large language model is performed based on the model fine-tuning prompt to obtain the supervised fine-tuned large language model.
3. The method according to claim 2, further comprising: Constructing a pre-training data set; wherein the pre-training data set includes: multiple contents in a pre-set vertical field; Generating a pre-training prompt described in natural language for each content in the pre-training data set; and Before the large language model is fine-tuned in a supervised manner, the large language model is further pre-trained based on the pre-training prompts.
4. The method according to claim 3, wherein: Continuing pre-training the large language model based on the pre-training prompt includes: continuing pre-training the large language model based on the pre-training prompt using a next word prediction technology.
5. The method according to claim 1, wherein: Building a training dataset includes: Generate self-enhancement training samples; Extracting a corresponding number of self-enhancement training samples and manually labeled training samples according to a pre-set mixing ratio; The extracted self-enhancement training samples and the manually labeled training samples are mixed to obtain the training data set.
6. The method according to claim 5, wherein: Generating self-enhancement training samples includes: Acquire a pre-annotated data sample; wherein the pre-annotated data sample includes: a query statement, a candidate text, and a first relevance label; Generate a task prompt described in natural language for the query statement and the candidate text; Inputting the task prompt into a supervised fine-tuned large language model to obtain an answer to the task prompt; wherein the answer includes a second relevance label and a relevance reason corresponding to the query sentence and the candidate text; and In response to determining that the first relevance label is consistent with the second relevance label, a self-enhancement training sample is generated based on the query sentence, the candidate text, the first relevance label and the relevance reason.
7. The method according to claim 2, wherein: Generating a model fine-tuning prompt described in natural language for each training sample in the training data set includes: Generate task descriptions in model fine-tuning prompts based on pre-set task description templates; Generating a model based on the query sentences in the training samples to fine-tune the search terms in the prompts; Fine-tune the text in the prompt based on the candidate text generation model in the training sample; and Generate an answer in a model fine-tuning prompt based on the relevance labels and relevance reasons in the training samples.
8. The method according to claim 2, wherein: Performing supervised fine-tuning on the large language model based on the model fine-tuning prompt includes: performing supervised fine-tuning on the large language model based on the model fine-tuning prompt using a next word prediction technique.
9. The method according to claim 2, further comprising: Use the test dataset to test the large language model after supervised fine-tuning; Obtaining error cases in which the relevance conclusion output by the large language model is inconsistent with the corresponding relevance label in the test set; Analyze the correlation reasons corresponding to the error cases to determine the error causes of the large language model; Collecting supplementary information related to the cause of the error based on the cause of the error; Generating targeted training samples based on the supplementary information; as well as Supplementary training is performed on the large language model based on the targeted training samples.
10. A correlation determination device, comprising: A query statement acquisition module is used to acquire the query statement input by the user; A candidate text acquisition module is used to acquire candidate texts from the information database to be searched; A query prompt generating module, generating prompts based on the query statement and the candidate text; The relevance determination module inputs the prompt into a large language model that has been fine-tuned in a supervised manner, and obtains a relevance conclusion and a relevance reason between the query statement and the candidate text output by the large language model.
11. An electronic device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the correlation determination method according to any one of claims 1 to 9 is implemented.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the correlation determination method according to any one of claims 1 to 9.
13. A computer program product, comprising computer program instructions, which, when executed on a computer, enable the computer to execute the correlation determination method according to any one of claims 1 to 9.