Negative sample construction method and device, and model training method and device

By constructing complex negative samples and using them for training, the problem of low accuracy of recall results when RAG systems are applied in different fields is solved, and the matching between content and user needs and the generalization ability of the model is improved.

CN120011551APending Publication Date: 2025-05-16CHINA UNIONPAY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510115808.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

When the existing search enhanced generation (RAG) system is used in different application fields, the search results of recalled models are relatively accurate, resulting in poor matching between generated content and user needs.

Method used

By building complex negative samples, the difficulty of distinguishing during recall model training is increased, and the recall model is trained using these negative samples to improve the accuracy of the search results. The specific method includes obtaining a positive sample set, calculating the similarity between the target positive sample and the remaining positive samples, building a negative sample set with a similarity greater than or equal to the preset threshold, and using it for model training.

Benefits of technology

It improves the accuracy of the search results of recalled models, improves the matching degree between the content generated by the RAG system and user needs, and enhances the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011551A_ABST
    Figure CN120011551A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides negative sample construction and model training methods and devices. The method comprises the steps that a positive sample set is obtained, the positive sample set comprises a plurality of positive samples, and each positive sample comprises a query statement and an associated text of the query statement; any positive sample in the positive sample set is used as a target positive sample, and for each remaining positive sample, the similarity between at least one of a target query statement and a target associated text of the target positive sample and query statements and associated texts of the remaining positive samples is calculated; and taking at least one other positive sample with the similarity greater than or equal to a preset threshold value as a candidate sample, extracting an associated text in each candidate sample as a first text, and constructing a first negative sample comprising the first text and the target query statement to obtain a first negative sample set. The distinction degree between the first negative sample and the positive sample is small, and the model trained by using the negative samples has higher detection precision and better generalization ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method and device for negative sample construction and model training. Background Art

[0002] Retrieval-Augmented Generation (RAG), as a method that combines retrieval and generation techniques, can be applied to multiple application scenarios of natural language processing. Retrieval-Augmented Generation mainly searches based on the input, and then generates content corresponding to the input based on the search results. Therefore, the degree of match between the content generated by the RAG method and user needs depends largely on the search results recalled by the recall model in the RAG method. The above recall model is used to vectorize the input, and search the information stored in the vectorized form for the vectorized input to obtain the search results.

[0003] The needs of users are different in different application fields. In order to ensure that the content generated by the RAG system is well matched with the needs of users when used in different application fields, the recall model needs to have a high accuracy when recalling in various fields. Therefore, for each application field, when applying the RAG system to that field, the recall model needs to be fine-tuned.

[0004] The recall results retrieved by the recall model obtained by the current fine-tuning method have a low accuracy rate, which leads to a poor match between the content generated by the RAG system and user needs when it is applied in different fields. Summary of the invention

[0005] The embodiments of the present application provide a negative sample construction, model training method and device. The negative samples constructed by this method increase the difficulty of distinguishing from positive samples. When the recall model is trained using the above negative samples, the accuracy of the retrieval results recalled by the recall model can be improved, which in turn helps to improve the matching degree between the content generated by the RAG system and user needs.

[0006] In a first aspect, an embodiment of the present application provides a negative sample construction method, comprising: obtaining a positive sample set, wherein the positive sample set includes multiple positive samples, each positive sample including a query statement and associated text of the query statement; taking any positive sample in the positive sample set as a target positive sample, and for each of the remaining positive samples, calculating the similarity between at least one of the target query statement and the target associated text of the target positive sample and the query statement and the associated text of the remaining positive samples; taking at least one of the remaining positive samples whose similarity is greater than or equal to a preset threshold as a candidate sample, extracting the associated text in each candidate sample as a first text, constructing a first negative sample including the first text and the target query statement, and obtaining a first negative sample set.

[0007] In a possible implementation, obtaining the positive sample set includes: processing a plurality of documents using a natural language processing model, and constructing the positive sample set according to the processing results.

[0008] In a possible implementation, the method of processing the multiple documents using a natural language processing model and constructing the positive sample set based on the processing results includes: processing the multiple documents using a natural language processing model to obtain multiple triples, wherein each triple includes a query statement, a reply content, and associated text; wherein the natural language model generates the reply content based on the associated text, and the reply content is used to reply to the query statement; for each triple, constructing a positive sample from the query statement and the associated text in the triple; and the positive sample set is composed of positive samples corresponding to multiple triples.

[0009] In a possible implementation, the multiple documents are documents in a preset knowledge base.

[0010] In a possible implementation, for each remaining positive sample, calculating the similarity between at least one of the target query statement and the target associated text of the target positive sample and the query statement and the associated text of the remaining positive sample, respectively, includes: calculating a first similarity between the target query statement and the query statement of the remaining positive sample, a second similarity between the target query statement and the associated text of the remaining positive sample, a third similarity between the target associated text and the query statement of the remaining positive sample, and a fourth similarity between the target associated text and the associated text of the remaining positive sample; and taking at least one remaining positive sample whose similarity is greater than or equal to a preset threshold as a candidate sample, includes: in response to at least one of the first similarity, the second similarity, the third similarity and the fourth similarity being greater than or equal to the preset threshold, taking the remaining positive sample as the candidate sample.

[0011] In a possible implementation, different training stages of the model correspond to different preset thresholds; and at least one remaining positive sample whose similarity is greater than or equal to the preset threshold is used as a candidate sample, and the associated text in each candidate sample is extracted as the first text, and a first negative sample including the first text and the target query statement is constructed to obtain a first negative sample set, including: for each training stage, at least one remaining positive sample whose similarity is greater than or equal to the preset threshold of the training stage is used as a candidate sample of the stage; the associated text of each candidate sample is extracted as the first text, and a first negative sample is constructed from the first text and the target query statement to obtain the first negative sample set of the training stage.

[0012] In a possible implementation, different training stages of the model are determined based on one of the following methods: training rounds, a value of a loss function, and an accuracy rate on a validation set.

[0013] In a possible implementation, the similarity is cosine similarity.

[0014] In a possible implementation, the method further includes: determining at least one second text from the document where the target associated text is located, constructing a second negative sample including the second text and the query statement for each second text, to obtain a second negative sample set; and integrating the first negative sample set and the second negative sample set to obtain a target negative sample set.

[0015] In a possible implementation, determining at least one second text from the document where the target associated text is located includes: determining a context of the target associated text in the document where the target associated text is located; and using the context as the second text.

[0016] In a second aspect, an embodiment of the present application provides a model training method, comprising: obtaining a positive sample set and a negative sample set; wherein the negative sample set is obtained by the method described in the first aspect and various possible implementation methods of the first aspect; constructing multiple training samples based on the positive samples in the positive sample set and the negative samples in the negative sample set; wherein each training sample includes the associated text of the query statement and one or more of the first text of the query statement, as well as the query statement; using multiple of the training samples to fine-tune the recall model to obtain a trained recall model.

[0017] In a possible implementation, constructing multiple training samples based on the positive samples in the positive sample set and the negative samples in the negative sample set includes: determining the first negative sample sets corresponding to different training stages from the negative sample set; wherein different training stages are used to determine different similarity thresholds for negative samples; for each training stage, constructing multiple training samples for that stage based on the positive samples in the positive sample set and the negative samples in the first negative sample set of that stage; and using the multiple training samples to fine-tune the recall model to obtain the trained recall model includes: in each training stage, using the multiple training samples of that stage to fine-tune the recall model to obtain the trained recall model of that stage.

[0018] In a third aspect, an embodiment of the present application provides a negative sample construction device, comprising: a first acquisition unit, used to acquire a positive sample set, wherein the positive sample set includes multiple positive samples, each positive sample includes a query statement and an associated text of the query statement; a similarity determination unit, used to take any positive sample in the positive sample set as a target positive sample, and for each remaining positive sample, calculate the similarity between at least one of a target query statement and a target associated text of the target positive sample and the query statement and the associated text of the remaining positive samples;

[0019] The extraction unit is used to take at least one remaining positive sample whose similarity is greater than or equal to a preset threshold as a candidate sample, extract the associated text in each candidate sample as the first text, construct a first negative sample including the first text and the target query statement, and obtain a first negative sample set.

[0020] In a possible implementation manner, the first acquiring unit is further configured to:

[0021] The multiple documents are processed using a natural language processing model, and the positive sample set is constructed according to the processing results.

[0022] In a possible implementation, the first acquisition unit is further used to: process the multiple documents using a natural language processing model to obtain multiple triples, wherein each triple includes a query statement, a reply content, and associated text; wherein the natural language model generates the reply content based on the associated text, and the reply content is used to reply to the query statement; for each triple, construct a positive sample from the query statement and the associated text in the triple; and the positive sample set is composed of positive samples corresponding to multiple triples.

[0023] In a possible implementation, the multiple documents are documents in a preset knowledge base.

[0024] In a possible implementation, the similarity determination unit is further used to calculate a first similarity between the target query statement and the query statements of the remaining positive samples, a second similarity between the target query statement and the associated texts of the remaining positive samples, a third similarity between the target associated text and the query statements of the remaining positive samples, and a fourth similarity between the target associated text and the associated texts of the remaining positive samples; and the extraction unit is further used to, in response to at least one of the first similarity, the second similarity, the third similarity and the fourth similarity being greater than or equal to the preset threshold, use the remaining positive samples as the candidate samples.

[0025] In a possible implementation, different training stages of the model correspond to different preset thresholds; and the extraction unit is further used to: for each training stage, use at least one remaining positive sample whose similarity is greater than or equal to the preset threshold of the training stage as a candidate sample for the stage; extract the associated text of each candidate sample as the first text, construct a first negative sample from the first text and the target query statement, and obtain the first negative sample set for the training stage.

[0026] In a possible implementation, different training stages of the model are determined based on one of the following methods: training rounds, a value of a loss function, and an accuracy rate on a validation set.

[0027] In a possible implementation, the similarity is cosine similarity.

[0028] In a possible implementation, the device also includes an integration unit, which is used to: determine at least one second text from the document where the target associated text is located, construct a second negative sample including the second text and the query statement for each second text, and obtain a second negative sample set; integrate the first negative sample set and the second negative sample set to obtain a target negative sample set.

[0029] In a possible implementation manner, the integration unit is further used to: determine the context of the target-associated text in the document where the target-associated text is located; and use the context as the second text.

[0030] In a fourth aspect, an embodiment of the present application provides a model training device, comprising: a second acquisition unit, used to acquire a positive sample set and a negative sample set; wherein the negative sample set is obtained by constructing a negative sample set by the negative sample construction device described in the third aspect and various possible implementation methods of the third aspect; a training sample construction unit, used to construct multiple training samples based on the positive samples in the positive sample set and the negative samples in the negative sample set; wherein each training sample includes the associated text of the query statement and one or more of the first text of the query statement, as well as the query statement; a fine-tuning unit, used to use multiple of the training samples to fine-tune the recall model to obtain a trained recall model.

[0031] In a possible implementation, the training sample construction unit is further used to: determine the first negative sample sets corresponding to different training stages from the negative sample set; wherein different training stages are used to determine different similarity thresholds for negative samples; for each training stage, construct multiple training samples for that stage based on the positive samples in the positive sample set and the negative samples in the first negative sample set of that stage; the fine-tuning unit is further used to: in each training stage, use the multiple training samples of that stage to fine-tune the recall model to obtain the recall model after training for that stage.

[0032] In a fifth aspect, an embodiment of the present application provides a negative sample construction device, including: a memory, a processor;

[0033] The memory stores computer-executable instructions;

[0034] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the above first aspect and / or various possible implementations of the first aspect.

[0035] In a sixth aspect, an embodiment of the present application provides a model training device, including: a memory, a processor;

[0036] The memory stores computer-executable instructions;

[0037] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the above second aspect and / or various possible implementations of the second aspect.

[0038] In the seventh aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer execution instructions are stored. When the computer execution instructions are executed by a processor, they are used to implement the first aspect and / or various possible implementation methods of the first aspect as above, or to implement the second aspect and / or various possible implementation methods of the second aspect as above.

[0039] In an eighth aspect, an embodiment of the present application provides a computer program product, comprising a computer program, which, when executed by a processor, implements the first aspect above and / or various possible implementations of the first aspect, or implements the second aspect above and / or various possible implementations of the second aspect.

[0040] The negative sample construction, model training method and device provided in the embodiment of the present application can calculate the similarity between at least one of the query statement and associated text in the positive sample and the query statement and associated text of any other positive sample for each positive sample in the positive sample set, and construct a negative sample with the associated text in the other positive samples whose similarity is greater than a preset threshold and the query statement of the positive sample. The negative sample including the query statement thus generated has a certain correlation with the positive sample, which reduces the distinction between the negative sample and the positive sample and increases the interference of the negative sample on the positive sample. When the above-mentioned negative samples and positive samples are used to train the recall model, the recall model can learn useful information more efficiently, thereby achieving higher recall accuracy and better generalization ability, and can improve the accuracy of the recall model recall data (retrieval results), which is conducive to improving the matching degree between the content generated by the RAG system and user needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0042] Figure 1 Schematic diagram of application scenarios provided for this application;

[0043] Figure 2 Schematic diagram of the negative sample construction method provided for this application Figure 1 ;

[0044] Figure 3 Schematic diagram of the negative sample construction method provided for this application Figure 2 ;

[0045] Figure 4 A flowchart of the model training method provided in this application;

[0046] Figure 5 A schematic diagram of the structure of the negative sample construction device provided in this application;

[0047] Figure 6 A schematic diagram of the structure of the model training device provided in this application;

[0048] Figure 7 A schematic diagram of the structure of a negative sample construction device or a model training device provided in this application.

[0049] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION

[0050] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. Instead, they are only examples of devices and methods consistent with some aspects of the present application.

[0051] First, the terms involved in this application are explained:

[0052] Retrieval-augmented generation (RAG) is a natural language processing model architecture that combines information retrieval and text generation techniques. Additional information is obtained from external databases to assist the model in generating content. RAG mainly includes the following stages: First, the retrieval stage. When given an input query or question, RAG first uses the retrieval module to retrieve one or more document fragments that are most relevant to the query from a large number of documents or knowledge bases. This process is similar to how search engines work, but it is usually more focused on specific tasks or fields. Second, the generation stage. Based on one or more retrieved relevant document fragments, RAG uses a pre-trained generative model to build a coherent and accurate answer. Generative models are able to understand the context and create new content based on the information provided.

[0053] Please refer to Figure 1 , Figure 1 The following is a schematic diagram of the RAG system generating response content based on the query statement. Figure 1 As shown, the RAG system includes a recall model 101 and a language model 103. The RAG system can receive a query statement 11, which is input into the recall model 101. The recall model 101 vectorizes the query statement and then searches it in the database. The text 12 matching the query statement can be retrieved from the database 102. The query statement 11 and the text 12 are then input into the language model 103. The language model 103 processes the text 12 according to the query statement 11 to obtain the reply content 13 corresponding to the query statement 11.

[0054] The recall model 101 in the RAG system mainly vectorizes the query statement and then retrieves the relevant text fragments from the database (the database can be a knowledge base corresponding to different fields). Since the language model generates the reply content based on the text fragments retrieved by the recall model, the performance of the recall model is closely related to the accuracy of the reply content generated by the RAG system. When the RAG system is applied to a specific field, in order to improve the accuracy of the reply content generated by the RAG system, the recall model of the RAG system needs to be fine-tuned. At present, the training samples for fine-tuning the recall model, especially the negative sample data, are relatively simple and cannot provide sufficient training information. Therefore, the accuracy of the data recalled by the recall model is low, resulting in the reply content generated by the RAG system being poorly matched with the user's needs.

[0055] The solution provided by the present disclosure determines negative samples through multi-dimensional similarity based on the positive sample set, and can obtain negative samples that are less distinguishable from the positive samples. The above positive and negative samples provide complex training information. Therefore, the retrieval results recalled by the recall model obtained from the above training data are more accurate, which helps to improve the matching degree between the reply content generated by the RAG system and the user's needs.

[0056] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0057] Please refer to Figure 2 , Figure 2 Schematic diagram of the negative sample construction method provided for this application Figure 1 ,like Figure 2 As shown, the method includes:

[0058] S201: Acquire a positive sample set, wherein the positive sample set includes a plurality of positive samples, and each positive sample includes a query statement and associated text of the query statement.

[0059] The execution subject of the negative sample construction method can be various electronic devices, such as servers and terminal devices. The above execution subject can obtain the positive sample set through various ways.

[0060] The positive sample set may include multiple positive samples. The positive sample may include a query statement and associated text. The query statement may be, for example, input data to the model, and the associated text may be, for example, text that the model is expected to retrieve from a preset knowledge base according to the query statement.

[0061] In some embodiments, each positive sample in the positive sample set may be determined based on data stored in a database. The database may store multiple documents. Each document may include one or more paragraphs of text. Specifically, for each document stored in the database, a corresponding query statement is constructed based on one or more paragraphs in the document, and the one or more paragraphs are used as associated texts corresponding to the query statement.

[0062] In some implementations, the above step S201 includes: processing multiple documents using a natural language processing model, and constructing a positive sample set according to the processing results.

[0063] In these embodiments, prompt information including document content can be input into a natural language processing model. The prompt information includes information indicating that the natural language processing model generates a query statement-associated text according to the document content, rules for generating a query statement and associated text from the document content, information on the output format of the query statement and associated text, etc. In addition, the prompt information can also include examples of generating a query statement and corresponding associated text according to the text content. The processing result of the natural language processing model processing each document can include at least one query statement-associated text pair. Each query statement-associated text pair can be used as a positive sample. A positive sample set can be constructed by multiple query statement-associated text pairs obtained by processing multiple documents.

[0064] A positive sample set is obtained by processing multiple documents through a natural language processing model, thereby improving the efficiency of obtaining the positive sample set.

[0065] In some embodiments, a natural language processing model is used to process multiple documents, and a positive sample set is constructed according to the processing results, including the following steps:

[0066] First, a natural language processing model is used to process multiple documents to obtain multiple triples, where each triple includes a query statement, a reply content, and associated text; wherein the natural language model generates the reply content according to the associated text, and the reply content is used to reply to the query statement.

[0067] Illustratively, the above-mentioned multiple documents may be stored in a database.

[0068] Each document can be read from the database, input into the natural language processing model, and instruct the natural language processing model to output a triple including a query statement (query), a reply content (answer), and an associated text (text). The natural language processing model here can be any type of model with natural language processing capabilities. The natural language processing model is used to automatically extract query-answer-text triples from multiple documents. The query can be a query statement generated by the natural language processing model, and the text can be associated text related to the query statement. The answer can be the reply content generated by the natural language processing model based on the text.

[0069] Secondly, for each triple, a positive sample is constructed from the query sentence and associated text in the triple.

[0070] Finally, the positive sample set is composed of the positive samples corresponding to multiple triplets.

[0071] For example, for the above triple query statement: "Who is the author of Journey to the West?", the response content is: "Wu Cheng'en"; the associated text is: "The author of Journey to the West is Wu Cheng'en", the positive sample obtained is: {query statement: "Who is the author of Journey to the West?", associated sample: "The author of Journey to the West is Wu Cheng'en"}.

[0072] After obtaining multiple triplets, positive samples corresponding to the multiple triplets can be generated, and a positive sample set can be constructed from the multiple positive samples.

[0073] By extracting triplets from documents through a large model and constructing positive samples based on the query and text in the triples, the time required to generate training samples is reduced, the efficiency of constructing training samples can be improved, and the cost and time required to generate training samples can be reduced.

[0074] In some embodiments, the method further comprises:

[0075] First, for each triple, a preset verification rule may be called to verify the accuracy of the triple.

[0076] Secondly, for the triplets that pass accuracy verification, positive samples are generated.

[0077] The above-mentioned preset verification rules may be pre-set and used to verify the triples output by the natural language processing model.

[0078] For example, the above-mentioned verification rules may include whether the above-mentioned associated text includes reply content; whether the above-mentioned associated sentence includes key words and phrases of the query sentence; whether the semantic similarity between the associated sentence and the query sentence is greater than a preset threshold, etc.

[0079] After the accuracy of the triple is verified, a positive sample can be constructed based on the triple. Specifically, the query statement of the triple is used as the query statement of the positive sample, and the associated text in the triple is used as the associated text of the positive sample to obtain the positive sample.

[0080] In these implementations, the extracted triples are verified to ensure that the corresponding reply content can be found in the corresponding associated text. Invalid or low-quality triples can be removed, and the quality of the formed positive samples can be improved.

[0081] In some embodiments, the plurality of documents may be documents in a preset knowledge base. The preset knowledge base may be a knowledge base in a preset field. The knowledge base stores a plurality of documents in the preset field.

[0082] By processing the documents in the above-mentioned preset knowledge base, a positive sample set and a negative sample set are generated for fine-tuning the recall model. When the recall model fine-tuned using the positive sample set and the negative sample set is applied to the above-mentioned preset field, it can return the retrieval results more accurately, thereby helping to improve the accuracy of the response content generated by the RAG system in the preset field.

[0083] S202: Taking any positive sample in the positive sample set as a target positive sample, and for each of the remaining positive samples, calculating the similarity between at least one of the target query statement and the target associated text of the target positive sample and the query statement and the associated text of the remaining positive samples.

[0084] The positive sample set can be obtained through the above step S201. For each positive sample in the positive sample set, the positive sample can be used as a seed sample, and candidate samples related to the seed can be found from the remaining positive samples, and negative samples can be constructed based on the candidate samples.

[0085] Specifically, for any positive sample in the positive sample set, the positive sample is used as a target positive sample, and multiple positive samples other than the target positive sample in the positive sample set are remaining positive samples.

[0086] The similarity between at least one sample element (sample element includes target query statement and target associated text) in the target positive sample and each sample element of any other positive sample can be calculated. That is, for any other positive sample, the similarity between the target query statement of the target positive sample and the query statement and associated text of the other positive sample can be calculated, and the similarity between the target associated text and the query statement and associated text of the other positive sample can also be calculated. Alternatively, the similarity between the target query statement and the query statement and associated text of the other positive sample can be calculated, and the similarity between the target associated text and the query statement and associated text of the other positive sample can also be calculated.

[0087] S203: taking at least one remaining positive sample whose similarity is greater than or equal to a preset threshold as a candidate sample, extracting the associated text in each candidate sample as the first text, constructing a first negative sample including the first text and the target query sentence, and obtaining a first negative sample set.

[0088] The remaining positive samples with any of the above similarities greater than a preset threshold can be used as candidate samples, and multiple candidate samples can be obtained.

[0089] The preset threshold value can be set according to specific application scenarios and is not limited here. Schematically, the preset threshold value can be 0.6, for example.

[0090] For each candidate sample, associated text may be extracted from the candidate sample as the first text, and a first negative sample including the first text and the target query statement may be constructed.

[0091] Through the above method, the first negative sample set can be composed of the first negative samples corresponding to each of the multiple candidate samples.

[0092] In this embodiment, for each positive sample in the positive sample set, the similarity between at least one of the query statement and the associated text in the positive sample and the query statement and the associated text of any other positive sample can be calculated, and the first negative sample can be constructed based on the associated text in the other positive samples whose similarity is greater than a preset threshold and the query statement of the positive sample. The negative sample thus generated has a certain correlation with the positive sample, which reduces the distinction between the negative sample and the positive sample and increases the interference of the negative sample on the positive sample. When the above-mentioned negative samples and positive samples are used to train the model, the model can learn from the above-mentioned negative samples the features, boundary conditions and / or confusing features that are ambiguous with the positive samples, thereby improving the retrieval accuracy of the recall model. The retrieval results recalled by the recall model obtained from the above-mentioned training data have a high accuracy, which helps to improve the matching degree between the reply content generated by the RAG system and the user's needs. In addition, the recall model trained by the above-mentioned negative sample set has good generalization ability.

[0093] In some embodiments, the above step S202 includes the following steps:

[0094] Calculate a first similarity between the target query statement and the query statements of the remaining positive samples, a second similarity between the target query statement and the associated texts of the remaining positive samples, a third similarity between the target associated text and the query statements of the remaining positive samples, and a fourth similarity between the target associated text and the associated texts of the remaining positive samples;

[0095] Accordingly, step S203 includes:

[0096] In response to at least one of the first similarity, the second similarity, the third similarity and the fourth similarity corresponding to the remaining positive samples being greater than or equal to a preset threshold, the remaining positive samples are taken as candidate samples.

[0097] In these embodiments, for any remaining positive samples, the first similarity, the second similarity, the third similarity, and the fourth similarity corresponding to the remaining positive samples may be calculated.

[0098] If one or more of the first similarity, the second similarity, the third similarity and the fourth similarity corresponding to the remaining positive samples is greater than or equal to a preset threshold, the remaining positive samples are candidate samples.

[0099] In one example, the first similarity, the second similarity, the third similarity, and the fourth similarity may correspond to the same preset threshold. In this example, if any of the first similarity, the second similarity, the third similarity, and the fourth similarity is greater than or equal to the preset threshold, the remaining positive samples may be candidate samples. If the first similarity, the second similarity, the third similarity, and the fourth similarity are all less than the preset threshold, the remaining positive samples are non-candidate samples.

[0100] The first similarity can reflect the similarity between the target associated text and the query statements of the remaining positive samples. By filtering out text segments completely irrelevant to the target positive sample according to whether the first similarity is greater than a preset threshold, the influence of noise data can be reduced.

[0101] The second similarity can reflect the similarity between the target query sentence and the associated text of the remaining positive samples. The above second similarity can be used to obtain candidate samples that are highly relevant to the target query sentence. Since the associated text of the above candidate samples is highly relevant to the target associated text in the target positive sample, it is helpful to help the model learn to distinguish text fragments that are highly relevant but not the correct answer.

[0102] The third similarity can reflect the similarity between the target associated text of the target positive sample and the query statement of the candidate sample. This can determine which queries have a high correlation with the associated text of the known correct answer, which helps the model learn a wider range of query patterns during the training process and improve its generalization ability.

[0103] The fourth similarity can reflect the similarity between the target associated text of the target positive sample and the associated text of the candidate sample, which can help determine which text fragments are not the associated text of the query sentence but have a high similarity with the associated text. In this way, adversarial samples can be generated to help the model better learn to distinguish between positive and negative samples.

[0104] The above similarity can be calculated using various text similarity algorithms, for example, a cosine similarity algorithm, an edit distance algorithm, a word embedding similarity algorithm, etc.

[0105] The first similarity, the second similarity, the third similarity and the fourth similarity may be cosine similarities. Taking the second similarity as an example, the second similarity may be calculated based on the following formula (1):

[0106] (1);

[0107] in," " is the vector dot product operator," ” is the Euclidean norm of vector X (i.e., the length of the vector). is the query sentence in the target positive sample, is the associated text in any of the remaining positive samples.

[0108] The above preset threshold can be any value greater than 0 and less than 1. In one example, the first threshold can be 0.6. In this example, the associated text in the first negative sample set of the query sentence in the Mth target positive sample is represented as follows:

[0109]

[0110] in, is the target query statement of the Mth target positive sample, is the query statement for any positive sample other than the Mth target positive sample; is the target associated text of the Mth target positive sample, It can be the associated text of any remaining positive sample except the Mth target positive sample.

[0111] In another example, different preset thresholds may be set for the first similarity, the second similarity, the third similarity, and the fourth similarity, for example, a first similarity threshold is set for the first similarity, a second similarity threshold is set for the second similarity, a third similarity threshold is set for the third similarity, and a fourth similarity threshold is set for the fourth similarity. In this example, if any one of the first similarity, the second similarity, the third similarity, and the fourth similarity is greater than or equal to the corresponding preset threshold, the remaining positive samples may be used as candidate samples. If the first similarity, the second similarity, the third similarity, and the fourth similarity are respectively less than the corresponding preset thresholds, the remaining positive samples are non-candidate samples.

[0112] In these implementations, based on the above target positive samples, a plurality of negative samples with low discrimination from the positive samples are generated in one or more documents using a multi-dimensional similarity metric. This helps to expand the negative sample set, and when the model is trained using the negative sample set, the model can learn more detailed discrimination, which helps to improve the accuracy and generalization ability of the model.

[0113] In some embodiments, different training stages correspond to different preset thresholds, and the above step S203 includes:

[0114] For each training stage, at least one remaining positive sample whose similarity is greater than or equal to a preset threshold of the training stage is used as a candidate sample of the stage;

[0115] The associated text of each candidate sample is extracted as the first text, and the first negative sample is constructed by the first text and the target query sentence to obtain the first negative sample set of the training stage.

[0116] In these embodiments, the training process of the model can be divided into at least two training stages, each training stage corresponds to a first negative sample set, and the preset thresholds for determining the first negative samples in each training stage are different.

[0117] The following is an example of a training phase including three phases, which includes an initial training phase, an intermediate training phase, and a late training phase. The preset threshold for determining the first negative sample in the initial training phase can be, for example, 0.6. That is, for any target positive sample, the first similarity between the target query statement of the target positive sample and the query statement of any other positive sample, the second similarity between the target query statement and the associated text of the other positive sample, the third similarity between the target associated text and the query statement of the other positive sample, and the fourth similarity between the target associated text and the associated text of the other positive sample can be calculated. If at least one of the first similarity, the second similarity, the third similarity, and the fourth similarity corresponding to any other positive sample is greater than or equal to the preset threshold of 0.6, the other positive sample is a candidate sample. The associated text of the candidate sample is used as the first text to construct the first negative sample of the initial training phase including the target query statement and the first text. Then, multiple first negative samples corresponding to each query statement are determined to obtain a first negative sample set of the initial training phase.

[0118] Similarly, the preset threshold for determining the first negative sample in the middle training stage may be, for example, 0.7; the preset threshold for determining the first negative sample in the later training stage may be, for example, 0.8. The first negative sample set of each stage may be obtained in a manner similar to that of the initial training stage.

[0119] In these embodiments, by using different preset thresholds for different training stages, the training data of different training stages can change dynamically. The model can gradually learn more interfering negative sample information as the training stage increases. Compared with using highly interfering negative sample information for training from the beginning of training, the training efficiency of the model can be improved.

[0120] In some embodiments, different training stages of the model are determined based on one of the following:

[0121] Training rounds, loss function values, and accuracy on the validation set.

[0122] In one example, different training stages of the model can be determined according to the training rounds. For example, if the total number of training rounds is fixed, such as 100 training rounds, rounds 1 to 30 can be used as the initial training stage, rounds 31 to 80 can be used as the intermediate training stage, and rounds 81 to 100 can be used as the late training stage.

[0123] Then the multiple first samples in the negative sample set of the current training stage are shown as follows:

[0124]

[0125] Wherein, k is an integer greater than or equal to 1 and less than or equal to the total number of training stages; a(k) is the similarity threshold corresponding to the kth training stage; is the target query statement of the Mth target positive sample; is the query statement for any positive sample other than the Mth target positive sample; is the target associated text of the Mth target positive sample; It can be the associated text of any remaining positive sample except the Mth target positive sample.

[0126] Through multiple rounds of iterative fine-tuning, the model gradually learns the ability to distinguish between positive and negative examples, and finally obtains an optimized recall model.

[0127] In one example, different training stages of the model can be determined according to the value of the loss function.

[0128] For any loss function, different training tasks use different ranges of loss function values ​​determined by the loss function. For a text recall task and a specified loss function, it can be determined that the range of the loss function value of the task can be predicted, so the training of the model can be divided into multiple training stages according to the value of the loss function. For example, assuming that the loss function ranges from 0 to 100, the loss function value greater than 70 can be used as the initial training stage, the loss function value range of 30 to 70 can be used as the intermediate stage, and the training stage with a loss function value less than 30 can be used as the late training stage.

[0129] In one example, the model training process can be divided into multiple training stages according to the accuracy value on the validation set. After one round of training of the model using one or more training samples, the validation data set can be used for validation. The training with a validation accuracy less than 0.5 can be regarded as the initial training stage, the training with a validation accuracy greater than 0.5 and less than 0.9 can be regarded as the intermediate training stage, and the training with a validation accuracy greater than 0.9 or more can be regarded as the late training.

[0130] The above embodiments provide content for dividing the training process into different training stages according to different methods, which helps to split the training process in different model training scenarios, so as to prepare different training data for training according to different stages.

[0131] Please refer to 3, Figure 3 Schematic diagram of the negative sample construction method provided for this application Figure 2 ,like Figure 3 As shown, the method includes Figure 2 In addition to the same steps S301 to S303 as steps S201 to S203 of the illustrated embodiment, the following steps are also included:

[0132] S304: Determine at least one second text from the document where the target associated text is located, and construct a second negative sample including the second text and the query statement for each second text to obtain a second negative sample set.

[0133] In this embodiment, the document containing the target associated text may be referred to as the original document. The target associated text may be removed from the original document to obtain the remaining text content of the original document, and at least one second text may be determined from the remaining text content according to a preset rule. For example, at least one second text may be determined in the remaining text content in units of paragraphs. Specifically, each paragraph in the remaining text content may be used as a second text; at least one second text may also be determined in the remaining text content according to semantics, and specifically, one or more paragraphs with the same semantics may be used as a second text.

[0134] In some implementations, step S305 includes the following steps:

[0135] First, in the document where the target associated text is located, the context of the target associated text is determined;

[0136] Second, use the context as the second text.

[0137] The context of the target associated text may be determined in the original document. The context of the target associated text may be, for example, text of a preset length or a preset number of paragraphs preceding the target associated text and / or text of a preset length or a preset number of paragraphs following the target associated text in the original document.

[0138] For example, if the target query is: Who is the author of novel S1? The target associated text is: The author of novel S1 is author A. The following context Tex1 of the target associated text can be obtained from the original document P: Author A also wrote novel B. Text2: Novel S1 has sold more than 500 million copies worldwide. The above Tex1 and Tex2 can be used as the second text of the target query "Who is the author of novel S1?"

[0139] In these implementations, the context of the target-associated text of the positive sample is used as the second text, and the second text has the characteristics of topic consistency and context similarity with the target-associated text, which reduces noise and improves the correlation between the negative sample and the positive sample.

[0140] S305: Integrate the first negative sample set and the second negative sample set to obtain a target negative sample set.

[0141] When integrating the first negative sample set and the second negative sample set, duplicate negative samples may be removed, thereby eliminating redundancy.

[0142] In this embodiment, by determining the second text in the document where the target associated text is located, so that the second text has thematic consistency and contextual similarity with the target associated text, irrelevant noise can be reduced, and the correlation between negative samples and positive samples can be further improved. When the target negative sample set is used to train the recall model, it can help to distinguish positive samples and negative samples with thematic consistency and contextual similarity during the model learning process. It is beneficial to improve the retrieval accuracy and generalization ability of the model.

[0143] Please refer to Figure 4 , Figure 4 A flow chart of the model training method provided in the present disclosure, such as Figure 4 As shown, the method comprises the following steps:

[0144] S401: Obtain a positive sample set and a negative sample set; wherein the negative sample set is composed of Figure 2~Figure 3 The negative sample construction method provided in the illustrated embodiment is constructed.

[0145] S402: construct multiple training samples according to the positive samples in the positive sample set and the negative samples in the negative sample set; wherein each training sample includes one or more of the associated text of the query statement and the first text of the query statement, and the query statement.

[0146] S403: Fine-tune the recall model using multiple training samples to obtain a trained recall model.

[0147] The executor of the model training method can be any electronic device, such as a server and a terminal device.

[0148] The recall model can be a vectorization model for vectorizing query statements and document contents in a database, and retrieving the vectorized document contents according to the vectorized query statements to determine the associated text of the query statements. The recall model can be a pre-trained model. When applying the recall model to a specific field, it is necessary to fine-tune the recall model using the knowledge of the field so that the recall model has a higher accuracy when applied to the field.

[0149] In the process of training the recall model, a preset loss function may be used to calculate the loss value. The loss function may be, for example, various loss functions, such as a cross entropy loss function, a mean square error loss function, a contrast loss function, and the like.

[0150] The positive sample set may include multiple positive samples, and each positive sample may include a query statement and associated text of the query statement. Specifically, reference may be made to the relevant descriptions in the embodiments shown in FIG. 2 and FIG. 3, which will not be repeated here.

[0151] The negative sample set may include multiple negative samples, and each negative sample may include a query statement and a first text of the query statement. Specifically, reference may be made to the relevant descriptions in the embodiments shown in FIG. 2 and FIG. 3, which will not be repeated here.

[0152] A plurality of training samples may be constructed according to a plurality of positive samples in a positive sample set and a plurality of negative samples in a negative sample set.

[0153] In one example, any positive sample can be extracted from the positive sample set, and the positive sample includes a query statement and an associated text of the query statement; one or more negative samples including the query statement can be found from the negative sample set. In addition to the query statement, each negative sample also includes a first text. For the query statement, one or more training samples including the query statement, the associated text and the first text can be constructed; the structure of any training sample is as follows {query statement, associated text, first text}.

[0154] Using the above training samples to train the recall model can achieve the following effects: First, enhance the model's ability to distinguish. By clearly indicating which texts are related to the query (associated text) and irrelevant (first text), the model can learn better feature representations to distinguish between relevant and irrelevant texts, thereby improving its ability to distinguish. Second, promote generalization. Introducing different types of positive and negative samples during training can help the model better understand the intent behind the query and learn to generalize to unseen data, which improves the model's generalization. Third, improve learning efficiency. Since each training sample contains information about the query, the correct answer, and the wrong answer, a single sample can provide more supervisory information, which helps speed up the model's learning process and may reduce the amount of training data required.

[0155] In this example, the loss functions used to train the recall model include but are not limited to: Triplet Loss, Hinge Loss, Contrastive Loss, and Binary Cross Entropy Loss.

[0156] In one example, a training sample may be composed of each positive sample and a positive sample label, wherein the positive sample label indicates that the sample is a positive sample. For example, a training sample composed of a positive sample and a positive sample label may be represented as follows: {query statement, associated text}-label 1, wherein label 1 includes any label indicating that the training sample is a positive sample, and label 1 may include one or more of numbers, symbols, and text. A training sample may be composed of each negative sample and a negative sample label, wherein the negative sample label indicates that the sample is a negative sample. For example, a training sample composed of a negative sample and a negative sample label may be represented as follows: {query statement, first text}-label 2, wherein label 2 includes any label indicating that the training sample is a negative sample, and label 2 may include one or more of numbers, symbols, and text.

[0157] Positive samples can be input into the recall model to increase the similarity score between the query statement and the associated text in the positive sample; negative samples can be input into the recall model to reduce the similarity score between the query statement and the associated text in the negative sample. After the above training, a trained recall model is obtained.

[0158] In this example, the recall model is trained by training samples consisting of positive samples and positive sample labels, and negative samples and negative sample labels, which can simplify data preparation for training samples.

[0159] When training the recall model in this example, the loss functions used include but are not limited to: SoftmaxLoss, Pairwise Ranking Loss, etc.

[0160] In this embodiment, since the associated text of the query statement has a certain correlation with the first text, the distinction between the associated text of the query statement and the first text is small, and the first text has a greater interference on the associated text. The recall model obtained after training with the training sample has higher detection accuracy and better generalization ability.

[0161] In some embodiments, the above step S402 includes the following sub-steps:

[0162] First, first negative sample sets corresponding to different training stages are determined from the negative sample set; wherein different training stages are used to determine different similarity thresholds of negative samples.

[0163] Secondly, for each training stage, multiple training samples of the stage are constructed based on the positive samples in the positive sample set and the negative samples in the first negative sample set of the stage.

[0164] The above-mentioned different training stages can be determined according to the training rounds, or according to the value of the loss function, or according to the range of the loss function value; or according to the value of the accuracy on the validation set.

[0165] In these embodiments, negative samples determined by different similarity preset thresholds are used in different training stages, that is, the training data in different training stages changes dynamically, and the recall model gradually learns more interfering negative sample information as the training stage increases. Compared with using highly interfering negative sample information for training from the beginning of training, the training efficiency of the model can be improved.

[0166] Figure 5 A schematic diagram of the structure of the negative sample construction device provided in this application, such as Figure 5 As shown, the negative sample construction device 50 provided in this embodiment includes:

[0167] A first acquisition unit 501 is used to acquire a positive sample set, wherein the positive sample set includes a plurality of positive samples, and each positive sample includes a query statement and associated text of the query statement;

[0168] A similarity determination unit 502 is used to take any positive sample in the positive sample set as a target positive sample, and for each of the remaining positive samples, calculate the similarity between at least one of the target query statement and the target associated text of the target positive sample and the query statement and the associated text of the remaining positive samples;

[0169] The extraction unit 503 is used to take at least one remaining positive sample with a similarity greater than or equal to a preset threshold as a candidate sample, extract the associated text in each candidate sample as the first text, construct a first negative sample including the first text and the target query statement, and obtain a first negative sample set.

[0170] In a possible implementation manner, the first acquiring unit 501 is further configured to:

[0171] Use the natural language processing model to process multiple documents and build a positive sample set based on the processing results.

[0172] In a possible implementation manner, the first acquiring unit 501 is further configured to:

[0173] A natural language processing model is used to process multiple documents to obtain multiple triples, wherein each triple includes a query statement, a reply content, and associated text; wherein the natural language model generates the reply content according to the associated text, and the reply content is used to reply to the query statement;

[0174] For each triple, construct a positive sample from the query sentence and associated text in the triple;

[0175] The positive sample set is composed of positive samples corresponding to multiple triplets.

[0176] In a possible implementation, the similarity determination unit 502 is further used to: calculate a first similarity between the target query statement and the query statements of the remaining positive samples, a second similarity between the target query statement and the associated texts of the remaining positive samples, a third similarity between the target associated text and the query statements of the remaining positive samples, and a fourth similarity between the target associated text and the associated texts of the remaining positive samples;

[0177] The extraction unit 503 is further configured to: in response to at least one of the first similarity, the second similarity, the third similarity and the fourth similarity being greater than or equal to a preset threshold, take the remaining positive samples as candidate samples.

[0178] In a possible implementation, different training stages of the model correspond to different preset thresholds, and the extraction unit 503 is further used to:

[0179] For each training stage, at least one remaining positive sample whose similarity is greater than or equal to a preset threshold of the training stage is used as a candidate sample of the stage;

[0180] The associated text of each candidate sample is extracted as the first text, and the first negative sample is constructed by the first text and the target query sentence to obtain the first negative sample set of the training stage.

[0181] In one possible implementation, different training stages of the model are determined based on one of the following methods:

[0182] Training rounds, loss function values, and accuracy on the validation set.

[0183] In a possible implementation, the similarity is cosine similarity.

[0184] In a possible implementation, the device 50 further includes an integration unit (not shown in the figure), and the integration unit is used to:

[0185] Determine at least one second text from the document where the target associated text is located, construct a second negative sample including the second text and the query statement for each second text, and obtain a second negative sample set;

[0186] The first negative sample set and the second negative sample set are integrated to obtain a target negative sample set.

[0187] In a possible implementation manner, the integration unit is further configured to:

[0188] In the document where the target associated text is located, determining the context of the target associated text;

[0189] Use the context as the second text.

[0190] The negative sample construction device provided in this embodiment can execute the method provided in the above negative sample construction method embodiment, and its implementation principle and technical effect are similar, which will not be described in detail in this embodiment.

[0191] Figure 6 A schematic diagram of the structure of the model training device provided in this application, such as Figure 6 As shown, the model training device 60 provided in this embodiment includes:

[0192] The second acquisition unit 601 is used to acquire a positive sample set and a negative sample set, wherein the negative sample set is composed of Figure 5 The negative sample construction device shown is used to construct a negative sample set;

[0193] A training sample construction unit 602 is used to construct a plurality of training samples according to the positive samples in the positive sample set and the negative samples in the negative sample set; wherein each training sample includes one or more of the associated text of the query statement and the first text of the query statement, and the query statement;

[0194] The fine-tuning unit 603 is used to fine-tune the recall model using multiple training samples to obtain a trained recall model.

[0195] In a possible implementation, the training sample construction unit 602 is further configured to:

[0196] Determine, from the negative sample set, first negative sample sets corresponding to different training stages, wherein different training stages are used to determine different similarity thresholds of negative samples;

[0197] For each training stage, construct multiple training samples of the stage based on positive samples in the positive sample set and negative samples in the first negative sample set of the stage; and

[0198] The fine-tuning unit 603 is further used for:

[0199] In each training stage, the recall model is fine-tuned using multiple training samples of that stage to obtain the recall model trained in that stage.

[0200] The model training device provided in this embodiment can execute the method provided in the above-mentioned model training method embodiment. Its implementation principle and technical effects are similar, and will not be described in detail in this embodiment.

[0201] Figure 7 This is a schematic diagram of the structure of the negative sample construction device or model training device provided in this application. Figure 7 As shown, the electronic device 70 provided in this embodiment includes: at least one processor 701 and a memory 702. Optionally, the device 70 further includes a communication component 703. The processor 701, the memory 702 and the communication component 703 are connected via a bus 704.

[0202] In a specific implementation process, at least one processor 701 executes the computer-executable instructions stored in the memory 702, so that at least one processor 701 executes the above method.

[0203] The specific implementation process of the processor 701 can be found in the above method embodiment, and its implementation principle and technical effect are similar, so this embodiment will not be repeated here.

[0204] In the above embodiments, it should be understood that the processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the invention can be directly implemented as a hardware processor, or can be implemented by a combination of hardware and software modules in the processor.

[0205] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (NVM), such as at least one disk storage.

[0206] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus in the drawings of this application is not limited to only one bus or one type of bus.

[0207] The present application also provides a computer program product, including a computer program, which implements the above method when executed by a processor.

[0208] The present application also provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the above method is implemented.

[0209] The above-mentioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The readable storage medium can be any available medium that can be accessed by a general or special-purpose computer.

[0210] An exemplary readable storage medium is coupled to a processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (Application Specific Integrated Circuits, referred to as: ASIC). Of course, the processor and the readable storage medium can also exist in the device as discrete components.

[0211] The division of units is only a logical function division, and there may be other divisions in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.

[0212] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0213] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0214] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0215] Those skilled in the art can understand that all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the aforementioned storage medium includes: ROM, RAM, disk or optical disk and other media that can store program codes.

[0216] Finally, it should be noted that those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses or adaptations of the present invention, which follow the general principles of the present invention and include common knowledge or customary technical means in the art not disclosed by the present invention, are not limited to the precise structure described above and shown in the drawings, and may be modified and changed in various ways without departing from the scope thereof. The scope of the present invention is limited only by the appended claims.

Claims

1. A negative sample construction method, comprising: Acquire a positive sample set, wherein the positive sample set includes a plurality of positive samples, and each positive sample includes a query statement and associated text of the query statement; Taking any positive sample in the positive sample set as a target positive sample, and for each of the remaining positive samples, calculating the similarity between at least one of the target query statement and the target associated text of the target positive sample and the query statement and the associated text of the remaining positive samples; At least one remaining positive sample whose similarity is greater than or equal to a preset threshold is taken as a candidate sample, the associated text in each candidate sample is extracted as the first text, a first negative sample including the first text and the target query statement is constructed, and a first negative sample set is obtained.

2. The method according to claim 1, characterized in that The obtaining of a positive sample set comprises: The multiple documents are processed using a natural language processing model, and the positive sample set is constructed according to the processing results.

3. The method according to claim 2, characterized in that The method of processing the plurality of documents using a natural language processing model and constructing the positive sample set according to the processing results includes: Processing the multiple documents using a natural language processing model to obtain multiple triples, wherein each triple includes a query statement, a reply content, and associated text; wherein the natural language model generates the reply content according to the associated text, and the reply content is used to reply to the query statement; For each triple, construct a positive sample from the query sentence and associated text in the triple; The positive sample set is composed of positive samples corresponding to a plurality of triplets.

4. The method according to claim 3, characterized in that The multiple documents are documents in a preset knowledge base.

5. The method according to claim 1, characterized in that The step of calculating, for each of the remaining positive samples, the similarity between at least one of the target query statement and the target associated text of the target positive sample and the query statement and the associated text of the remaining positive samples, comprises: Calculating a first similarity between the target query and the query of the remaining positive samples, a second similarity between the target query and the associated text of the remaining positive samples, a third similarity between the target associated text and the query of the remaining positive samples, and a fourth similarity between the target associated text and the associated text of the remaining positive samples; And taking at least one remaining positive sample whose similarity is greater than or equal to a preset threshold as a candidate sample comprises: In response to at least one of the first similarity, the second similarity, the third similarity and the fourth similarity being greater than or equal to the preset threshold, the remaining positive samples are taken as the candidate samples.

6. The method according to claim 1, characterized in that Different training stages of the model correspond to different preset thresholds respectively; and at least one remaining positive sample whose similarity is greater than or equal to the preset threshold is used as a candidate sample, and the associated text in each candidate sample is extracted as the first text, and a first negative sample including the first text and the target query sentence is constructed to obtain a first negative sample set, including: For each training stage, at least one remaining positive sample whose similarity is greater than or equal to a preset threshold of the training stage is used as a candidate sample of the stage; The associated text of each candidate sample is extracted as the first text, and a first negative sample is constructed by the first text and the target query sentence to obtain a first negative sample set in the training stage.

7. The method according to claim 6, characterized in that The different training stages of the model are determined based on one of the following methods: Training rounds, loss function values, and accuracy on the validation set.

8. The method according to any one of claims 1 to 7, characterized in that: The similarity is cosine similarity.

9. The method according to any one of claims 1 to 7, characterized in that: The method further comprises: Determine at least one second text from the document where the target associated text is located, construct a second negative sample including the second text and the query statement for each second text, and obtain a second negative sample set; The first negative sample set and the second negative sample set are integrated to obtain a target negative sample set.

10. The method according to claim 9, characterized in that The step of determining at least one second text from the document where the target associated text is located includes: In the document where the target associated text is located, determining the context of the target associated text; The context is used as the second text.

11. A model training method, comprising: Obtain a positive sample set and a negative sample set; wherein the negative sample set is obtained by the method described in any one of claims 1 to 10; Constructing a plurality of training samples according to the positive samples in the positive sample set and the negative samples in the negative sample set; wherein each training sample includes one or more of the associated text of the query statement and the first text of the query statement, as well as the query statement; The recall model is fine-tuned using a plurality of the training samples to obtain a trained recall model.

12. The method according to claim 11, characterized in that The constructing a plurality of training samples according to the positive samples in the positive sample set and the negative samples in the negative sample set comprises: Determine, from the negative sample set, first negative sample sets corresponding to different training stages, wherein different training stages are used to determine different similarity thresholds of negative samples; For each training stage, constructing a plurality of training samples of the stage based on the positive samples in the positive sample set and the negative samples in the first negative sample set of the stage; And the use of the plurality of training samples to fine-tune the recall model to obtain a trained recall model includes: In each training stage, the recall model is fine-tuned using a plurality of training samples in that stage to obtain a recall model trained in that stage.

13. A negative sample construction device, comprising: A first acquisition unit is used to acquire a positive sample set, wherein the positive sample set includes a plurality of positive samples, and each positive sample includes a query statement and associated text of the query statement; A similarity determination unit, configured to take any positive sample in the positive sample set as a target positive sample, and for each of the remaining positive samples, calculate the similarity between at least one of a target query statement and a target associated text of the target positive sample and the query statement and the associated text of the remaining positive samples; The extraction unit is used to take at least one remaining positive sample whose similarity is greater than or equal to a preset threshold as a candidate sample, extract the associated text in each candidate sample as the first text, construct a first negative sample including the first text and the target query statement, and obtain a first negative sample set.

14. The device according to claim 13, characterized in that The first acquisition unit is further used for: The multiple documents are processed using a natural language processing model, and the positive sample set is constructed according to the processing results.

15. The device according to claim 14, characterized in that The first acquisition unit is further used for: Processing the multiple documents using a natural language processing model to obtain multiple triples, wherein each triple includes a query statement, a reply content, and associated text; wherein the natural language model generates the reply content according to the associated text, and the reply content is used to reply to the query statement; For each triple, construct a positive sample from the query sentence and associated text in the triple; The positive sample set is composed of positive samples corresponding to a plurality of triplets.

16. The device according to claim 15, characterized in that The multiple documents are documents in a preset knowledge base.

17. The device according to claim 13, characterized in that The similarity determination unit is further configured to: Calculating a first similarity between the target query and the query of the remaining positive samples, a second similarity between the target query and the associated text of the remaining positive samples, a third similarity between the target associated text and the query of the remaining positive samples, and a fourth similarity between the target associated text and the associated text of the remaining positive samples; And the extraction unit is further used for: In response to at least one of the first similarity, the second similarity, the third similarity and the fourth similarity being greater than or equal to the preset threshold, the remaining positive samples are taken as the candidate samples.

18. The device according to claim 13, characterized in that Different training stages of the model correspond to different preset thresholds; and the extraction unit is further used to: For each training stage, at least one remaining positive sample whose similarity is greater than or equal to a preset threshold of the training stage is used as a candidate sample of the stage; The associated text of each candidate sample is extracted as the first text, and a first negative sample is constructed by the first text and the target query sentence to obtain a first negative sample set in the training stage.

19. The device according to claim 18, characterized in that The different training stages of the model include determining based on one of the following methods: Training rounds, loss function values, and accuracy on the validation set.

20. The device according to any one of claims 13 to 19, characterized in that The similarity is cosine similarity.

21. The device according to any one of claims 13 to 19, characterized in that The device further comprises an integration unit, which is used for: Determine at least one second text from the document where the target associated text is located, construct a second negative sample including the second text and the query statement for each second text, and obtain a second negative sample set; The first negative sample set and the second negative sample set are integrated to obtain a target negative sample set.

22. The device according to claim 21, characterized in that The integration unit is further used for: In the document where the target associated text is located, determining the context of the target associated text; The context is used as the second text.

23. A model training device, comprising: A second acquisition unit, used to acquire a positive sample set and a negative sample set; wherein the negative sample set is constructed by the method described in any one of claims 1 to 10; A training sample construction unit, configured to construct a plurality of training samples according to the positive samples in the positive sample set and the negative samples in the negative sample set; wherein each training sample includes one or more of the associated text of the query statement and the first text of the query statement, and the query statement; A fine-tuning unit is used to fine-tune the recall model using the multiple training samples to obtain a trained recall model.

24. The device according to claim 23, characterized in that The training sample construction unit is further used to: Determine, from the negative sample set, first negative sample sets corresponding to different training stages, wherein different training stages are used to determine different similarity thresholds of negative samples; For each training stage, constructing a plurality of training samples of the stage based on the positive samples in the positive sample set and the negative samples in the first negative sample set of the stage; The fine-tuning unit is further used for: In each training stage, the recall model is fine-tuned using a plurality of training samples in that stage to obtain a recall model trained in that stage.

25. A negative sample construction device, characterized in that: include: Memory, processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the method according to any one of claims 1 to 10.

26. A model training device, characterized in that: include: Memory, processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the method according to claim 11 or 12.

27. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1-10 or 11-12 when executed by a processor.

28. A computer program product, comprising a computer program, which, when executed by a processor, implements the method of any one of claims 1-10 or 11-12.