Text search model training method, text search method, device and equipment
The text retrieval model is trained to understand complex queries by using text marks, enhancing its ability to provide accurate search results for legal provisions, addressing the limitations of existing systems in handling ambiguous or complex queries.
Patent Information
- Application Number
- JP2024163313
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2024-04-25
- Filing Date
- 2024-09-20
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2044-09-20
AI Technical Summary
Existing legal provision search systems struggle to provide accurate search results for complex or ambiguous queries, particularly for non-professionals due to difficulties in understanding fine-grained lexical relationships and complex problems.
A method for training a text retrieval model by inputting a sample text into a first retrieval model to obtain text marks, and training the model based on these marks to enhance its ability to understand complex queries, using a large-scale language model and fine-tuning techniques to improve accuracy.
The trained model effectively handles complex queries, providing accurate search results by leveraging lexical relationships and reducing computational complexity, making it suitable for non-experts to find relevant legal provisions.
Smart Images

Figure 0007787263000001 
Figure 0007787263000002 
Figure 0007787263000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to the field of artificial intelligence, particularly to technical fields such as deep learning, large-scale language models, natural language understanding and search, etc. More specifically, the present disclosure provides a method for training a text search model, a text processing method, an apparatus, an electronic device, and a storage medium. [Background technology]
[0002] With the development of artificial intelligence technology, large language models (LLMs) may be used to perform search tasks. Summary of the Invention
[0003] The present disclosure provides a method for training a text search model, a text processing method, an apparatus, an electronic device, a storage medium, and a program.
[0004] According to one aspect of the present disclosure, there is provided a method for training a text retrieval model, the method including: inputting a sample text to be queried into a first text retrieval model to obtain first output text marks; and training the first text retrieval model based on text mark tags of the sample text to be queried and the first output text marks to obtain a target text retrieval model, wherein the sample text to be queried corresponds to a first target reference text in a plurality of reference texts, the text mark tags are reference text marks of the first target reference text, the plurality of reference text marks are determined based on a reference text sequence, and the reference text sequence is determined based on a vocabulary of the plurality of reference texts; and the first text retrieval model is obtained by training an initial text retrieval model using the plurality of reference texts and the corresponding plurality of reference text marks, thereby outputting reference text marks corresponding to the reference text.
[0005] According to another aspect of the present disclosure, there is provided a text search method, including: inputting a text to be queried into a target text search model to obtain a current output result; and in response to determining that the current output result hits a second target reference text in the plurality of reference texts, determining the second target reference text as a search result corresponding to the text to be queried, wherein the target text search model is obtained by training a first text search model using a sample text to be queried and a text mark tag, the sample text to be queried corresponds to the first target reference text in the plurality of reference texts, and the text mark tag is a reference text mark of the first target reference text.
[0006] According to another aspect of the present disclosure, there is provided an apparatus for training a text retrieval model, the apparatus including: a first obtaining module for inputting a sample text to be queried into a first text retrieval model and obtaining a first output text mark; and a training module for training the first text retrieval model based on text mark tags of the sample text to be queried and the first output text mark to obtain a target text retrieval model, wherein the sample text to be queried corresponds to a first target reference text in a plurality of reference texts, the text mark tags are reference text marks of the first target reference text, the plurality of reference text marks are determined based on a reference text sequence, and the reference text sequence is determined based on a vocabulary of the plurality of reference texts, and the first text retrieval model is obtained by training an initial text retrieval model using the plurality of reference texts and the corresponding plurality of reference text marks, thereby outputting the reference text marks corresponding to the reference text.
[0007] According to another aspect of the present disclosure, there is provided a text retrieval device including: a fourth obtaining module that inputs a text to be queried into a target text retrieval model to obtain a current output result; and a determining module that, in response to determining that the current output result hits a second target reference text in the plurality of reference texts, determines the second target reference text as a search result corresponding to the text to be queried, wherein the target text retrieval model is obtained by training a first text retrieval model using a sample text to be queried and a text mark tag, the sample text to be queried corresponds to the first target reference text in the plurality of reference texts, and the text mark tag is a reference text mark of the first target reference text.
[0008] According to another aspect of the present disclosure, there is provided an electronic device including at least one processor and a memory communicatively coupled to the at least one processor, the memory storing instructions executable by the at least one processor, the instructions being executed by the at least one processor such that the at least one processor can perform a method provided in the present disclosure.
[0009] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium having computer instructions stored thereon, the computer instructions causing a computer to perform the methods provided in the present disclosure.
[0010] According to another aspect of the present disclosure, there is provided a computer program product which, when executed by a processor, implements the methods provided in the present disclosure.
[0011] It should be understood that the subject matter described in this section is not intended to identify key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will be readily apparent from the following specification. [Brief explanation of the drawings]
[0012] The drawings are for a better understanding of the invention and are not intended to limit the disclosure. [Figure 1] FIG. 1 is a flowchart of a method for training a text retrieval model according to one embodiment of the present disclosure. [Figure 2A] FIG. 2A is a schematic diagram of training an initial text retrieval model according to one embodiment of the present disclosure. [Figure 2B] FIG. 2B is a schematic diagram of training a first text retrieval model according to one embodiment of the present disclosure. [Figure 3] FIG. 3 is a flowchart of a text search method according to another embodiment of the present disclosure. [Figure 4] FIG. 4 is a schematic diagram of a text search method according to an embodiment of the present disclosure. [Figure 5] FIG. 5 is a block diagram of an apparatus for training a text retrieval model according to one embodiment of the present disclosure. [Figure 6] FIG. 10 is a block diagram of a text search device according to another embodiment of the present disclosure. [Figure 7] FIG. 7 is a block diagram of an electronic device to which a method for training a text search model and / or a method for text search according to an embodiment of the present disclosure can be applied. DETAILED DESCRIPTION OF THE INVENTION
[0013]
[0023] The following description of exemplary embodiments of the present disclosure will be made with reference to the drawings. For ease of understanding, various details of the embodiments of the present disclosure are included, but these are merely illustrative. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, the following description will omit descriptions of known functions and structures.
[0014] With the widespread use of regulatory documents such as laws and policies, users have an increasingly strong need to inquire about legal provisions and policies. Corresponding legal provision search systems and policy content search systems have also been developed. Taking legal provision search systems as an example, legal provision search systems may perform precise or fuzzy searches based on the content of a portion of a legal provision, or may provide the full text of a law.
[0015] However, some legal provision search systems are primarily aimed at legal professionals, such as lawyers. If these professionals are familiar with a certain provision of a certain law, they can use the search system to search for legal provisions with full authority. However, these search systems are essentially databases containing all laws, making it difficult for non-professionals with little or no legal knowledge to search for relevant laws.
[0016] Furthermore, the search system provides search results based on the lexical similarity between the problem and the legal provisions. However, if the problem is simple or contains legal keywords, the search system can provide good search results, but it is difficult to mine the fine-grained lexical relationships between the problem and the legal provisions. This search system has difficulty understanding complex problems. If the problem is complex, ambiguous, or unclear, the search system has difficulty providing good search results.
[0017] Therefore, to perform searches based on relatively complex problems, the present disclosure provides a method for training a text search model, which is described below.
[0018] FIG. 1 is a flowchart of a method for training a text retrieval model according to one embodiment of the present disclosure.
[0019] As shown in FIG. 1, the method 100 may include operations S110 to S120.
[0020] In operation S110, a sample text to be queried is input into a first text retrieval model to obtain a first output text mark.
[0021] In the embodiments of the present disclosure, the sample text to be queried may be a complex problem or a simple problem. For example, the sample text to be queried may be "building a steel-clad building in the west to block the sunlight from the low-rise residential area, and whether there is a national legal provision." The sample text to be queried may be a complex problem.
[0022] In an embodiment of the present disclosure, the first text retrieval model may be a generative artificial intelligence model.
[0023] In an embodiment of the present disclosure, the first text search model may be obtained by training an initial search model using a plurality of reference texts and a corresponding plurality of reference text marks. For example, the reference text may be a legal article. A legal article of a first reference text in the plurality of reference texts may be, for example, "Article 293 of the Civil Code of the People's Republic of China: When constructing a building, it shall not violate the relevant construction guidelines of the state and shall not obstruct ventilation, lighting, and sunlight of adjacent buildings." The reference text mark of the first reference text may be "55648."
[0024] In an embodiment of the present disclosure, the plurality of reference text marks may be determined based on the reference text sequence. The reference text sequence may be determined based on the vocabulary of the plurality of reference texts. For example, the reference text sequence may be a plurality of reference texts arranged in a certain order. This order may be determined based on the vocabulary of the reference texts. If the lexical similarity between two reference texts is high, the difference between the corresponding two reference text marks is small. The plurality of reference texts may include current legal provisions and repealed legal provisions. A second reference text in the plurality of reference texts may be a legal provision stating, "Article 89 of the Property Law of the People's Republic of China [Regarding Ventilation, Lighting, and Solar Radiation] When constructing a building, it shall not violate relevant state construction guidelines and shall not obstruct ventilation, lighting, and solar radiation of adjacent buildings," and the reference text mark of the second reference text may be "55942."
[0025] In an embodiment of the present disclosure, the first text retrieval module can output a reference text mark corresponding to the reference text. For example, when the first reference text is input into the first text retrieval model, the reference text mark "55648" is obtained. When the second reference text is input into the first text retrieval model, the reference text mark "55942" is obtained.
[0026] In an embodiment of the present disclosure, the first output text mark may be a string of characters that is similar to or matches the reference mark.
[0027] In operation S120, a first text retrieval model is trained based on the text mark tags of the sample text to be queried and the first output text marks to obtain a target retrieval model.
[0028] In an embodiment of the present disclosure, the sample text to be queried may correspond to a first target reference text in a plurality of reference texts. For example, the sample text to be queried may correspond to the first reference text. The first reference text may be the first target reference text.
[0029] In an embodiment of the present disclosure, the text mark tag may be the reference text mark of the first target reference text, for example, the text mark tag of the sample text to be queried may be the reference text mark "55648".
[0030] In embodiments of the present disclosure, various loss functions can be used to determine the difference between the text mark tag and the first output text mark, and the first text retrieval model can be trained based on this difference.
[0031] According to an embodiment of the present disclosure, the reference text mark is determined based on the reference text sequence related to vocabulary height, which helps the model to efficiently understand the corresponding training task. The first text retrieval model is trained using the reference text and the corresponding text, which ensures that the model parameters of the first text retrieval model have a very strong correlation with the reference text, allowing the first text retrieval model to search for the target reference text from multiple reference texts. This effectively avoids the generated text being different from the actual reference text when the text retrieval model directly generates text, and fully improves the authority of the text retrieval model. In addition, by training the first text retrieval model using the sample text to be queried and the corresponding text mark tag, sufficient interaction between the sample text to be queried and the model parameters related to the reference text can be achieved, which helps the model to understand complex problems, allows the model to accurately determine the relationship between complex problems and legal provisions, and reduces computational complexity.
[0032] Having described the training method of the present disclosure, the reference text marks of the present disclosure will now be further described.
[0033] In some embodiments, the multiple reference texts are derived from N reference text sets. A reference text set may include at least one reference text, where N may be an integer greater than 1. For example, if the reference texts are legal provisions, the entire text of the law to which the legal provisions belong may be one reference text set. The Civil Code - Property Rights Section may be one reference text set. Each legal provision in the Civil Code - Property Rights Section may be one reference text.
[0034] In some embodiments, the reference text sequence is determined based on the vocabulary of the multiple reference texts by the following operation: clustering N reference text sets respectively, and adjusting the order of the reference texts in the N reference text sets to obtain N processed text sets. For example, the full texts of laws corresponding to the N reference text sets include the "Civil Code - Property Rights Section," the "Civil Code - Contract Section," and the "Education Law." Many of the legal provisions in the "Civil Code - Contract Section" are related to contracts. Many of the legal provisions in the "Education Law" are related to education. Therefore, different reference texts in the same reference text set have close lexical relationships. Based on the vocabulary of each reference text in the reference text set, clustering can be performed on one or more reference text sets respectively, forming one or more clusters in each clustered reference text set, and adjusting the order of the reference texts in these reference text sets to obtain one or more processed text sets.
[0035] In some embodiments, the reference text sequence is determined based on the vocabulary of the multiple reference texts by the following operation: obtain a fusion text set based on the N post-processing text sets. For example, after performing a clustering process on the N reference text sets, the obtained N post-processing text sets can be N subsets of the fusion text set, respectively. Each subset may include all the reference texts in one post-processing text set.
[0036] In some embodiments, the reference text sequence is determined based on the vocabulary of the multiple reference texts by the following operation: based on the vocabulary of each of the multiple reference texts, a clustering process is performed on the fused text set, and the order of the multiple reference texts in the fused text set is adjusted to obtain a reference text sequence. For example, different legal provisions from different laws may have a close lexical relationship. Based on the vocabulary of each of the multiple reference texts in the fused text set, a clustering process is performed on the fused text set, and the order of the multiple reference texts is adjusted again to obtain a reference text sequence, so that the distance between different legal provisions with close or similar vocabulary in different laws is close. Next, marks for the multiple reference texts may be determined sequentially based on the order of the multiple reference texts in the reference text sequence.
[0037] In some other embodiments, it may be understood that random integer marking may be performed on all reference texts in the set of N reference texts. That is, the marking of each reference text may be a single random integer. However, the generative AI model employs an output method generated by individual elements (tokens). If random integer marking is performed on the reference texts, it may be difficult for the generative AI model to learn complex relationships between multiple reference texts.
[0038] According to an embodiment of the present disclosure, a clustering process is performed on N reference text sets to fully utilize the lexical relationships between the reference texts in the text sets. Furthermore, by performing a clustering process again on the multiple clustered text sets, the lexical relationships between the reference texts in different text sets can be fully utilized, and the multiple reference texts can be sorted in lexically related order so that different reference texts with high lexical similarity have the same or similar reference text mark prefixes, allowing the model to fully utilize the lexical relationships between different reference texts.
[0039] Having described the reference text and reference text marks of the present disclosure above, the first text retrieval model of the present disclosure will now be further described.
[0040] In some embodiments, the first text retrieval model may be a large-scale language model, such as a sentence-minded word model, a conversational generative pre-training (ChatGPT) model, or any other large-scale language model.
[0041] In an embodiment of the present disclosure, the first text retrieval model is obtained by training an initial text retrieval model. The initial text retrieval model may be a pre-trained model that has been trained using a large amount of text data and has strong vocabulary understanding ability. A method for training the initial text retrieval model is described below with reference to FIG. 2A.
[0042] FIG. 2A is a schematic diagram of training an initial text retrieval model according to one embodiment of the present disclosure.
[0043] In some embodiments, the first text retrieval model is obtained by training an initial text retrieval model using a plurality of reference texts and a corresponding plurality of reference text marks through the following operations: inputting the reference texts into the initial text retrieval model, and obtaining second output text marks for the reference texts. As shown in Figure 2A, reference texts 201, 202, 203, and 204 are input into an initial text retrieval model M200, respectively, to obtain an output text mark 2011 for reference text 201, an output text mark 2021 for reference text 202, an output text mark 2031 for reference text 203, and an output text mark 2041 for reference text 204. The output text marks 2011 to 2041 may be second output text marks, respectively.
[0044] In some embodiments, the first text retrieval model may be obtained by training an initial text retrieval model using a plurality of reference texts and a corresponding plurality of reference text marks by the following operation: fully fine-tuning the initial text retrieval model based on the second output text marks and the reference text marks. For example, the reference text 201 may be the first reference text. Thus, the reference text mark of the reference text 201 may be "55648". Also, for example, the reference text 202 may be the second reference text. The reference text mark of the reference text 202 may be "55942".
[0045] In an embodiment of the present disclosure, exhaustively fine-tuning the initial text retrieval model based on the second output text mark and the reference text mark includes determining a cross-entropy loss based on the second output text mark and the reference text mark. The initial text retrieval model is exhaustively fine-tuned based on the cross-entropy loss. For example, the initial text retrieval model may output multiple strings and multiple correspondence probabilities. The string with the highest correspondence probability is set as the second output text mark. The training goal of the model may be maximizing the inter-target likelihood. That is, the model parameters may be adjusted to maximize the correspondence probability of the string matching the reference text mark. A cross-entropy loss function may be used to determine the cross-entropy loss between the second output text mark and the reference text mark. The initial text retrieval model is exhaustively fine-tuned based on the cross-entropy loss corresponding to the reference text to obtain a first text retrieval model. For example, the cross-entropy loss between the output text mark 2011 of the reference text 201 and the reference text mark "55648" may be determined, and the model may be exhaustively fine-tuned. The cross-entropy loss between the output text mark 2021 of the reference text 202 and the reference text mark "55942" can also be determined to fine-tune the model again.
[0046] According to an embodiment of the present disclosure, an initial text retrieval model is trained using reference text and reference text marks, and the trained first text retrieval model can output text marks based on the text, and a target reference text can be searched for from multiple reference texts based on the output text marks, which contributes to improving the model's authority and avoids the degradation of model accuracy caused by model "optical illusions."
[0047] Having described above how this disclosure trains an initial text retrieval model, we now turn to a description of how the first text retrieval model is trained.
[0048] FIG. 2B is a schematic diagram of training a first text retrieval model according to one embodiment of the present disclosure.
[0049] In some embodiments, a sample text to be queried can be input into a first text retrieval model to obtain a first output text mark. As shown in FIG. 2B , the sample text to be queried 210 can be the above sample text to be queried: "Building a steel-clad building in the west to block the sunlight from the low-rise residential area, and whether there are national legal regulations." The sample text to be queried 210 can be input into a first text retrieval model M201 to obtain an output text mark 211 of the sample text to be queried 210. The output text mark 211 can be the first output text mark.
[0050] In an embodiment of the present disclosure, the sample text to be queried may correspond to a first target reference text. For example, multiple training text pairs may be obtained. Taking legal provisions as an example, the training text pair may include a sample text to be queried and a corresponding answer legal provision. The corresponding answer legal provision may be the first target reference text. The reference text mark of the corresponding answer legal provision may be a text mark tag of the sample text to be queried. The sample text to be queried 210 may correspond to the reference text 201. The reference text 201 may be the first target reference text corresponding to the sample text to be queried 210. The reference text mark "55648" of the reference text 201 may be a text mark tag of the sample text to be queried 210.
[0051] In some implementations, a first text retrieval model may be trained based on the text mark tags of the sample text to be queried and the first output text tags.
[0052] In an embodiment of the present disclosure, training the first text retrieval model may include full-scale fine-tuning the first text retrieval model. For example, a loss value can be determined using various loss functions (e.g., the above-mentioned cross-entropy loss function) based on the text mark tag "55648" of the sample text to be queried 210 and the output text mark 211 of the sample text to be queried 210. The loss value can be used to full-scale fine-tune the first text retrieval model. After training the first text retrieval model using all of the sample text to be queried and some of the corresponding text mark tags, a second file retrieval model can be obtained. After training the model using all of the sample text to be queried and the corresponding text mark tags, the obtained model can be used as a target text retrieval model.
[0053] According to the embodiments of the present disclosure, after training a model using a plurality of reference texts and a plurality of reference text marks, the model is further trained again using a sample text to be queried and corresponding text mark tags, which effectively improves the model's ability to handle complex problems and helps the model accurately output text marks corresponding to the problems.
[0054] The above description of the present disclosure takes the example where the sample text to be queried corresponds to one first target reference text, but the present disclosure is not limited thereto, and the sample text to be queried may correspond to multiple first target reference texts, as will be described below.
[0055] In some embodiments, a reference text mark may include a first reference text submark and a second reference text submark. For example, the reference text mark "55648" may include a first reference text submark "55" and a second reference text submark "648." The reference text mark "55942" may include a first reference text submark "55" and a second reference text submark "942." As can be appreciated, the first reference text submark may be a prefix of the reference text mark.
[0056] In some embodiments, the first reference text submarks of the first target reference texts corresponding to the sample text to be queried are the same, for example, the reference text mark "55648" and the reference text mark "55942" have the same first reference text submark "55".
[0057] For ease of understanding, the present disclosure has been described above in terms of examples where the referenced text is a legal document, but the present disclosure is not limited thereto and will be described hereinafter.
[0058] In some embodiments, the reference text may be rule text, which may include provisions in files created by various authorities, such as legal provisions, policy file provisions, and standard file provisions.
[0059] The method for training a text retrieval model of the present disclosure has been described above. The text retrieval method of the present disclosure will now be described.
[0060] FIG. 3 is a flowchart of a text search method according to another embodiment of the present disclosure.
[0061] As shown in FIG. 3, the method 300 may include operations S310 to S320.
[0062] In operation S310, the text to be queried is input into the target text retrieval model, and the current output result is obtained.
[0063] In an embodiment of the present disclosure, the text to be queried may be a question entered by the user.
[0064] In an embodiment of the present disclosure, the target text retrieval model is obtained by training a first text retrieval model using a sample text to be queried and text mark tags. The first text retrieval model may be a generative artificial intelligence model. The sample text to be queried may be a complex problem or a simple problem.
[0065] In an embodiment of the present disclosure, the sample text to be queried corresponds to a first target reference text in the plurality of reference texts, and the text mark tag is a reference text mark of the first target reference text. The reference text mark may be a character string.
[0066] In the embodiment of the present disclosure, the current output result may be a string.
[0067] In operation S320, in response to determining that the current output result hits the second target reference text in the plurality of reference texts, the second target reference text is taken as a search result corresponding to the text to be queried.
[0068] In an embodiment of the present disclosure, a search can be performed among the multiple reference texts based on the current output result to determine whether the current output result hits a second target reference text. If the current output result hits one of the multiple reference texts, the hit reference text can be the second target reference text. The second target reference text can be returned to the user as a search result.
[0069] According to the embodiments of the present disclosure, a target search model can be obtained by training using sample texts to be queried and corresponding text mark tags, and the accuracy of search results for complex problems can be efficiently improved. Furthermore, by using the target text search model to determine search results, it can consume less resources and reduce the cost for non-experts to accurately search texts in specialized fields.
[0070] It can be understood that the target text retrieval model according to the above method 300 is obtained by training the first text retrieval model using the method 100. The above detailed descriptions regarding the first text retrieval model, the reference text, the reference text mark, and the sample text to be queried can be similarly applied to the first text retrieval model, the reference text, the reference text mark, and the sample text to be queried in the method 300, and the disclosure thereof will not be repeated here.
[0071] The text search method of the present disclosure has been described above. The text search method of the present disclosure will be further described below.
[0072] FIG. 4 is a schematic diagram of a text search method according to an embodiment of the present disclosure.
[0073] In an embodiment of the present disclosure, inputting the query text into the target text retrieval model and obtaining a current output result may include inputting the query text into an encoder of the target text retrieval model and obtaining an encoding result. The encoding result is input into a decoder of the target text retrieval model, and beam search decoding is performed on the encoding result to obtain multiple current output results. As shown in FIG. 4, the encoder of the target text retrieval model M403 can encode the query text 420 and obtain the encoding result. Then, the decoder of the target text retrieval model M403 can perform beam search decoding on the encoding result to obtain current output result 421 and current output result 422. If the above sample text 210 to be queried is the query text 420, the current output result 421 may be "55648," and the current output result 422 may be "55942." According to an embodiment of the present disclosure, beam search decoding is performed on the encoding result, allowing the model to output multiple results and comprehensively search for reference texts highly relevant to the query text.
[0074] In an embodiment of the present disclosure, the current output result may be an integer string, for example, the output result of the target text search model is arranged so as to reduce the likelihood probability of a non-integer string and sufficiently improve the probability of an integer string.
[0075] In an embodiment of the present disclosure, a search can be performed among multiple reference texts based on the current output result. For example, a search can be performed among multiple reference texts based on the current output result 421 and the current output result 422. It may be determined that the current output result 421 hits the first reference text, and that the current output result 422 hits the second reference text. In this case, the first reference text and the second reference text can be search results of the text 420 to be queried.
[0076] It can be understood that the present disclosure has been described above using an example in which beam search decoding is performed on the encoding result. However, the present disclosure is not limited to this, and the encoding result may be decoded using other decoding methods.
[0077] It can be understood that the present disclosure has been described above using an example in which the current output result hits one or more reference texts. Below, an example in which the current output result does not hit any reference texts will be described.
[0078] In some embodiments, the target text retrieval model may be a large-scale language model, which may introduce a degree of randomness, for example, in the encoding and / or decoding processes.
[0079] In some embodiments, the method 300 may further include, in response to determining that the current output result does not match any of the multiple reference texts, repeatedly inputting the query text into the target text retrieval model until a subsequent output result of the current output result matches a second target reference text in the multiple reference texts. For example, if the current output result of the query text does not match any of the multiple reference texts, the query text can be input into the target text retrieval model again to obtain a subsequent output result of the current output result. Then, it is determined whether the subsequent output result matches one or more of the multiple reference texts. After one or more iterations, if the subsequent output result matches a second target reference text in the multiple reference texts, the second target reference text is used as a search result for the query text. According to embodiments of the present disclosure, if the output result does not match any of the reference texts, the robustness of the model can be further improved by using the model to perform re-inference.
[0080] Having thus far been understood to have described the method of the present disclosure, the apparatus of the present disclosure will now be described.
[0081] FIG. 5 is a block diagram of an apparatus for training a text retrieval model according to one embodiment of the present disclosure.
[0082] As shown in FIG. 5, the apparatus 500 may include a first acquisition module 510 and a training module 520 .
[0083] The first obtaining module 510 inputs the sample text to be queried into the first text retrieval model and obtains a first output text mark.
[0084] The training module 520 trains a first text retrieval model based on the text mark tags of the sample text to be queried and the first output text marks to obtain a target text retrieval model.
[0085] In an embodiment of the present disclosure, the sample text to be queried corresponds to a first target reference text in the plurality of reference texts, the text mark tag is a reference text mark of the first target reference text, the plurality of reference text marks are determined based on the reference text sequence, and the reference text sequence is determined based on the vocabulary of the plurality of reference texts.
[0086] In an embodiment of the present disclosure, the first text retrieval model is obtained by training an initial text retrieval model using a plurality of reference texts and a plurality of corresponding reference text marks, thereby outputting reference text marks corresponding to the reference texts.
[0087] In some embodiments, the first text retrieval model is obtained by training an initial text retrieval model using a plurality of reference texts and a corresponding plurality of reference text marks by the following modules: a second obtaining module, which inputs the reference texts into the initial text retrieval model and obtains second output text marks of the reference texts; and a first fine-tuning module, which fully fine-tunes the initial text retrieval model based on the second output text marks and the reference text marks to obtain the first text retrieval model.
[0088] In some embodiments, the first fine-tuning module includes a first determining sub-module for determining a cross-entropy loss based on the second output text mark and the reference text mark, and a first fine-tuning sub-module for fully fine-tuning the initial text retrieval model based on the cross-entropy loss.
[0089] In some embodiments, the multiple reference texts are from N reference text sets, each reference text set including at least one reference text, where N is an integer greater than 1. The reference text sequence is determined based on the vocabulary of the multiple reference texts by the following modules: a first clustering processing module, which performs a clustering process on the N reference text sets respectively, and adjusts the order of the reference texts in the N reference text sets to obtain N processed text sets; a third obtaining module, which obtains a fused text set based on the N processed text sets; and a second clustering processing module, which performs a clustering process on the fused text set based on the vocabulary of each of the multiple reference texts, and adjusts the order of the multiple reference texts in the fused text set to obtain the reference text sequence.
[0090] In some embodiments, the sample text to be queried corresponds to a plurality of first target reference texts, the reference text marks include first reference text submarks and second reference text submarks, and the first reference text submarks of the plurality of first target reference texts are the same.
[0091] In some embodiments, the training module includes a second fine-tuning sub-module for thoroughly fine-tuning the first text retrieval model.
[0092] In some embodiments, the reference text is a rule text and the first text retrieval model is a large-scale language model.
[0093] It can be understood that the training device of the present disclosure has been described above, and the text search device of the present disclosure will now be described.
[0094] FIG. 6 is a block diagram of a text search device according to another embodiment of the present disclosure.
[0095] As shown in FIG. 6, the apparatus 600 may include a fourth acquisition module 610 and a determination module 620 .
[0096] The fourth acquisition module 610 inputs the text to be queried into the target text search model and obtains the current output result.
[0097] In response to determining that the current output result hits the second target reference text in the plurality of reference texts, the determination module 620 determines the second target reference text as a search result corresponding to the text to be queried.
[0098] In an embodiment of the present disclosure, the target text retrieval model is obtained by training a first text retrieval model using a sample text to be queried and a text mark tag, where the sample text to be queried corresponds to a first target reference text in a plurality of reference texts, and the text mark tag is a reference text mark of the first target reference text.
[0099] In an embodiment of the present disclosure, the target text retrieval model may be trained by the device 500, for example.
[0100] In some embodiments, the target text retrieval model is a large-scale language model. In response to determining that the current output result does not hit any of the plurality of reference texts, the apparatus further includes an execution module that repeatedly performs an operation of inputting the text to be queried into the target text retrieval model until an output result after the current output result hits a second target reference text in the plurality of reference texts.
[0101] In the technical solution disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure and other processing of such user personal information shall all comply with the provisions of relevant laws and shall not violate public order and morals.
[0102] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program.
[0103] 7 illustrates an exemplary block diagram for implementing an example electronic device 700 according to an embodiment of the present disclosure. The electronic device is intended to represent various types of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The electronic device may also represent various types of mobile devices, such as personal digital assistants, mobile phones, smartphones, wearable devices, and other similar computing devices. The components, their connections and relationships, and their functions illustrated herein are exemplary only and do not limit the implementation of the present disclosure as described and / or claimed herein.
[0104] 7, the device 700 includes a computing unit 701, which can perform various appropriate operations and processes based on a computer program stored in a read-only memory (ROM) 702 or loaded from a storage unit 708 into a random access memory (RAM) 703. The RAM 703 can further store various programs and data necessary for the operation of the device 700. The computing unit 701, the ROM 702, and the RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0105] Several components in the device 700 are connected to an I / O interface 705, including an input unit 706, such as a keyboard, a mouse, etc., an output unit 707, such as various types of displays, speakers, etc., a storage unit 708, such as a magnetic disk, an optical disk, etc., and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices via a computer network, such as the Internet, and / or various telecommunication networks.
[0106] The computing unit 701 may be any of a variety of general-purpose and / or specialized processing modules having processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units for executing machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the methods and processes described above, such as the text retrieval model training method and / or the text retrieval method. For example, in some embodiments, the text retrieval model training method and / or the text retrieval method may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 708. In some embodiments, some or all of the computer program may be loaded and / or installed into the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into RAM 703 and executed by the computing unit 701, it may perform one or more steps of the above-described method for training a text retrieval model and / or the text retrieval method. Alternatively, in another embodiment, the computing unit 701 may be configured to perform the method for training a text retrieval model and / or the text retrieval method in any other suitable form (e.g., via firmware).
[0107] Various embodiments of the systems and techniques described herein may be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may be embodied in one or more computer programs that can be executed and / or interpreted by a programmable system that includes at least one programmable processor, which may be a special purpose or general purpose programmable processor, and that can receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0108] Program codes for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, so that when the program code is executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are performed. The program code may be executed entirely on a device, partially on a device, partially on a device as a separate software package and partially on a remote device, or entirely on a remote device or server.
[0109] In the context of this disclosure, a machine-readable medium may be a tangible medium, and may contain or store a program for use in or in connection with an instruction execution system, device, or electronic device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or electronic device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include an electrical connection of one or more wires, a portable computer diskette, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM) or flash memory, an optical fiber, a compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0110] To provide interaction with a user, a computer may implement the systems and techniques described herein and include a display device (e.g., a cathode ray tube (CRT) display or a liquid crystal display (LCD)) for displaying information to a user, and a keyboard and pointing device (e.g., a mouse or trackball) through which a user can provide input to the computer. Other types of devices may also provide interaction with a user; for example, the feedback provided to the user may be any form of sensing feedback (e.g., visual feedback, auditory feedback, or tactile feedback) and may receive input from the user in any form (including voice input, speech input, or tactile input).
[0111] The systems and techniques described herein can be implemented in a computing system that includes background components (e.g., a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with embodiments of the systems and techniques described herein), or a computing system that includes any combination of such background, middleware, or front-end components. The components of the system can be connected to each other by any form or medium of digital data communication (e.g., a communications network). Examples of communications networks include, by way of example, a local area network (LAN), a wide area network (WAN), and the Internet.
[0112] The computer system may include clients and servers. Clients and servers are generally remote and typically interact through a communication network. The relationship of client and server is created by computer programs running on the corresponding computers and having the client-server relationship.
[0113] It should be understood that various types of flows shown above may be used, and steps may be rearranged, added, or deleted. For example, the steps described in the present invention may be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present disclosure can be achieved, and the present specification is not limited thereto.
[0114] The specific embodiments described above do not limit the scope of protection of the present disclosure. Those skilled in the art should understand that various modifications, combinations, subcombinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present disclosure should be included within the scope of protection of the present disclosure.
Claims
1. A method for training a text retrieval model by a processor, comprising: inputting a sample text to be queried into a first text retrieval model and obtaining a first output text mark; training the first text retrieval model based on the text mark tags of the sample text to be queried and the first output text marks to obtain a target text retrieval model; wherein the sample text to be queried corresponds to a first target reference text in a plurality of reference texts, the text mark tag is a reference text mark of the first target reference text, the plurality of reference text marks are determined based on a reference text sequence, and the reference text sequence is determined based on a vocabulary of the plurality of reference texts; The first text retrieval model is obtained by training an initial text retrieval model using a plurality of the reference texts and a corresponding plurality of the reference text marks, thereby outputting the reference text marks corresponding to the reference texts. How to train a text search model.
2. The first text retrieval model is obtained by training an initial text retrieval model using a plurality of reference texts and a corresponding plurality of reference text marks by the following operations: The operation is inputting the reference text into an initial text retrieval model to obtain a second output text mark of the reference text; and fine-tuning the initial text retrieval model based on the second output text mark and the reference text mark to obtain the first text retrieval model. The method of claim 1.
3. and fine-tuning the initial text retrieval model based on the second output text marks and the reference text marks, determining a cross-entropy loss based on the second output text mark and the reference text mark; and fine-tuning the initial text retrieval model based on the cross-entropy loss. The method of claim 2.
4. the plurality of reference texts are from N reference text sets, each reference text set including at least one of the reference texts, N being an integer greater than 1; The reference text sequence is determined based on a vocabulary of a plurality of the reference texts by: The operation is Perform a clustering process on each of the N reference text sets, adjust the order of the reference texts in the N reference text sets, and obtain N processed text sets; Obtaining a fused text set based on the N processed text sets; performing a clustering process on the fused text set based on the vocabulary of each of the plurality of reference texts, and adjusting the order of the plurality of reference texts in the fused text set to obtain the reference text sequence; The method of claim 1.
5. The sample text to be queried corresponds to a plurality of the first target reference texts, and the reference text marks include a first reference text submark and a second reference text submark, and the first reference text submarks of the plurality of first target reference texts are the same. The method of claim 1.
6. Training the first text retrieval model comprises: and fine-tuning the first text retrieval model. The method of claim 1.
7. The reference text is a rule text and the first text retrieval model is a large-scale language model. The method of claim 1.
8. A method of text searching by a processor, comprising: Inputting a text to be queried into a target text retrieval model and obtaining a current output result; in response to determining that the current output result is a hit for a second target reference text in the plurality of reference texts, determining the second target reference text as a search result corresponding to the text to be queried; Wherein, the target text retrieval model is obtained by training a first text retrieval model using a sample text to be queried and a text mark tag, the sample text to be queried corresponds to a first target reference text in a plurality of reference texts, and the text mark tag is a reference text mark of the first target reference text. Text search methods.
9. the target text retrieval model is a large-scale language model; The method comprises: In response to determining that the current output result does not hit any of the plurality of reference texts, the method further includes repeatedly performing an operation of inputting the text to be queried into a target text retrieval model until a post-output result of the current output result hits a second target reference text in the plurality of reference texts. The method of claim 8.
10. a first obtaining module for inputting a sample text to be queried into a first text retrieval model and obtaining a first output text mark; a training module for training the first text retrieval model based on the text mark tags of the sample text to be queried and the first output text marks to obtain a target text retrieval model; wherein the sample text to be queried corresponds to a first target reference text in a plurality of reference texts, the text mark tag is a reference text mark of the first target reference text, the plurality of reference text marks are determined based on a reference text sequence, and the reference text sequence is determined based on a vocabulary of the plurality of reference texts; The first text retrieval model is obtained by training an initial text retrieval model using a plurality of the reference texts and a corresponding plurality of the reference text marks, thereby outputting the reference text marks corresponding to the reference texts. A training device for text retrieval models.
11. The first text retrieval model is a second obtaining module for inputting the reference text into an initial text retrieval model and obtaining a second output text mark of the reference text; a first fine-tuning module for fine-tuning the initial text retrieval model based on the second output text mark and the reference text mark to obtain the first text retrieval model, obtained by training an initial text retrieval model using a plurality of reference texts and a corresponding plurality of reference text marks.
11. The apparatus of claim 10.
12. The first fine-tuning module: a first determining sub-module for determining a cross-entropy loss based on the second output text mark and the reference text mark; a first fine-tuning sub-module that fine-tunes the initial text retrieval model based on the cross-entropy loss.
12. The apparatus of claim 11.
13. the plurality of reference texts are from N reference text sets, each reference text set including at least one of the reference texts, N being an integer greater than 1; The reference text sequence is a first clustering processing module for performing clustering processing on each of the N reference text sets, adjusting the order of the reference texts in the N reference text sets, and obtaining N processed text sets; a third acquisition module for acquiring a fused text set based on the N processed text sets; a second clustering processing module that performs a clustering process on the fused text set based on the vocabulary of each of the plurality of reference texts, and adjusts the order of the plurality of reference texts in the fused text set to obtain the reference text sequence, determined based on the vocabulary of the plurality of reference texts 11. The apparatus of claim 10.
14. The training module comprises: a second fine-tuning submodule for fully fine-tuning the first text retrieval model; 11. The apparatus of claim 10.
15. The reference text is a rule text and the first text retrieval model is a large-scale language model.
11. The apparatus of claim 10.
16. a fourth obtaining module for inputting the query text into the target text retrieval model to obtain a current output result; a determining module, in response to determining that the current output result hits a second target reference text in a plurality of reference texts, determining that the second target reference text is a search result corresponding to the text to be queried; wherein the target text retrieval model is obtained by training a first text retrieval model using a sample text to be queried and a text mark tag, the sample text to be queried corresponds to a first target reference text in a plurality of reference texts, and the text mark tag is a reference text mark of the first target reference text; Text search device.
17. the target text retrieval model is a large-scale language model; The device comprises: and an execution module that, in response to determining that the current output result does not match any of the plurality of reference texts, repeatedly executes an operation of inputting the text to be queried into a target text retrieval model until a subsequent output result of the current output result matches a second target reference text in the plurality of reference texts.
17. The apparatus of claim 16.
18. at least one processor; a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor such that the at least one processor can perform the method of any one of claims 1 to 7. electronic equipment.
19. A non-transitory computer-readable storage medium having computer instructions stored thereon, comprising: The computer instructions cause a computer to carry out the method of any one of claims 1 to 7. A non-transitory computer-readable storage medium.
20. Implementing the method of any one of claims 1 to 7 when executed by a processor Computer program.
Citation Information
Patent Citations
Retrieval method, device, electronic apparatus and storage medium
JP2022046759A
Method and apparatus for training retrieval model, device, computer storage medium, and computer program
JP2022054389A
Program, method, information processing apparatus, and system
JP2022086279A
Data management apparatus, data management method, and program
JP2023120862A