Semantic recall method and device based on large language model and electronic equipment
By decomposing the problem to be queried into sub-problems and performing semantic coding fusion, and generating query semantic vectors, the problem of the deviation and insufficient performance of the information recall results in the prior art depend on language model coding is solved, and a more accurate and comprehensive information recall is achieved.
Patent Information
- Application Number
- CN202510251261.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-24
AI Technical Summary
In the prior art, there are problems in the information recall in which the recall results depend on the language model's deviation and insufficient performance of the problem encoding, resulting in insufficient accuracy and coverage of the recall results.
By using the problem decomposition prompt word to decompose the problem to be queried into multiple sub-problems, the embedding model is used to semantic code the query problem and sub-problems, the query semantic vector is fused, and the query semantic vector is recalled based on the similarity between the query semantic vector and the semantic vector of knowledge in the knowledge database.
Improve the accuracy and coverage of information recalls, avoiding the instability of recall results due to biased language model understanding and insufficient performance.
Smart Images

Figure CN120196740A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a semantic recall method, device and electronic equipment based on a large language model. Background Art
[0002] In the field of natural language processing, a variety of innovative technical frameworks and models have emerged in recent years, bringing significant progress to tasks such as text processing, semantic understanding, and information retrieval. Among them, the Retrieval-Augmented Generation (RAG) framework provides knowledge guidance for large language models through an external knowledge base, effectively alleviating the problem of difficulty in updating model parameters and ensuring the accuracy, timeliness, and completeness of answers. Figure 4 The semantic similarity calculation model shown in the figure is a dual-tower architecture (Deep Structured Semantic Model, DSSM). It uses an encoder to vectorize the text and calculate the cosine value of the vector angle to quickly determine the semantic similarity of the text. In the field of data retrieval, the AutoFAS framework uses a multi-stage serial form to obtain retrieval results, which effectively improves the efficiency and accuracy of large-scale retrieval. In the initial stage, the framework uses the efficiency of the dual-tower model to obtain preliminary recall results from a large-scale knowledge base, and then uses the accuracy of the single-tower model to perform multi-stage sorting on the recall results, and finally obtains accurate retrieval results and their sorting.
[0003] However, despite the remarkable achievements of these technologies, there are still some deficiencies in existing technologies. The recall results of the dual-tower recall technology are highly dependent on the encoding results of the language model for the question. The deviation of the question and the insufficient performance of the language model will significantly affect the accuracy and coverage of the recall results. When applied to the large language model RAG module, this deficiency will seriously weaken the knowledge enhancement effect and reduce the credibility and authority of the answer.
[0004] Therefore, how to efficiently and accurately recall information based on retrieval questions is a technical problem that needs to be solved urgently. Summary of the invention
[0005] In view of the above problems existing in the prior art, the present invention provides a semantic recall method, device and electronic device based on a large language model, so as to efficiently and accurately recall information according to retrieval questions.
[0006] The present invention provides a semantic recall method based on a large language model, comprising the following steps.
[0007] Using a problem decomposition prompt to guide a large language model to decompose the problem to be queried into multiple sub-problems; using an embedding model to semantically encode the problem to be queried to obtain a semantic vector of the problem to be queried; using the embedding model to semantically encode each of the multiple sub-problems respectively to obtain a semantic vector of each sub-problem; fusing the semantic vector of the problem to be queried and the semantic vectors of each sub-problem to obtain a query semantic vector; recalling a retrieval result of the problem to be queried according to the similarity between the query semantic vector and the semantic vectors of the knowledge in the knowledge database.
[0008] According to a semantic recall method based on a large language model provided by the present invention, the fusing the semantic vector of the problem to be queried and the semantic vectors of each sub-problem to obtain a query semantic vector includes: performing a weighted linear operation on the semantic vector of the problem to be queried and the semantic vectors of each sub-problem to obtain the query semantic vector.
[0009] According to a semantic recall method based on a large language model provided by the present invention, the method further includes: splitting the retrieval result into clauses to obtain multiple sub-results included in the retrieval result; splitting the reference answer for this retrieval into clauses to obtain multiple sub-answers included in the reference answer; using the embedding model to process the multiple sub-results respectively to obtain a semantic vector of each sub-result in the multiple sub-results; using the embedding model to process the multiple sub-answers respectively to obtain a semantic vector of each sub-answer in the multiple sub-answers; for the semantic vector of each sub-answer, determining the similarity between it and the semantic vector of each sub-result respectively to obtain multiple similarities corresponding to the semantic vector of each sub-answer; determining an evaluation score of the retrieval result according to the multiple similarities corresponding to the semantic vector of each sub-answer.
[0010] According to a semantic recall method based on a large language model provided by the present invention, the determining an evaluation score of the retrieval result according to the multiple similarities corresponding to the semantic vector of each sub-answer includes: for each sub-answer, taking the maximum similarity among the multiple similarities corresponding to its semantic vector as the coverage score of the sub-answer; in the case where the coverage score is greater than a preset score threshold, determining that the sub-answer is covered in this retrieval; determining the ratio of the number of the sub-answers covered in this retrieval to the total number of the multiple sub-answers to obtain the evaluation score of the retrieval result.
[0011] A semantic recall method based on a large language model provided by the present invention, wherein the embedding model is trained in the following manner: obtaining a training data set; wherein, the training data set includes problem texts as sample data and correct retrieval results as labels of the sample data; using problem decomposition prompt words to guide the large language model to decompose the problem text into multiple sub-problems; using an initial embedding model to perform semantic encoding on the problem text to obtain a semantic vector of the problem text; using the initial embedding model to perform semantic encoding on each of the multiple sub-problems respectively to obtain a semantic vector of each sub-problem; fusing the semantic vector of the problem text with the semantic vector of each sub-problem to obtain a training semantic vector; obtaining a predicted retrieval result of the problem text according to the similarity between the training semantic vector and the semantic vector of knowledge in a knowledge database; using a loss function to determine the difference between the correct retrieval result corresponding to the problem text and the predicted retrieval result; using a preset optimization algorithm to sequentially adjust the parameters of the initial embedding model to reduce the difference until a training end condition is reached, thereby obtaining the embedding model.
[0012] A semantic recall method based on a large language model provided by the present invention, wherein fusing the semantic vector of the problem text with the semantic vector of each sub-problem to obtain a training semantic vector includes: performing a weighted linear operation on the semantic vector of the problem text and the semantic vector of each sub-problem to obtain the training semantic vector.
[0013] The present invention also provides a semantic recall device based on a large language model, including the following modules: A first decomposition module, configured to use problem decomposition prompt words to guide the large language model to decompose a problem to be queried into multiple sub-problems; a second decomposition module, configured to use an embedding model to perform semantic encoding on the problem to be queried to obtain a semantic vector of the problem to be queried; an encoding module, configured to use the embedding model to perform semantic encoding on each of the multiple sub-problems respectively to obtain a semantic vector of each sub-problem; a fusion module, configured to fuse the semantic vector of the problem to be queried and the semantic vector of each sub-problem to obtain a query semantic vector; a recall module, configured to recall a retrieval result of the problem to be queried according to the similarity between the query semantic vector and the semantic vector of knowledge in a knowledge database.
[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the semantic recall method based on a large language model as described in any one of the above is implemented.
[0015] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the semantic recall method based on a large language model as described in any one of the above is implemented.
[0016] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the semantic recall method based on a large language model as described in any one of the above is implemented.
[0017] The semantic recall method based on a large language model provided by the present invention utilizes the powerful semantic understanding ability of the large language model to conduct a more comprehensive analysis of the query problem, and obtains multiple sub-problems that can fully reflect the details and hidden intentions of the query problem. The semantic vector of the query problem and the semantic vectors of each sub-problem are fused to obtain a query semantic vector that can fully reflect the complete and deep semantic meaning of the query problem. Furthermore, according to the similarity between the query semantic vector and the semantic vectors of the knowledge in the knowledge database, retrieval results with better coverage and completeness for the query problem can be recalled. Therefore, the semantic recall method provided by the present invention can avoid the influence of performance fluctuations caused by the deviation of the embedding model's understanding of the query problem and the insufficient performance of the embedding model on the accuracy and coverage of the recall results. Thus, information can be recalled efficiently and accurately according to the retrieval problem. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0019] Figure 1 It is a flowchart of the semantic recall method based on a large language model provided by the present invention.
[0020] Figure 2 It is a flowchart of the evaluation method for retrieval results provided by the present invention.
[0021] Figure 3 It is a flowchart of the training method for the embedding model provided by the present invention.
[0022] Figure 4 It is a structural schematic diagram of the DSSM two-tower model in the prior art.
[0023] Figure 5 It is a structural schematic diagram of the multi-channel two-tower recall architecture provided by the present invention.
[0024] Figure 6It is a schematic structural diagram of the semantic recall device based on the large language model provided by the present invention.
[0025] Figure 7 It is a schematic structural diagram of the electronic device provided by the present invention. Specific embodiments
[0026] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts fall within the scope of protection of the present invention.
[0027] The following is combined with Figures 1 - 5 Describe the semantic recall method based on the large language model of the present invention.
[0028] Figure 1 It is a schematic flow diagram of the semantic recall method based on the large language model provided by the present invention. As Figure 1 shown, the method includes the following steps.
[0029] Step 101, use the problem decomposition prompt word to guide the large language model to decompose the problem to be queried into multiple sub-problems.
[0030] The semantic recall method provided by the present invention can be applied to a variety of target tasks, such as an intelligent customer service system processing user consultations, or a medical intelligent agent answering medical questions, etc. For different target tasks, the content of the problem to be queried is also different. For example, the problem to be queried is a medical problem proposed by the user, etc.
[0031] The problem decomposition prompt word is used to guide the large language model to decompose the problem to be queried into multiple sub-problems.
[0032] The sub-problem is multiple problems obtained by decomposing the problem to be queried into smaller, more specific, easier-to-process or answer problems. These problems are usually part of the original problem to be queried or a refinement of a certain aspect.
[0033] The problem decomposition prompt word is a prompt word used to guide the large language model to decompose the problem to be queried into multiple sub-problems.
[0034] Only as an example, the problem decomposition prompt word is: "Please decompose the following problem into 5 sub-problems according to semantics; 'the problem to be queried'".
[0035] The large language model can be an open-source large language model, and can be a large language model after continued training, which is not limited by the description in this specification.
[0036] In the specific implementation process, the output of the large language model can be parsed to obtain multiple decomposed sub-questions.
[0037] Step 102: Use an embedding model to perform semantic encoding on the question to be queried, obtaining the semantic vector of the question to be queried.
[0038] The core idea of the embedding model is to map high-dimensional data to a low-dimensional embedding space while preserving the features and semantic information of the original data.
[0039] In the specific implementation process, multiple embedding models can be used to perform semantic encoding on the question to be queried. For example, the BGE (BAAI General Embedding) model, BERT (Bidirectional Encoder Representations from Transformers), CBOW (Continuous Bag of Words), etc.
[0040] In the specific implementation process, the question to be queried can be input into the embedding model, and the embedding model outputs the semantic vector of the question to be queried.
[0041] Step 103: Use the embedding model to perform semantic encoding on each of the multiple sub-questions respectively, obtaining the semantic vector of each sub-question.
[0042] In the specific implementation process, each sub-question can be input into the embedding model respectively, and the embedding model outputs the semantic vector of each sub-question.
[0043] Step 104: Fuse the semantic vector of the question to be queried and the semantic vectors of each sub-question to obtain the query semantic vector.
[0044] In the specific implementation process, various methods can be used to fuse the semantic vector of the question to be queried and the semantic vectors of each sub-question to obtain the query semantic vector, which is not limited by the description in this specification.
[0045] In some embodiments, different weights can be assigned to the semantic vector of the question to be queried and the semantic vectors of each sub-question, and weighted linear operations can be performed to obtain the query semantic vector.
[0046] Only as an example, the large language model decomposes the question to be queried into 5 sub-questions. Subsequently, the BGE model is used to encode the question to be queried and the decomposed sub-questions respectively. A weight of 0.5 is assigned to the semantic vector of the question to be queried, and weights of 0.1 are assigned to the semantic vectors of the 5 sub-questions. Linear weighting is performed on all the semantic vectors to obtain the query semantic vector.
[0047] Step 105: Recall the search results of the query question based on the similarity between the query semantic vector and the semantic vector of the knowledge in the knowledge database.
[0048] In a specific implementation process, a variety of methods can be used to determine the similarity between the query semantic vector and the semantic vector of the knowledge in the knowledge database, for example, by calculating cosine similarity, Euclidean distance, and the like.
[0049] In the specific implementation process, Figure 5 As shown in the figure, the knowledge in the knowledge database can also be represented as a semantic vector. After obtaining the query semantic vector and the semantic vector of the knowledge, the semantic relevance is evaluated by calculating the similarity between them. The higher the similarity value, the stronger the semantic relevance between the query question and the knowledge. According to the similarity value, the knowledge in the knowledge database is sorted. The knowledge with the highest similarity is recalled first as the retrieval result of the query question.
[0050] In the existing retrieval technology, the evaluation system for retrieval results is not perfect. The main problem is that positive samples rely on manual annotation, and in actual scenarios with huge amounts of knowledge and high repetition rates, it is easy to generate erroneous negative samples, resulting in deviations in the evaluation results. In order to solve the above problems, the present invention provides the following retrieval result evaluation method for accurately and reasonably evaluating the recalled retrieval results.
[0051] Figure 2 is a flow chart of the search result evaluation method provided by the present invention, such as Figure 2 As shown, the method includes the following steps.
[0052] Step 201: Segment the search result to obtain multiple sub-results contained in the search result.
[0053] Step 202: Sentence the reference answer for this search to obtain multiple sub-answers contained in the reference answer.
[0054] In the specific implementation process, the search results and reference answers can be cleaned separately to remove the garbled characters and redundant characters that may exist in the search results and reference answers, and reduce the interference with subsequent evaluation. For example, the isalnum function can be used for text cleaning. Then, the search results and reference answers are separated into sentences using punctuation marks (for example, commas) to obtain multiple sub-results contained in the search results and multiple sub-answers contained in the reference answers. The sub-results and sub-answers obtained by the sentence separation can be used to refine the granularity of the evaluation, prevent misjudgment caused by minor differences in details, and effectively improve the accuracy of the evaluation.
[0055] Step 203: Use the embedding model to process the multiple sub-results separately to obtain a semantic vector of each of the multiple sub-results.
[0056] Step 204: Use the embedding model to process multiple sub-answers respectively to obtain the semantic vectors of each sub-answer among the multiple sub-answers.
[0057] As an example only, the BGE model can be used to process multiple sub-results respectively to obtain the semantic vectors of each sub-result among the multiple sub-results. And use the BGE model to process multiple sub-answers respectively to obtain the semantic vectors of each sub-answer among the multiple sub-answers.
[0058] Step 205: For the semantic vector of each sub-answer, determine the similarity between it and the semantic vector of each sub-result respectively to obtain multiple similarities corresponding to the semantic vector of each sub-answer.
[0059] As an example only, the cosine similarity between the semantic vector of each sub-answer and the semantic vector of each sub-result can be calculated in the form of a Cartesian product to obtain multiple similarities corresponding to the semantic vector of each sub-answer.
[0060] Step 206: Determine the evaluation score of the retrieval result according to the multiple similarities corresponding to the semantic vector of each sub-answer.
[0061] In some embodiments, for each sub-answer, the maximum similarity among the multiple similarities corresponding to its semantic vector can be used as the coverage score of the sub-answer; in the case where the coverage score is greater than a preset score threshold (for example, 0.9), it is determined that the sub-answer is covered in this retrieval; determine the ratio of the number of sub-answers covered in this retrieval to the total number of multiple sub-answers to obtain the evaluation score of the retrieval result.
[0062] The above evaluation method for retrieval results provided by the present invention can realize the automatic evaluation of retrieval results. Since the coverage of the retrieval results to the reference answers is evaluated at the semantic level of clauses, compared with the prior art, the matching granularity is refined to clauses, effectively improving the accuracy of the evaluation.
[0063] Figure 3 It is a schematic flowchart of the training method of the embedding model provided by the present invention, as Figure 3 shown, the method includes the following steps.
[0064] Step 301: Obtain a training data set; wherein, the training data set includes the question text as sample data and the correct retrieval result as the label of the sample data.
[0065] In the specific implementation process, the training data set can be obtained in various ways, which is not limited by the description of this specification.
[0066] Step 302: Use the question decomposition prompt words to guide the large language model to decompose the question text into multiple sub-questions.
[0067] Step 303: Use the initial embedding model to perform semantic encoding on the question text to obtain the semantic vector of the question text.
[0068] Step 304: Use the initial embedding model to perform semantic encoding on each of the multiple sub-questions respectively to obtain the semantic vector of each sub-question.
[0069] Step 305: Fuse the semantic vector of the question text with the semantic vector of each sub-question to obtain the training semantic vector.
[0070] In some embodiments, a weighted linear operation can be performed on the semantic vector of the question text and the semantic vector of each sub-question to obtain the training semantic vector.
[0071] Step 306: Obtain the predicted retrieval result of the question text according to the similarity between the training semantic vector and the semantic vector of the knowledge in the knowledge database.
[0072] For the detailed description of steps 302 to 306, refer to the relevant content in steps 101 to 105, which will not be elaborated here.
[0073] Step 307: Use the loss function to determine the difference between the correct retrieval result and the predicted retrieval result corresponding to the question text.
[0074] Step 308: Use the preset optimization algorithm to sequentially adjust the parameters of the initial embedding model to reduce the difference until the training end condition is reached to obtain the embedding model.
[0075] In the specific implementation process, use the loss function, for example, Cross-Entropy Loss, to determine the difference between the correct retrieval result and the predicted retrieval result corresponding to the question text; use the preset optimization algorithm, for example, the gradient descent optimization algorithm, to adjust the parameters of the initial embedding model to reduce the value of the loss function until the training end condition is reached to obtain the embedding model. The training end condition can include conditions such as the convergence of the loss function and reaching the preset number of training times.
[0076] In the embodiments provided by the present invention, by fusing the semantic vector of the question text with the semantic vector of each sub-question to obtain the training semantic vector, and using the training semantic vector to train the initial embedding model, an embedding model that can comprehensively and accurately understand the semantics of the retrieval question can be trained, effectively stimulating the retrieval potential of the embedding model.
[0077] The semantic recall device based on the large language model provided by the present invention will be described below. The semantic recall device based on the large language model described below can be correspondingly referred to the semantic recall method based on the large language model described above.
[0078] Figure 6 It is a schematic structural diagram of the semantic recall device based on the large language model provided by the present invention. As Figure 6 shown, the device 600 includes the following modules.
[0079] The first decomposition module 610 is used to guide the large language model to decompose the problem to be queried into multiple sub-problems by using the problem decomposition prompt words.
[0080] The second decomposition module 620 is used to perform semantic encoding on the problem to be queried by using the embedding model to obtain the semantic vector of the problem to be queried.
[0081] The encoding module 630 is used to perform semantic encoding on each of the multiple sub-problems by using the embedding model to obtain the semantic vector of each sub-problem.
[0082] The fusion module 640 is used to fuse the semantic vector of the problem to be queried and the semantic vectors of each sub-problem to obtain the query semantic vector.
[0083] The recall module 650 is used to recall the retrieval result of the problem to be queried according to the similarity between the query semantic vector and the semantic vectors of the knowledge in the knowledge database.
[0084] Figure 7 Illustrates a schematic structural diagram of an electronic device, such as Figure 7As shown, the electronic device may include: a processor 710, a communications interface 720, a memory 730, and a communication bus 740. Among them, the processor 710, the communications interface 720, and the memory 730 complete communication with each other through the communication bus 740. The processor 710 may call the logical instructions in the memory 730 to execute a semantic recall method based on a large language model. The method includes: using a problem decomposition prompt to guide the large language model to decompose the problem to be queried into multiple sub-problems; using an embedding model to perform semantic encoding on the problem to be queried to obtain a semantic vector of the problem to be queried; using the embedding model to perform semantic encoding on each of the multiple sub-problems respectively to obtain a semantic vector of each sub-problem; fusing the semantic vector of the problem to be queried and the semantic vectors of each sub-problem to obtain a query semantic vector; recalling a retrieval result of the problem to be queried according to the similarity between the query semantic vector and the semantic vectors of the knowledge in the knowledge database.
[0085] In addition, when the logical instructions in the above-mentioned memory 730 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0086] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the semantic recall method based on a large language model provided by the above-mentioned various methods. The method includes: using a problem decomposition prompt to guide the large language model to decompose the query problem into multiple sub-problems; using an embedding model to perform semantic encoding on the query problem to obtain a semantic vector of the query problem; using the embedding model to perform semantic encoding on each of the multiple sub-problems respectively to obtain a semantic vector of each sub-problem; fusing the semantic vector of the query problem and the semantic vectors of each sub-problem to obtain a query semantic vector; and recalling a retrieval result of the query problem according to the similarity between the query semantic vector and the semantic vectors of the knowledge in the knowledge database.
[0087] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the semantic recall method based on a large language model provided by the above-mentioned various methods. The method includes: using a problem decomposition prompt to guide the large language model to decompose the query problem into multiple sub-problems; using an embedding model to perform semantic encoding on the query problem to obtain a semantic vector of the query problem; using the embedding model to perform semantic encoding on each of the multiple sub-problems respectively to obtain a semantic vector of each sub-problem; fusing the semantic vector of the query problem and the semantic vectors of each sub-problem to obtain a query semantic vector; and recalling a retrieval result of the query problem according to the similarity between the query semantic vector and the semantic vectors of the knowledge in the knowledge database.
[0088] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0089] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0090] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A semantic recall method based on large language models, characterized in that, include: Use question decomposition prompt words to guide the large language model to decompose the query question into multiple sub-questions; Using the embedding model, semantic encoding is performed on the question to be queried to obtain a semantic vector of the question to be queried; Using the embedding model, semantic encoding is performed on each of the multiple sub-problems to obtain a semantic vector of each sub-problem; The semantic vector of the question to be queried and the semantic vector of each sub-question are merged to obtain a query semantic vector; According to the similarity between the query semantic vector and the semantic vector of the knowledge in the knowledge database, the retrieval result of the question to be queried is recalled.
2. The semantic recall method based on a large language model according to claim 1, wherein The step of fusing the semantic vector of the question to be queried and the semantic vector of each sub-question to obtain a query semantic vector includes: A weighted linear operation is performed on the semantic vector of the question to be queried and the semantic vector of each sub-question to obtain the query semantic vector.
3. The semantic recall method based on a large language model according to claim 1 or 2, characterized in that, The method further comprises: Segment the search result to obtain multiple sub-results contained in the search result; The reference answer of this search is divided into sentences to obtain multiple sub-answers contained in the reference answer; Using the embedding model, the plurality of sub-results are processed respectively to obtain a semantic vector of each of the plurality of sub-results; Using the embedding model, the multiple sub-answers are processed respectively to obtain a semantic vector of each of the multiple sub-answers; For each of the semantic vectors of the sub-answer, determine the similarity between the semantic vector and the semantic vector of each of the sub-results, and obtain multiple similarities corresponding to the semantic vector of each of the sub-answers; The evaluation score of the search result is determined according to the multiple similarities corresponding to the semantic vector of each sub-answer.
4. The semantic recall method based on the large language model according to claim 3, wherein Determining the evaluation score of the search result according to the multiple similarities corresponding to the semantic vector of each sub-answer includes: For each of the sub-answers, the maximum similarity among the multiple similarities corresponding to the semantic vector is used as the coverage score of the sub-answer; When the coverage score is greater than a preset score threshold, determining that the sub-answer is covered in this search; The ratio of the number of the sub-answers covered in this search to the total number of the multiple sub-answers is determined to obtain an evaluation score of the search result.
5. The semantic recall method based on a large language model according to claim 1, wherein The embedding model is trained in the following way: Acquire a training data set; wherein the training data set includes a question text as sample data and a correct retrieval result as a label of the sample data; Using question decomposition prompt words, guiding the large language model to decompose the question text into multiple sub-questions; Using the initial embedding model, semantic encoding is performed on the question text to obtain a semantic vector of the question text; Using the initial embedding model, semantic encoding is performed on each of the multiple sub-problems to obtain a semantic vector of each sub-problem; The semantic vector of the question text is merged with the semantic vector of each sub-question to obtain a training semantic vector; Obtain the predicted retrieval result of the problem text according to the similarity between the trained semantic vector and the semantic vectors of the knowledge in the knowledge database; Use a loss function to determine the difference between the correct retrieval result corresponding to the problem text and the predicted retrieval result; Use a preset optimization algorithm to successively adjust the parameters of the initial embedding model to reduce the difference until the training end condition is reached, and obtain the embedding model.
6. The semantic recall method based on a large language model according to claim 5, wherein The fusing the semantic vector of the problem text with the semantic vectors of each sub-question to obtain a trained semantic vector includes: Perform a weighted linear operation on the semantic vector of the problem text and the semantic vectors of each sub-question to obtain the trained semantic vector.
7. A semantic recall device based on a large language model, characterized in that, Includes: A first decomposition module for guiding a large language model to decompose a query problem into multiple sub-questions by using a problem decomposition prompt word; A second decomposition module for semantically encoding the query problem by using an embedding model to obtain a semantic vector of the query problem; An encoding module for semantically encoding each sub-question in the multiple sub-questions respectively by using the embedding model to obtain a semantic vector of each sub-question; A fusion module for fusing the semantic vector of the query problem and the semantic vectors of each sub-question to obtain a query semantic vector; A recall module for recalling the retrieval result of the query problem according to the similarity between the query semantic vector and the semantic vectors of the knowledge in the knowledge database.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the semantic recall method based on a large language model according to any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the semantic recall method based on a large language model according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the semantic recall method based on a large language model according to any one of claims 1 to 6.