Large language model and multi-modal model retrieval method, electronic equipment and medium

By combining preprocessing methods of generating vectors and inverted index databases with large language models and multimodal models, the problem of inaccurate understanding of user intentions in the prior art is solved, and more accurate and highly granular answers are achieved, improving user experience.

CN120353945APending Publication Date: 2025-07-22STATE GRID BUSINESS TRAVEL CLOUD TECH CO LTD

Patent Information

Application Number
CN202510825151.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The prior art cannot accurately understand user intentions in Q&A scenarios, has low answer accuracy, cannot understand multi-format documents, especially pictures and videos, and relies on manpower to build a knowledge base inefficient efficiency.

Method used

By preprocessing the knowledge base, a vector database and an inverted index database are generated, combined with a large language model and a multimodal model, search terms are preprocessed and retrieved, and a collection of vector recall slices and text recall slices are generated, irrelevant content is filtered, and the prompt word engineering is used to improve the accuracy of answers.

Benefits of technology

It achieves an accurate understanding of user intentions, improves the accuracy and granularity of answers, and can locate specific paragraphs or sentences of the document, improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353945A_ABST
    Figure CN120353945A_ABST
Patent Text Reader

Abstract

The invention provides a retrieval method for a large language model and a multi-modal model, electronic equipment and a medium. The method comprises the steps of performing first preprocessing on a knowledge base to obtain a vector database; performing second preprocessing on the knowledge base to obtain a reverse index database; obtaining a search word; performing third preprocessing on the search word to obtain a first search word set, performing retrieval by a vector database according to the first search word set, and obtaining a vector recall slice set according to the vector database; performing fourth preprocessing on the search word to obtain a second search word set, performing retrieval by a reverse index database according to the second search word set, and recalling a text slice set according to feedback of the reverse index database; and receiving the vector recall slice set and the text slice set, combining the vector recall slice set and the text slice set into a recall slice set, processing the recall slice set, and inputting the processed recall slice set into a large language model to obtain a retrieval result. Therefore, the retrieval result provided by the invention is more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of retrieval, and particularly relates to a retrieval method for large language models and multi-modal models, an electronic device, and a readable storage medium. Background Art

[0002] In the existing question-and-answer scenarios, firstly, a large number of users consult repetitive or similar questions; secondly, some questions appear in the form of pictures or videos; thirdly, a large number of questions are frequent, repetitive, and similar questions, which take a lot of time for question answerers. The existing technologies often rely on question-and-answer pairs, have a low reply accuracy rate, a slow knowledge base construction process, and rely heavily on manpower. The existing technologies mainly rely on traditional natural language processing technologies and search technologies, cannot accurately understand the user's intention, have a limited answer accuracy rate, cannot give targeted answers to questions, and the user experience is poor. The existing methods have limited support for multi-format documents and cannot understand complex structures, pictures, videos, etc. in the documents. The existing methods have insufficient reply granularity and cannot accurately locate complete document fragments, being limited to question-and-answer pairs or document levels. Summary of the Invention

[0003] In order to solve the technical defect that the large language model and multi-modal model in the existing technology answer inaccurately, the present invention adopts the following technical solutions: In the first aspect of the present invention, a retrieval method for large language models and multi-modal models is provided. The retrieval method for large language models and multi-modal models of the present invention includes: Performing a first preprocessing on a knowledge base to obtain a vector database; Performing a second preprocessing on the knowledge base to obtain an inverted index database; Obtaining a search term; Performing a third preprocessing on the search term to obtain a first retrieval term set, retrieving according to the first retrieval term set from the vector database, and obtaining a vector recall slice set according to the vector database; Performing a fourth preprocessing on the search term to obtain a second retrieval term set, retrieving according to the second retrieval term set from the inverted index database, and obtaining a text recall slice set according to the feedback of the inverted index database; Receiving the vector recall slice set and the text recall slice set, combining them into a recall slice set, and after processing the recall slice set, inputting it into a large language model to obtain a retrieval result.

[0004] Further, the step of after processing the recall slice set and inputting it into a large language model to obtain a retrieval result includes: Performing re-ranking on the recall slice set to obtain a corpus set; Creating a prompt engineering according to the corpus set and then inputting it into the large language model to obtain a retrieval result.

[0005] Further, the first preprocessing of the knowledge base to obtain a vector database includes: Performing text slicing processing on the knowledge base to obtain text slices; Performing vectorization processing on the text slices to obtain a vector database.

[0006] Further, performing text slicing processing on the knowledge base to obtain text slices includes: Obtaining first text information in the knowledge base and performing first text slicing processing on the first text information; Obtaining video information and image information in the knowledge base, generating second text information after processing the video information and the image information through a multi-modal model, and performing second text slicing processing on the second text information.

[0007] Further, the second preprocessing of the knowledge base to obtain an inverted index database includes: Performing word segmentation processing on the knowledge base to obtain an inverted index database.

[0008] Further, the third preprocessing of the search term to obtain a first set of retrieval terms includes: Performing vectorization on the search term to obtain a first set of retrieval terms.

[0009] Further, the fourth preprocessing of the search term to obtain a second set of retrieval terms includes: Performing text slicing processing on the search term to obtain a second set of retrieval terms.

[0010] Further, the re-ranking of the recall slice set to obtain a corpus set includes: Filtering out irrelevant content according to the relevance between the elements in the recall slice set and the search term according to the re-ranking algorithm.

[0011] In the second aspect of the present invention, an electronic device is provided. The above-mentioned electronic device includes a processor and a memory, and the processor is used to execute a computer program stored in the memory to implement the retrieval methods of the above-mentioned large language model and multi-modal model.

[0012] In the third aspect of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium stores at least one instruction, and the at least one instruction is executed by the processor for the retrieval methods of the above-mentioned large language model and multi-modal model.

[0013] Compared with the prior art, the beneficial effects of the present invention are as follows: The retrieval method of the large language model and multi-modal model of the present invention vectorizes the retrieval terms, and the obtained first retrieval term set is vectorized. Thus, the first retrieval term can obtain retrieval results in the vector database of the present invention. The present invention obtains a second retrieval term set by performing text slicing on the retrieval terms, so that the second retrieval term maintains text characteristics and can obtain retrieval results in the inverted index database. Furthermore, in the subsequent question and answer content, the feedback of both the vector database and the inverted index database can be considered. Therefore, the retrieval method of the large language model and multi-modal model of the present invention can more accurately understand the user's intention. The present invention analyzes the relevance between the recall slice set and the search terms of the present invention, thereby filtering out some irrelevant content and making the acquisition of the retrieval intention more accurate and rich. An electronic device and a computer-readable storage medium for a large language model and multi-modal model provided by the present invention also solve the problems raised in the background art section. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The specification drawings forming a part of this application are used to provide a further understanding of the present invention. The schematic embodiments and descriptions of the present invention are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings: Figure 1 is a flowchart of an embodiment of the present invention; Figure 2 is a flowchart of obtaining a vector database by performing a first preprocessing on the knowledge base of the present invention; Figure 3 is a flowchart of obtaining text slices by performing text slicing on the knowledge base of the present invention; Figure 4 is a flowchart of recall and retrieval of an embodiment of the present invention; Figure 5 is a schematic diagram of a user-side retrieval interface of an embodiment of the present invention; Figure 6 is a schematic diagram of a retrieval module of an embodiment of the present invention.

[0015] Figure 7 is a structural block diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] The present invention will be described in detail below with reference to the drawings and in conjunction with the embodiments. It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other.

[0017] The following detailed descriptions are all exemplary descriptions, aiming to provide further detailed explanations for the present invention. Unless otherwise specified, all technical terms adopted in the present invention have the same meanings as those commonly understood by those of ordinary skill in the art to which this application belongs. The terms used in the present invention are only for describing specific embodiments, and are not intended to limit the exemplary embodiments according to the present invention.

[0018] Embodiment 1 Figure 1 is a flowchart of an embodiment of the present invention. The process of the retrieval method for the large language model and multi-modal model of the present invention includes Step 100 - Step 600.

[0019] Step 100, perform a first preprocessing on the knowledge base to obtain a vector database.

[0020] The present invention collects various formats of information as the information in the knowledge base of the present invention. The information in the knowledge base includes both text information and video information and image information. Specifically, the text information (structured data) includes: text materials, question-and-answer pairs, and the video information (unstructured data) and image information (unstructured data) of the present invention include: video materials and image materials. After data cleaning and sorting of the above text information, video information and image information, a knowledge base is obtained.

[0021] The present invention obtains a vector database by performing a first preprocessing on the knowledge base. Specifically, Step 100 includes two sub-steps: Step 110 and Step 120.

[0022] Please refer to Figure 2 , Figure 2 which is a flowchart of the present invention for performing a first preprocessing on the knowledge base to obtain a vector database.

[0023] Step 110, perform text slicing processing on the knowledge base to obtain text slices; Step 120, perform vectorization processing on the text slices to obtain a vector database.

[0024] In the present invention, text slices are obtained by performing text slicing processing on the information in the knowledge base. Specifically, the present invention can perform text slicing processing on the information in the database in various ways, including: fixed-length chunking, sentence-based chunking, paragraph chunking, document chunking, sliding window chunking, semantic chunking, recursive chunking, context-enhanced chunking, etc.

[0025] The sliced text can be vectorized only after text slicing processing, so as to obtain a vector database. Only by creating a vector database in the present invention can a basis for text processing be provided for subsequent user questions and answers, thereby solving the problem that large language models in the existing natural language processing field and search field cannot accurately understand user intentions, resulting in low answer accuracy.

[0026] The above step 110 includes two sub-steps: step 111 and step 112.

[0027] Please refer to Figure 3 , Figure 3 which is a flowchart of the present invention for obtaining text slices by performing text slicing processing on the knowledge base.

[0028] Step 111: Obtain the first text information in the knowledge base and perform first text slicing processing on the first text information. Thus, text slices corresponding to the first text information can be obtained through step 111.

[0029] Step 112: Obtain the video information and image information in the knowledge base, generate second text information after processing the video information and the image information through a multi-modal model, and perform second text slicing processing on the second text information. Thus, text slices corresponding to the second text information can be obtained through step 112.

[0030] Existing methods have limited support for documents in multiple formats. To solve this technical defect, the retrieval method of the large language model and multi-modal model of the present invention needs to generate a vector database through text information, and also generate a vector database through image and video information.

[0031] The specific method is to perform step 111 on the text information and perform step 112 on the video information and image information. Since step 111 itself is the first text information, the first text slicing processing can be directly performed on the first text information.

[0032] Since step 112 obtains video information and image information, it is necessary to first process these video information and image information through a multi-modal model to generate second text information, and then perform second text slicing processing on the second text information.

[0033] Since the present invention processes text - type information in step 111, processes video information and image information in step 112, and then combines step 120 to perform vectorization processing on the text slices to obtain a vector database. Thus, the retrieval method of the large - language model and multi - modal model of the present invention can not only understand text information, but also understand complex image and video information. Therefore, the method of the embodiments of the present invention can better understand the user's intention. In other words, this is because when creating the vector database, the retrieval method of the large - language model and multi - modal model of the present invention comprehensively considers text information, image information, and video information.

[0034] Please continue to refer to Figure 1 。

[0035] The method of the embodiments of the present invention performs a second pre - processing on the knowledge base through step 200 to obtain an inverted index database. Specifically, the present invention obtains an inverted index database by performing word - segmentation processing on the text information in the knowledge base. The present invention performs word - segmentation processing on the knowledge base to obtain an inverted index database. Combining the previous step 100, the present invention obtains two databases through the knowledge base: one database is a vector database, and the other database is an inverted index database. This provides two different types of databases for the retrieval method of the present invention. Thus, in subsequent processing steps, since the present invention has two databases, the two databases (vector database and inverted index database) of the present invention can recall slice sets respectively. After these slice sets are combined with the subsequent processing of prompt engineering and then input into the large - language model for questioning, it can better match the true intention of the questioner. In other words, since the vector database contains content from videos, images, and texts, when the user asks a question, if the question contains pictures or videos, the present invention, because the vector content corresponding to the pictures or videos has been stored in the vector database, combined with subsequent processing links, can accurately understand the questions provided by the user with pictures or videos. That is to say, the present application solves the defect that the large - language models of the prior art cannot understand the complex structure of documents, pictures, and videos.

[0036] Step 300, obtain a search term. The search term of the present invention is obtained by rewriting the user's original question. Specifically, the rewriting includes: cleaning, word - segmentation, recombination, and rewriting of the user's original question. The step of obtaining a search term by rewriting the user's original question can improve the expression ability of the user's original question.

[0037] Step 400, perform a third pre - processing on the search term to obtain a first retrieval term set, retrieve according to the first retrieval term set from the vector database, and obtain a vector recall slice set according to the vector database.

[0038] In step 400, the specific operation of obtaining the first retrieval word set by performing the third preprocessing on the search word is as follows: vectorize the search word to obtain the first retrieval word set. Retrieve according to the first retrieval word set from the vector database, and obtain a vector recall slice set according to the vector database.

[0039] Step 500: Perform the fourth preprocessing on the search word to obtain the second retrieval word set, retrieve according to the second retrieval word set from the inverted index database, and recall a text slice set according to the feedback of the inverted index database.

[0040] The specific operation of performing the fourth preprocessing on the search word in step 500 to obtain the second retrieval word set is as follows: perform text slicing processing on the search word to obtain the second retrieval word set. The text slicing processing can specifically adopt a word segmentation algorithm. Perform word segmentation on the search word through the word segmentation algorithm to obtain the second retrieval word set. Retrieve according to the second retrieval word set from the inverted index database, and recall a text slice set according to the feedback of the inverted index database.

[0041] Through step 400 and step 500, in the present invention, by vectorizing the retrieval word, the obtained first retrieval word set is vectorized, and thus the first retrieval word can obtain a retrieval result in the vector database of the present invention. The present invention obtains the second retrieval word set by performing text slicing processing on the retrieval word, so that the second retrieval word maintains its text characteristics, and thus can obtain a retrieval result in the inverted index database. Furthermore, in the subsequent question and answer content, both the feedback of the vector database and the feedback of the inverted index database can be considered. Therefore, the retrieval method of the present invention can more accurately understand the user's intention.

[0042] Step 600: Receive the vector recall slice set and the text recall slice set, combine them into a recall slice set, and input the processed recall slice set into a large language model to obtain a retrieval result.

[0043] Since the slice sets recalled by the present invention include both vector-recalled slice sets and text-recalled slice sets, this provides a richer slice set for the large language model of the present application. Compared with the questions directly asked by users, the slice sets recalled through the vector database and the inverted index database of the present application can more accurately express the user's intention. Therefore, through the subsequent processing steps of the present invention, the large language model of the present invention can give more accurate answers. First, through the retrieval results fed back by the vector database and the inverted index database, compared with the original user's question, the retrieval results of the present invention are some slice sets, which can more accurately express the intention of the original questioner, thus overcoming the defects of low answer accuracy and poor user experience in the prior art. Second, since the feedback results are some slice sets, and the slice sets include both text-recalled slice sets and vector-recalled slice sets, this provides richer corpus for the large language model of the present application before processing the corpus, and the answer content is more accurate. Third, since the text-recalled slice sets of the present invention come from the inverted index database, and the inverted index database is obtained by slicing the knowledge base. The slicing process slices articles, paragraphs, and sentences, and after slicing, they are stored in the inverted index database. In this way, for the original user's question, the text-recalled slices can locate the article, paragraph, and / or sentence, and the beneficial effect is a higher granularity. In other words, according to a specific question of the original user, since the prior art does not adopt the method of the present application, the prior art can only locate the article and cannot give the answer of the specific paragraph and specific sentence in the article according to a specific question of the original user. The user can only read the whole article to find the answer. However, the retrieval methods of the large language model and the multi-modal model of the present application can locate the paragraph and / or sentence, and the retrieval answers given by the present application have a high granularity and are more accurate, improving the defect that the user in the prior art can only read the whole article to find the answer.

[0044] Specifically, step 600 includes two sub-steps: step 610 and step 620.

[0045] Please refer to Figure 4 , Figure 4 which is the recall and retrieval flow chart of the embodiment of the present invention.

[0046] The present invention includes both recall by the vector database and recall by the inverted index database. The specific recall method is to retrieve the first set of retrieval terms by the vector database to obtain the recall by the vector database; retrieve the second set of retrieval terms by the inverted index database to obtain the recall by the inverted index database. Combine the recall by the vector database and the recall by the inverted index database to obtain the recall slice set. That is to say, there are both vectors and text in the recall slice set of the present invention. The text in the recall slice set includes: articles, paragraphs, sentences, etc.

[0047] Step 610: Reorder the recalled slice set to obtain a corpus set.

[0048] The specific method for reordering the recalled slice set in the present invention is as follows: Analyze the relevance between the recalled slice set and the search term of the present invention, so as to filter out some irrelevant content. After analyzing the relevance between the recalled slice set and the search term and filtering out the parts irrelevant to the search term, the corpus set is obtained.

[0049] Step 620: Create a prompt engineering based on the corpus set and input it into a large language model to obtain a retrieval result.

[0050] In the present invention, the corpus is combined with prompt engineering and input into a large language model to obtain a retrieval result. The prompt engineering of the present invention can adopt different frameworks, specifically including: CRISPE framework, BROKE framework, ICIO framework, CoT framework. In the embodiments of the present invention, the ICIO and Cot frameworks are preferably used.

[0051] By combining the corpus with prompt engineering and inputting it into a large language model, the large language model can better understand the intention of the questioner, so as to give an answer.

[0052] Please refer to Figure 5 , Figure 5 which is a schematic diagram of the user - side retrieval interface of the embodiment of the present invention.

[0053] The retrieval interface of the user - side in the embodiment of the present invention is as Figure 5 shown. The user inputs the original user question in the question dialog box 1000. Through the retrieval method of the large language model and the multi - modal model used in steps 100 - 600 of the present invention, the answer can be fed back to the user in the retrieval result feedback dialog box 2000.

[0054] Please refer to Figure 6 , Figure 6 which is a schematic diagram of the retrieval module of the embodiment of the present invention.

[0055] The retrieval module of the embodiment of the present invention includes a knowledge base module 700 for storing the knowledge base of the present invention. The knowledge base of the present invention can be obtained by performing data cleaning and sorting on texts, question-and-answer pairs, and structured data. A slicing module 710. The slicing module 710 of the present invention is used to execute step 110 of the present invention. In step 110, text slicing processing is performed on the knowledge base to obtain text slices. Specifically, the slicing module 710 includes a text slicing module 711 and a multimodal processing module 712. Among them, the slicing module 710 executes step 111, and the multimodal processing module 712 is used to execute step 112. Step 112: Obtain video information and image information in the knowledge base, generate second text information after processing the video information and the image information through a multimodal model, and perform second text slicing processing on the second text information. By means of the method of the slicing module 710 executing step 110, the embodiment of the present invention can process both text data and image and video data. A first vectorization module 730, a vector database 750. Step 120 is executed on the first vectorization module 730. Step 120: Perform vectorization processing on the text slices to obtain a vector database. The present invention performs vectorization processing on the text slices by executing step 120 on the first vectorization module 730, and then stores each vector obtained after the vectorization processing in the vector database 750. Since the present invention executes step 112 through the multimodal processing module 712 before vector processing, step 112 is to generate second text information by processing image information and video information through a multimodal model, and then perform vectorization processing on the second text information and store it in the vector database 750. Thus, the vector database 750 of the present invention stores the content vectorized from image information and video information, and thus can feedback richer content through the vector database in the subsequent user question session.

[0056] The first word segmentation module 720 is used to execute step 200. In step 200, a second preprocessing is performed on the knowledge base to obtain an inverted index database. Specifically, the present invention obtains the inverted index database by performing word segmentation processing on the knowledge base. That is to say, in addition to obtaining the vector database 750 through vectorization processing of the information in the knowledge base, the present invention also performs word segmentation processing on the text information in the knowledge base, thereby obtaining the inverted index database 740. The vector database 750 and the inverted index database 740 of the present invention are two different databases. The purpose of creating the inverted index database 740 is to provide a text retrieval database. For example, information in formats such as industry materials, technical manuals, books, and other text types is collected through the inverted index database 740. Furthermore, in subsequent user question sessions, richer and more professional content can be fed back through the inverted index database. By setting two databases (vector database and inverted index database), the retrieval method of the large language model and the multi-modal model of the present invention can not only understand text information, but also understand complex image and video information. Thus, for some questions when the user asks, in addition to text descriptions, there are also images and videos. The method of the embodiment of the present invention can not only understand the intention of the words in the user's question, but also understand the image intention and video intention of the user's question. That is to say, the embodiment of the present invention can better understand the user's intention.

[0057] The search term module 760 is used to execute step 300. In step 300, search terms are obtained. The search terms of the present invention are obtained by rewriting the user's original question. Specifically, please refer to Figure 5, the questioner inputs the original user question in the question dialog box 1000 on the retrieval interface of the client. Step 300 is executed by the search term module 760, so as to rewrite the original user question to obtain search terms. Since the search terms of the present invention are obtained after being rewritten, compared with the original user question, the expressiveness of the search terms is richer. The vectorized retrieval module 770 is used to execute step 400. In step 400, the search terms are subjected to third preprocessing to obtain a first retrieval term set, and the vector database 750 retrieves according to the first retrieval term set, and a vector recall slice set is obtained according to the vector database. The word segmentation retrieval module 780 is used to execute step 500. In step 500, the search terms are subjected to fourth preprocessing to obtain a second retrieval term set, and the inverted index database 740 retrieves according to the second retrieval term set, and a text slice set is recalled according to the feedback of the inverted index database. By vectorizing the retrieval terms, the first retrieval term set obtained by the present invention is vectorized, and thus the first retrieval term can obtain retrieval results in the vector database of the present invention. The present invention obtains a second retrieval term set by performing text slicing processing on the retrieval terms, so that the second retrieval term maintains text characteristics, and thus can obtain retrieval results in the inverted sorting database. Furthermore, in the subsequent Q&A content, both the feedback of the vector database and the feedback of the inverted index database can be considered. Therefore, the retrieval method of the present invention can better understand the user's intention.

[0058] The retrieval processing module 790 includes a recall module 791, a re-ranking module 792, and a large language model processing module 793. Among them, the retrieval processing module 790 is used to execute step 600. In step 600, the vector recall slice set and the text slice set are received, combined into a recall slice set, and the recall slice set is input into the large language model after being processed to obtain a retrieval result. Specifically, the recall module 791 is used to receive and store the recall slice set. The re-ranking module 792 is used to execute step 610. In step 610, the recall slice set is re-ranked to obtain a corpus set. The advantage of the re-ranking module 792 is that the recall slice set is analyzed for relevance with the search terms of the present invention, so as to filter out some irrelevant content. The large language model processing module 793 is used to execute step 620. In step 620, a prompt engineering is created according to the corpus set and then input into the large language model to obtain a retrieval result. The retrieval result is sent to the retrieval output module 820. Figure 5 , the retrieval output module 820 sends the retrieval result to the retrieval result feedback dialog box 2000, and feedbacks the answer to the user through the retrieval result feedback dialog box 2000.

[0059] Embodiment 2 As Figure 7As shown in the figure, the present invention also provides an electronic device 100 for implementing a retrieval method of a large language model and a multimodal model; The electronic device 100 includes a memory 101, at least one processor 102, a computer program 103 stored in the memory 101 and executable on the at least one processor 102, and at least one communication bus 104.

[0060] The memory 101 can be used to store the computer program 103. The processor 102 realizes the steps of the enhanced retrieval based on the large language model in Embodiment 1 by running or executing the computer program stored in the memory 101 and calling the data stored in the memory 101.

[0061] The memory 101 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the electronic device 100 (such as audio data, etc.). In addition, the memory 101 can include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices.

[0062] The at least one processor 102 can be a Central Processing Unit (CPU), and can also be other general-purpose processors, a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 102 can be a microprocessor or the processor 102 can also be any conventional processor, etc. The processor 102 is the control center of the electronic device 100, and connects various parts of the entire electronic device 100 through various interfaces and lines.

[0063] The memory 101 in the electronic device 100 stores multiple instructions to implement a retrieval method of a large language model and a multimodal model. The processor 102 can execute the multiple instructions to implement: Perform a first preprocessing on the knowledge base to obtain a vector database; Perform a second preprocessing on the knowledge base to obtain an inverted index database; Obtain a search term; Perform a third preprocessing on the search term to obtain a first set of retrieval terms, retrieve according to the first set of retrieval terms by a vector database, and obtain a vector recall slice set according to the vector database; Perform a fourth preprocessing on the search term to obtain a second set of retrieval terms, retrieve according to the second set of retrieval terms by the inverted index database, and recall a text slice set according to the feedback of the inverted index database; Receive the vector recall slice set and the text slice set, combine them into a recall slice set, and input the processed recall slice set into a large language model to obtain a retrieval result.

[0064] Embodiment 3 If the modules / units integrated in the electronic device 100 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above embodiment methods of the present invention, it can also be completed by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, and read-only memory (ROM, Read-Only Memory).

[0065] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0066] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0067] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction means that implements the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0068] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0069] In the description of this specification, the descriptions referring to the terms "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0070] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific implementation manners of the present invention, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the protection scope of the claims of the present invention.

Claims

1. A retrieval method for large language models and multimodal models, characterized in that, Including: Performing a first preprocessing on the knowledge base to obtain a vector database; Performing a second preprocessing on the knowledge base to obtain an inverted index database; Obtaining a search term; Performing a third preprocessing on the search term to obtain a first set of retrieval terms, retrieving according to the first set of retrieval terms from the vector database, and obtaining a set of vector recall slices according to the vector database; Performing a fourth preprocessing on the search term to obtain a second set of retrieval terms, retrieving according to the second set of retrieval terms from the inverted index database, and obtaining a set of text recall slices according to the feedback of the inverted index database; Receiving the set of vector recall slices and the set of text recall slices, combining them into a set of recall slices, and inputting the processed set of recall slices into a large language model to obtain a retrieval result.

2. The retrieval method of the large language model and the multimodal model according to claim 1, characterized in that, The inputting the processed set of recall slices into a large language model to obtain a retrieval result includes: Re-ranking the set of recall slices to obtain a corpus set; Creating a prompt engineering according to the corpus set and inputting it into a large language model to obtain a retrieval result.

3. The retrieval method of the large language model and the multimodal model according to claim 1, characterized in that, The performing a first preprocessing on the knowledge base to obtain a vector database includes: Performing text slicing on the knowledge base to obtain text slices; Performing vectorization on the text slices to obtain a vector database.

4. The retrieval method of the large language model and the multimodal model according to claim 3, characterized in that, The performing text slicing on the knowledge base to obtain text slices includes: Obtaining first text information in the knowledge base and performing a first text slicing on the first text information; Obtaining video information and image information in the knowledge base, generating second text information after processing the video information and the image information through a multi-modal model, and performing a second text slicing on the second text information.

5. The retrieval method of the large language model and the multimodal model according to claim 1, characterized in that, The performing a second preprocessing on the knowledge base to obtain an inverted index database includes: Performing word segmentation on the knowledge base to obtain an inverted index database.

6. The retrieval method of the large language model and the multimodal model according to claim 1, characterized in that, The performing a third preprocessing on the search term to obtain a first set of retrieval terms includes: Performing vectorization on the search term to obtain a first set of retrieval terms.

7. The retrieval method of the large language model and the multimodal model according to claim 1, characterized in that, The performing a fourth preprocessing on the search term to obtain a second set of retrieval terms includes: Performing text slicing on the search term to obtain a second set of retrieval terms.

8. The retrieval method of the large language model and the multimodal model according to claim 2, characterized in that, The re-ranking the set of recall slices to obtain a corpus set includes: Filtering out irrelevant content according to the relevance between the elements in the set of recall slices and the search term according to a re-ranking algorithm.

9. An electronic device, characterized in that, Including a processor and a memory, the processor is used to execute a computer program stored in the memory to implement the retrieval method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one instruction, and the at least one instruction is executed by a processor to implement the retrieval method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Vector knowledge base-based large-model question-answer dialogue method and system and storage medium

    CN118484526A

  • Knowledge base question and answer method and device and computer readable storage medium

    CN119128096A

  • Medical AI question and answer method based on retrieval enhancement

    CN119541887A

  • Inverted indexes with multiple language support

    US20120158718A1

Cited By

  • Knowledge base construction method, knowledge retrieval method, display method, equipment, storage medium and program product

    CN120973952A