Performance optimization method and system for big language model knowledge base retrieval
The query conditions entered by the user are analyzed semantically through a large language model, and combined with inverted index and approximate closest search algorithm to optimize the knowledge base search, solving the problem of inconsistent search results and improving the search efficiency and answer quality.
Patent Information
- Application Number
- CN202411993679.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
AI Technical Summary
In the existing knowledge base search process, the search results are prone to inconsistent conditions, resulting in inefficiency and low quality of answers.
The query conditions entered by the user are analyzed through a large language model, combined with inverted indexing and approximate closest search algorithms, relevant information is retrieved from the knowledge base, and the search process is optimized by diversity strategies, rearrangement strategies, parallel computing and cache technology, and answers are generated using multiple embedding models.
It improves the search speed and efficiency, enhances the accuracy and fit of generated answers, and optimizes the search experience.
Smart Images

Figure CN119938833A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of AI, and in particular to a performance optimization method and system for large language model knowledge base retrieval. Background Art
[0002] RAG (Retrieval Augmented Generation) is a technology that combines knowledge base retrieval with large language model generation for tasks such as question answering and text generation. This technology can make full use of the existing knowledge in the knowledge base and combine it with the generation capability of the large language model to provide more accurate, comprehensive and fluent answers or texts.
[0003] AI knowledge base is a knowledge management platform based on artificial intelligence technology. It establishes an efficient knowledge management system by learning and analyzing massive data. AI knowledge base can integrate, classify and optimize scattered knowledge resources, and realize automatic extraction, storage and query of knowledge through technologies such as natural language processing, thereby providing intelligent knowledge management services for enterprises. Large Language Model (LLM) refers to a neural network model trained with a large amount of text data, which can generate text similar to human language. Large language models can be used for various tasks such as question answering, text summarization, and machine translation.
[0004] In the existing search process, it is easy for the search results to be inconsistent. Summary of the invention
[0005] To achieve the above-mentioned purpose and other related purposes, the present invention discloses a performance optimization method for large language model knowledge base retrieval, comprising: The user enters the query criteria; The large model performs semantic analysis on the query conditions entered by the user to understand the user's intention; According to the semantic analysis results of the large model, relevant information is retrieved from the knowledge base, wherein an inverted index and an approximate nearest neighbor search algorithm are used to retrieve relevant information from the knowledge base; Generate answers based on the retrieved information and the semantic analysis results of the large model.
[0006] Furthermore, the retrieval module retrieves relevant information from the knowledge base according to the semantic analysis result of the large model, including: Use pre-trained Embedding models.
[0007] Furthermore, the retrieval module retrieves relevant information from the knowledge base according to the semantic analysis result of the large model, including: Use a sparse embedding model.
[0008] Furthermore, the retrieval module retrieves relevant information from the knowledge base according to the semantic analysis result of the large model, including: Diversity strategy and rearrangement strategy were used for retrieval.
[0009] Furthermore, the retrieval module retrieves relevant information from the knowledge base according to the semantic analysis result of the large model, including: Parallel computing is used to perform retrieval, and cache is used to store retrieval results.
[0010] Furthermore, the generation module uses a multiple embedding model to generate answers.
[0011] In another aspect, the present invention provides a performance optimization system for large language model knowledge base retrieval, comprising: The query condition input module is used for the user to input the query condition; Semantic analysis module, which is used by the large model to perform semantic analysis on the query conditions input by the user and understand the user's intention; A retrieval module is used to retrieve relevant information from the knowledge base according to the semantic analysis results of the large model, wherein the retrieval module uses an inverted index and an approximate nearest neighbor search algorithm to retrieve relevant information in the knowledge base; The generation module is used to generate answers based on the retrieved information and the semantic analysis results of the large model.
[0012] By adopting the above technical solution, the retrieval speed and efficiency are improved, and the efficiency of generating answers is improved, thereby optimizing the retrieval experience and improving the fit of the final generated answers. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, among which:
[0014] Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION
[0015] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0016] Reference Figure 1The embodiment of the present invention provides a performance optimization method for large language model knowledge base retrieval, including the following steps: S1: The user enters the query conditions.
[0017] S2: The large model performs semantic analysis on the query conditions entered by the user to understand the user's intention.
[0018] S3: Retrieve relevant information from the knowledge base based on the semantic analysis results of the large model.
[0019] Among them, the specific process of retrieval is optimized in the embodiment of the present invention: Use more effective retrieval algorithms. Traditional retrieval algorithms, such as the BM25 algorithm, can effectively retrieve relevant information, but their computational efficiency is relatively low. In order to improve retrieval efficiency, the following methods can be used:
[0020] Use inverted index: Inverted index is a data structure that indexes document content into a dictionary, which can effectively improve retrieval speed.
[0021] Use the approximate nearest neighbor search algorithm: The approximate nearest neighbor search algorithm can quickly find documents similar to the query vector, thereby improving retrieval efficiency.
[0022] Use an efficient Embedding model. The Embedding model is a technology that converts text into vector representation. The efficiency of the Embedding model directly affects the efficiency of the retrieval module. To improve retrieval efficiency, you can use the following methods:
[0023] Use pre-trained Embedding models: Pre-trained Embedding models have been trained and can be used directly, which can effectively improve the retrieval speed.
[0024] Use sparse Embedding model: Sparse Embedding model can effectively reduce the size of Embedding model, thereby improving retrieval speed.
[0025] Optimize the search strategy. The search strategy refers to how to select and sort the search results. In order to improve the search efficiency, you can optimize the search strategy, for example:
[0026] Use diversity strategies: Diversity strategies can ensure the diversity of search results and avoid duplicate results.
[0027] Use re-ranking strategies: Re-ranking strategies can re-rank search results based on the relevance of queries and documents to improve the accuracy of search results.
[0028] Parallel computing can effectively improve retrieval efficiency. In order to improve retrieval efficiency, parallel computing technology can be used, such as:
[0029] Use distributed indexes: Distributed indexes can spread index data across multiple servers, thereby increasing retrieval speed.
[0030] Use multithreading: Multithreading can execute multiple retrieval tasks at the same time, thereby increasing the retrieval speed.
[0031] Use cache, which can store search results, thereby reducing the number of repeated searches and improving search efficiency.
[0032] S4: Generate answers based on the retrieved information and the semantic analysis results of the large model.
[0033] Specifically, when generating answers, the present invention adopts a multiple embedding method. Multiple embedding refers to using multiple embedding models in one model to represent the same text. Multiple embedding can represent text from different angles, thereby improving the performance of the model.
[0034] The main benefits of multiple embeddings include: Improve the robustness of the model: Different embedding models have different advantages and disadvantages. Using multiple embeddings can make up for the shortcomings of each embedding model and improve the robustness of the model.
[0035] Improve the accuracy of the model: Multiple embeddings can provide richer semantic information, thereby improving the accuracy of the model.
[0036] Improve the generalization ability of the model: Multiple embeddings can help the model learn more general features, thereby improving the generalization ability of the model.
[0037] An embodiment of the present invention further provides a system, comprising: The query condition input module is used for the user to input the query condition; Semantic analysis module, which is used by the large model to perform semantic analysis on the query conditions input by the user and understand the user's intention; A retrieval module is used to retrieve relevant information from the knowledge base according to the semantic analysis results of the large model, wherein the retrieval module uses an inverted index and an approximate nearest neighbor search algorithm to retrieve relevant information in the knowledge base; The generation module is used to generate answers based on the retrieved information and the semantic analysis results of the large model.
[0038] Those skilled in the art will appreciate that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those generally understood by those skilled in the art in the art to which the present invention belongs. It should also be understood that terms such as those defined in general dictionaries should be understood to have meanings consistent with the meanings in the context of the prior art, and will not be interpreted with idealized or overly formal meanings unless specifically defined.
[0039] For the method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should know that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.
[0040] It can be known from the description of the above implementation modes that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application can be essentially or partly contributed to the prior art in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes several instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute the methods described in the various implementation modes of the present application or certain parts of the implementation modes.
[0041] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A performance optimization method for large language model knowledge base retrieval, characterized in that: include: The user enters the query criteria; The large model performs semantic analysis on the query conditions entered by the user to understand the user's intention; According to the semantic analysis results of the large model, relevant information is retrieved from the knowledge base, wherein an inverted index and an approximate nearest neighbor search algorithm are used to retrieve relevant information from the knowledge base; Generate answers based on the retrieved information and the semantic analysis results of the large model.
2. The method according to claim 1, characterized in that The retrieval module retrieves relevant information from the knowledge base according to the semantic analysis results of the large model, including: Use a pre-trained Embedding model.
3. The method according to claim 1, characterized in that The retrieval module retrieves relevant information from the knowledge base according to the semantic analysis results of the large model, including: Use a sparse embedding model.
4. The method according to claim 1, characterized in that: The retrieval module retrieves relevant information from the knowledge base according to the semantic analysis results of the large model, including: Diversity strategy and rearrangement strategy were used for retrieval.
5. The method according to claim 1, characterized in that The retrieval module retrieves relevant information from the knowledge base according to the semantic analysis results of the large model, including: Parallel computing is used to perform retrieval, and cache is used to store retrieval results.
6. The method according to claim 1, characterized in that The generation module uses a multiple embedding model to generate answers.
7. A performance optimization system for large language model knowledge base retrieval, characterized in that: include: The query condition input module is used for the user to input the query condition; Semantic analysis module, which is used by the large model to perform semantic analysis on the query conditions input by the user and understand the user's intention; A retrieval module is used to retrieve relevant information from the knowledge base according to the semantic analysis results of the large model, wherein the retrieval module uses an inverted index and an approximate nearest neighbor search algorithm to retrieve relevant information from the knowledge base; The generation module is used to generate answers based on the retrieved information and the semantic analysis results of the large model.