Retrieval method and device, electronic equipment, storage medium and computer program product

By storing and processing vector data on local devices, and employing keyword reordering algorithms and local vector indexes, the issues of real-time performance, privacy protection, and cost control in cloud-based vector databases are resolved, resulting in efficient and accurate retrieval results.

CN120873173AActive Publication Date: 2025-10-31UNIONTECH SOFTWARE TECH CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202511393823.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2025-10-31
Estimated Expiration
2045-09-26

AI Technical Summary

Technical Problem

The existing RAG technology framework relies on cloud-based vector databases for vector retrieval, which has technical limitations in terms of real-time performance, privacy protection, and cost control.

Method used

Vector data is stored and processed on local devices, using a keyword reordering algorithm. It is retrieved through vector indexes in local non-volatile memory and managed using a predefined buffer and non-volatile memory, reducing reliance on the cloud.

Benefits of technology

It improves the accuracy of search results, avoids network latency and high costs, ensures privacy protection, and reduces reliance on vectorization models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120873173A_ABST
    Figure CN120873173A_ABST
Patent Text Reader

Abstract

The invention relates to a retrieval method and device, electronic equipment, a storage medium and a computer program product. The method comprises the following steps: acquiring a query vector of a query problem; utilizing a vector index in a local nonvolatile memory to determine a first preset number of vectors with the similarity higher than the query vector, and taking the vectors as vector retrieval results; obtaining a text block corresponding to the vector retrieval result from a nonvolatile memory; performing word segmentation on the text block and the query problem corresponding to the vector retrieval result; creating a keyword index based on the word segmentation result of the text block corresponding to the vector retrieval result; keyword retrieval is conducted on the basis of the word segmentation result of the query question and the keyword index, a second preset number of text blocks with the high matching degree with the query question are obtained, and the second preset number of text blocks are used for assisting a local large language model in generating answers for the query question.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of deep learning, and more particularly to a retrieval method and apparatus, electronic device, storage medium, and computer program product. Background Technology

[0002] With the rapid development of artificial intelligence technology, especially Natural Language Processing (NLP), large language models (LLMs) based on deep learning have achieved remarkable results in various fields such as text generation, dialogue systems, and machine translation. However, the knowledge bases of these large-scale pre-trained language models often suffer from lag and insufficient specialization, failing to cover the latest information in specific domains or in-depth knowledge of specific industries. To address this issue, Retrieval-Augmented Generation (RAG) technology has emerged. By introducing an external knowledge base onto LLM, it enhances the output of the generative model through vector retrieval, thereby improving the accuracy and relevance of the generated results.

[0003] Because traditional databases are inefficient at processing vector data and similarity searches, existing RAG (Rapid Algorithm for Generative Data Retrieval) frameworks generally rely on cloud-based vector databases to handle large-scale data retrieval needs. However, cloud-based vector database solutions have technical limitations in terms of real-time performance, privacy protection, and cost control. Summary of the Invention

[0004] This disclosure provides a retrieval method and apparatus, electronic device, storage medium, and computer program product to at least address the technical limitations of related technologies that rely on cloud-based vector databases for vector retrieval in terms of real-time performance, privacy protection, and cost control.

[0005] According to a first aspect of the present disclosure, a retrieval method is provided, comprising: obtaining a query vector for a query question; using a vector index in a local non-volatile memory to determine a first predetermined number of vectors with high similarity to the query vector as vector retrieval results, wherein the non-volatile memory is used to store vector indexes of external documents and text blocks corresponding to each vector under the vector indexes, the text blocks being obtained by segmenting the text content of the external documents; obtaining text blocks corresponding to the vector retrieval results from the non-volatile memory; performing word segmentation on the text blocks corresponding to the vector retrieval results and the query question; creating a keyword index based on the word segmentation results of the text blocks corresponding to the vector retrieval results; performing keyword retrieval based on the word segmentation results of the query question and the keyword index to obtain a second predetermined number of text blocks with high matching degree to the query question, wherein the second predetermined number of text blocks are used to assist a local large language model in generating an answer to the query question.

[0006] Optionally, before determining a first predetermined number of vectors with high similarity to the query vector using the vector index in local non-volatile memory and using them as vector retrieval results, the method further includes: creating a planar vector index for the external document; writing the planar vector index into a predetermined buffer in local memory, wherein the predetermined buffer is a region in memory used to store the planar vector index of the external document; and writing the planar vector index in the predetermined buffer into non-volatile memory in response to the amount of data in the predetermined buffer reaching a preset threshold. The determination of the first predetermined number of vectors with high similarity to the query vector using the vector index in local non-volatile memory and using them as vector retrieval results includes: determining the first predetermined number of vectors with high similarity to the query vector using the vector index in the predetermined buffer and non-volatile memory, and using them as vector retrieval results.

[0007] Optionally, in response to the data volume in the predetermined buffer reaching a preset threshold, writing the planar vector index in the predetermined buffer to non-volatile memory includes: in response to the data volume in the predetermined buffer reaching the preset threshold and the number of vectors under the planar vector index in the predetermined buffer being greater than a third predetermined number, creating a non-planar vector index based on the planar vector index in the predetermined buffer; writing the non-planar vector index to non-volatile memory, and clearing the predetermined buffer.

[0008] Optionally, creating a planar vector index for an external document includes: in response to the external document being a plain text document, dividing the text content of the external document into blocks using a sliding window approach to obtain multiple text blocks; in response to the external document being a structured document, dividing the text content of the external document into blocks based on the hierarchical information of the external document to obtain multiple text blocks, wherein each text block contains hierarchical information of all levels to which it belongs in the external document; vectorizing each text block to obtain a vector for each text block; and creating a planar vector index based on the vector of each text block.

[0009] Optionally, the system receives a list of files to be deleted, wherein the list includes identifiers of at least one document to be deleted; queries the index creation status of at least one document to be deleted; for documents to be deleted whose index creation status is "index creation successful," queries a predetermined vector index from memory based on the identifier of the document to be deleted; in response to the existence of the predetermined vector index in memory, deletes the predetermined vector index from memory; in response to the absence of the predetermined vector index in memory, queries the predetermined vector index from non-volatile memory and sets the deletion flag corresponding to the predetermined vector index to 1, wherein the predetermined vector index is the vector index of the document to be deleted; for documents to be deleted whose index creation status is "in the index creation process," stops the index creation process of the document to be deleted; for documents to be deleted whose index creation status is "index creation failed," skips the document to be deleted.

[0010] Optionally, using vector indexes in a predetermined buffer and non-volatile memory, a first predetermined number of vectors with the highest similarity to the query vector are determined and used as vector retrieval results. This includes: using vector indexes in a predetermined buffer of memory to determine the similarity between each vector under the vector index in the predetermined buffer and the query vector; using vector indexes in non-volatile memory to determine the similarity between each vector under the vector index in non-volatile memory and the query vector; merging the vectors under the vector indexes in the predetermined buffer and non-volatile memory and sorting them in descending order of similarity, and determining the first predetermined number of vectors with the highest similarity as vector retrieval results.

[0011] Optionally, the similarity between each vector under the vector index in the non-volatile memory and the query vector is determined using the vector index in the non-volatile memory, including: constructing a bit map based on the deletion flag bit of the vector index in the non-volatile memory; and using the bit map and the vector index in the non-volatile memory, determining the similarity between each vector under the vector index in the non-volatile memory where the deletion flag bit is 0 and the query vector.

[0012] According to a second aspect of the present disclosure, a retrieval apparatus is provided, comprising: a first acquisition unit configured to acquire a query vector of a query question; a first retrieval unit configured to use a vector index in a local non-volatile memory to determine a first predetermined number of vectors with high similarity to the query vector as vector retrieval results, wherein the non-volatile memory is used to store a vector index of an external document and text blocks corresponding to each vector under the vector index, the text blocks being obtained by segmenting the text content of the external document; a second acquisition unit configured to acquire text blocks corresponding to the vector retrieval results from the non-volatile memory; a word segmentation unit configured to segment the text blocks corresponding to the vector retrieval results and the query question; a creation unit configured to create a keyword index based on the word segmentation results of the text blocks corresponding to the vector retrieval results; and a second retrieval unit configured to perform keyword retrieval based on the word segmentation results of the query question and the keyword index to obtain a second predetermined number of text blocks with high matching degree to the query question, wherein the second predetermined number of text blocks are used to assist a local large language model in generating an answer to the query question.

[0013] Optionally, the first retrieval unit is further configured to: create a planar vector index of the external document before determining a first predetermined number of vectors with high similarity to the query vector using the vector index in local non-volatile memory and using them as vector retrieval results; write the planar vector index into a predetermined buffer in local memory, wherein the predetermined buffer is a region in memory used to store the planar vector index of the external document; write the planar vector index in the predetermined buffer into non-volatile memory in response to the amount of data in the predetermined buffer reaching a preset threshold; and determine a first predetermined number of vectors with high similarity to the query vector using the vector index in the predetermined buffer and non-volatile memory and use them as vector retrieval results.

[0014] Optionally, the first retrieval unit is further configured to, in response to the data volume in the predetermined buffer reaching a preset threshold and the number of vectors under the planar vector index in the predetermined buffer being greater than a third predetermined number, create a non-planar vector index based on the planar vector index in the predetermined buffer; write the non-planar vector index to non-volatile memory; and clear the predetermined buffer.

[0015] Optionally, the first retrieval unit is further configured to, in response to the external document being a plain text document, divide the text content of the external document into blocks using a sliding window method to obtain multiple text blocks; in response to the external document being a structured document, divide the text content of the external document into blocks based on the hierarchical information of the external document to obtain multiple text blocks, wherein each text block contains hierarchical information of all levels to which it belongs in the external document; vectorize each text block to obtain a vector for each text block; and create a planar vector index based on the vector of each text block.

[0016] Optionally, the above apparatus further includes a deletion unit configured to receive a list of deleted files, wherein the list of deleted files includes an identifier of at least one document to be deleted; query the index creation status of at least one document to be deleted; for a document to be deleted whose index creation status is successful, query a predetermined vector index from memory based on the identifier of the document to be deleted; in response to the existence of the predetermined vector index in memory, delete the predetermined vector index from memory; in response to the absence of the predetermined vector index in memory, query the predetermined vector index from non-volatile memory and set the deletion flag corresponding to the predetermined vector index to 1, wherein the predetermined vector index is the vector index of the document to be deleted; for a document to be deleted whose index creation status is in the index creation process, stop the index creation process of the document to be deleted; for a document to be deleted whose index creation status is failed, skip the document to be deleted.

[0017] Optionally, the first retrieval unit is further configured to: utilize the vector index in the predetermined buffer of memory to determine the similarity between each vector under the vector index in the predetermined buffer and the query vector; utilize the vector index in non-volatile memory to determine the similarity between each vector under the vector index in non-volatile memory and the query vector; merge the vectors under the vector index in the predetermined buffer and non-volatile memory and sort them in descending order of similarity, and determine the first predetermined number of vectors with the highest similarity as the vector retrieval results.

[0018] Optionally, the first retrieval unit is further configured to construct a bitmap based on the deletion flag bit of the vector index in the non-volatile memory; and to determine the similarity between each vector under the vector index in the non-volatile memory where the deletion flag bit is 0 and the query vector using the bitmap and the vector index in the non-volatile memory.

[0019] According to a third aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute instructions to implement a retrieval method according to the present disclosure.

[0020] According to a sixth aspect of the present disclosure, a computer-readable storage medium is provided that, when instructions in the computer-readable storage medium are executed by at least one processor, causes at least one processor to perform the retrieval method as described above according to the present disclosure.

[0021] According to a seventh aspect of the present disclosure, a computer program product is provided, including computer instructions that, when executed by a processor, implement the retrieval method according to the present disclosure.

[0022] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects: According to the retrieval method, apparatus, electronic device, storage medium, and computer program product disclosed herein, vector data is stored locally and processed locally, eliminating reliance on cloud-based vector databases. This eliminates dependence on the network, avoids common cloud service issues such as network latency and bandwidth bottlenecks, avoids high storage costs and pay-per-access fees associated with cloud services, and avoids the security risks of transmitting sensitive data to the cloud. Furthermore, this disclosure proposes a keyword reordering algorithm, which first obtains a large number of retrieval results for later use during vector retrieval. Then, based on the retrieval results, it retrieves source text information from a predetermined cache and non-volatile memory, and performs word segmentation on both the query question and the source text information. The word segmentation results of the source text are used to create a keyword index. The word segmentation results of the query question are used to search the keyword index to obtain a second predetermined number of text blocks with the highest matching degree to the query question, which serve as the final retrieval results. This reduces reliance on the vectorization model and improves the accuracy of the retrieval results.

[0023] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0024] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.

[0025] Figure 1 This is a flowchart illustrating a retrieval method according to an exemplary embodiment; Figure 2 This is a schematic diagram illustrating a flowchart of document parsing and text segmentation according to an exemplary embodiment; Figure 3 This is a schematic diagram illustrating a vector index creation process according to an exemplary embodiment; Figure 4 This is a schematic diagram illustrating the relationship between three tables according to an exemplary embodiment; Figure 5 This is a schematic diagram illustrating a vector index deletion process according to an exemplary embodiment; Figure 6 This is a schematic diagram illustrating a vector retrieval process according to an exemplary embodiment; Figure 7 This is a system architecture diagram of a local vector database according to an exemplary embodiment; Figure 8 This is a block diagram illustrating a retrieval device according to an exemplary embodiment; Figure 9This is a block diagram of an electronic device 900 according to an embodiment of the present disclosure. Detailed Implementation

[0026] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0027] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following examples do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0028] It should be noted that the phrase "at least one of several items" in this disclosure refers to three parallel cases: "any one of the several items", "a combination of any number of the several items", and "all of the several items". For example, "including at least one of A and B" includes the following three parallel cases: (1) including A; (2) including B; (3) including A and B. Another example is "performing at least one of step one and step two", which indicates the following three parallel cases: (1) performing step one; (2) performing step two; (3) performing both step one and step two.

[0029] RAGs demonstrate outstanding performance in applications such as question-answering systems, intelligent customer service, knowledge base management, and complex task-oriented dialogue systems. By introducing a retrieval step, RAGs not only improve the accuracy of generated content but also enrich the generated content with more detail and knowledge connections. RAG application scenarios primarily involve two process steps: Retrieval steps: The information retrieval module acquires external or internal knowledge related to the user input, generates queries, and finds similar data fragments.

[0030] Generation steps: Input the search results into the generative model to generate responses or content that better match actual needs and context.

[0031] The information retrieval module in the above retrieval steps typically requires efficient vector similarity search to quickly find information relevant to the user's query from the knowledge base. Therefore, it needs to process and store large amounts of vector data to achieve similarity retrieval of high-dimensional data. Since traditional databases are inefficient in processing vector data and performing similarity searches, existing RAG technology frameworks generally rely on cloud-based vector databases to handle large-scale data retrieval needs. Cloud-based vector databases are typically based on distributed computing architectures to store and retrieve large amounts of vector data. The following are common implementation schemes for current mainstream cloud-based vector database technologies: Distributed storage: Vector data is stored in a distributed storage system. Data distribution is managed through sharding to handle large-scale datasets. These sharded data can be distributed across multiple nodes to achieve load balancing and efficient access.

[0032] Vector Indexing: Cloud-based vector databases typically employ specific vector indexing structures, such as HNSW (Hierarchical Navigable Small World): an approximate nearest neighbor search algorithm often used for fast high-dimensional vector retrieval; IVF-PQ (Inverted File Product Quantization): significantly reduces vector storage space through block partitioning and quantization to support fast matching of large-scale vector data; and LSH (Locality-Sensitive Hashing): a hashing technique used to accelerate vector similarity search, particularly suitable for high-dimensional data.

[0033] High-concurrency retrieval: The cloud-based vector database is designed with multi-level caching and high-concurrency processing mechanisms, combined with load balancing strategies, to support a large number of concurrent retrieval requests and ensure that the system response time is within a reasonable range.

[0034] API Interface: The cloud-based vector database interfaces with application systems through application programming interfaces (APIs) such as REST and gRPC, providing standardized query, insert, and update operations for use in various environments.

[0035] However, the aforementioned cloud-based vector database solution has the following technical limitations: Real-time performance: Cloud-based vector databases are limited by the quality and latency of network connections, and these delays can affect user experience, especially in applications with high real-time requirements.

[0036] Privacy Protection: In applications involving sensitive user information, such as medical Q&A and financial customer service, cloud data storage and retrieval pose privacy and security risks. The process of uploading user data to the cloud may lead to data leaks or privacy violations.

[0037] Cost control: Cloud services are typically charged based on the number of accesses or storage capacity. In application scenarios with high-frequency access, the cost can rise sharply, making them unsuitable for enterprises or organizations with limited budgets or who are cost-sensitive.

[0038] To address the aforementioned issues, this disclosure proposes a local vector database, where vector data is stored locally. Users can perform vector data storage and fast retrieval on their local devices, eliminating reliance on cloud-based vector databases. This eliminates network dependence, avoids network latency and bandwidth bottlenecks common in cloud services, significantly improves real-time performance, and avoids the high costs of cloud service storage and pay-per-access billing. It also avoids the security risks of transmitting sensitive data to the cloud. Furthermore, this disclosure proposes a keyword reordering algorithm. During vector retrieval, a large number of search results are first obtained for backup. Then, based on the search results, source text information is retrieved from a predetermined cache and non-volatile memory. Both the query question and the source text information are segmented. The segmentation results of the source text are used to create a keyword index. The segmentation results of the query question are used to search the keyword index, retrieving the top two predetermined number of text blocks with the highest matching degree to the query question as the final search result. This reduces reliance on the vectorization model and improves the accuracy of the search results. Furthermore, this disclosure also segments the vector data for storage. When a new vector index is written, it is first stored in a predetermined cache in local memory. Only when the amount of data in the predetermined cache reaches a preset threshold is the vector index in the predetermined cache written to non-volatile memory. This avoids the frequent file I / O operations caused by directly writing to non-volatile memory, and also avoids the problem of excessive local memory consumption caused by storing the entire vector index in memory. In addition, this disclosure uses a structured method for parsing, dividing the document into blocks according to its structural information. Each block includes hierarchical information about all levels it belongs to in the document, such as multi-level headings in a docx file and table information in an xslx table. This structural information can provide more information to the text blocks during segmentation, making the search results more accurate. In summary, compared with related technologies, this disclosure provides a more economical, efficient, and stable solution.

[0039] The following will describe in detail, with reference to the accompanying drawings, retrieval methods and apparatuses, electronic devices, storage media, and computer program products according to exemplary embodiments of the present disclosure.

[0040] Figure 1 This is a flowchart illustrating a retrieval method according to an exemplary embodiment, such as... Figure 1 As shown, the retrieval method includes the following steps: In step 101, the query vector for the query question is obtained.

[0041] As an example, a vector retrieval interface can be pre-configured. In response to the invocation of the vector retrieval interface, a query request input by the user can be received. This query request may include a query question. By vectorizing the query question, a query vector can be obtained. As an example, the above query request may also include the number of returned results, that is, indicating how many results related to the query question are included in the final returned search results, such as selecting the first number of results after sorting by relevance to the query question. This disclosure does not limit this.

[0042] As an example, the HTTP interface of the client-side vectorized model can be called, the text of the query question can be input, and the corresponding query vector can be obtained. This disclosure does not limit this.

[0043] return Figure 1 In step 102, the vector index in the local non-volatile memory is used to determine a first predetermined number of vectors with the highest similarity to the query vector, and these vectors are used as vector retrieval results. The non-volatile memory is used to store the vector index of the external document and the text block corresponding to each vector under the vector index. The text block is obtained by dividing the text content of the external document into blocks.

[0044] As an example, Euclidean distance can be used to calculate the similarity distance between two vectors. The smaller the distance, the more similar the semantics of the corresponding texts. This disclosure does not limit this.

[0045] According to an exemplary embodiment of this disclosure, prior to step S102, a planar vector index of an external document may be created; the planar vector index is written into a predetermined buffer in local memory, wherein the predetermined buffer is a region in memory used to store the planar vector index of the external document; in response to the amount of data in the predetermined buffer reaching a preset threshold, the planar vector index in the predetermined buffer is written into non-volatile memory; based on this, using the vector index in the local non-volatile memory, a first predetermined number of vectors with the highest similarity to the query vector are determined and used as vector retrieval results, which may include: using the vector index in the predetermined buffer and non-volatile memory to determine a first predetermined number of vectors with the highest similarity to the query vector and using them as vector retrieval results.

[0046] In this embodiment, the vector index is stored in segments. When a new vector index is written, it is first stored in a predetermined buffer in local memory. Only when the amount of data in the predetermined buffer reaches a preset threshold is the vector index in the predetermined buffer written to non-volatile memory. This avoids the frequent file I / O operations caused by directly writing to non-volatile memory, and also avoids the problem of excessive local memory consumption caused by storing all vector indexes in memory.

[0047] Specifically, to facilitate subsequent vector retrieval, external documents can be introduced in advance, i.e., vector indexes of external documents can be built in advance. After building the vector indexes of external documents, it is necessary to save the generated vector index data. However, if all the vector index data of external documents is stored in memory, memory consumption will be excessive; if all the vector index data of external documents is directly written to disk, it will lead to frequent file I / O operations. To avoid huge performance overhead, this disclosure proposes a vector caching mechanism.

[0048] As an example, when new vector index data needs to be written, this new vector index data is first stored in a pre-defined buffer, which serves as a temporary cache area and is typically stored in memory. When the amount of data in the pre-defined buffer reaches a preset threshold, the vector index data in the pre-defined buffer can be written to local non-volatile memory.

[0049] It should be noted that the aforementioned preset threshold can be set as needed. The aforementioned non-volatile memory can be a hard disk or other forms of memory, and this disclosure does not limit it.

[0050] As an example, when storing using the above-mentioned vector caching mechanism, subsequent vector retrieval needs to be performed simultaneously in a predetermined buffer and local non-volatile memory. For example, during vector retrieval, the vector index in the predetermined buffer and non-volatile memory can be used to determine the first predetermined number of vectors with the highest similarity to the query vector, and these vectors can be used as the vector retrieval results.

[0051] As an example, retrieving the text block corresponding to the vector retrieval result from non-volatile memory may include retrieving the text block corresponding to the vector retrieval result from a predetermined buffer and non-volatile memory. That is, if the predetermined buffer not only stores vector index data, the predetermined buffer may also be considered when retrieving the text block.

[0052] According to an exemplary embodiment of the present disclosure, the above-mentioned response to the data volume in the predetermined buffer reaching a preset threshold, writing the planar vector index in the predetermined buffer to non-volatile memory, may include: in response to the data volume in the predetermined buffer reaching the preset threshold and the number of vectors under the planar vector index in the predetermined buffer being greater than a third predetermined number, creating a non-planar vector index based on the planar vector index in the predetermined buffer; writing the non-planar vector index to non-volatile memory, and clearing the predetermined buffer.

[0053] In this embodiment, when there are many vectors under the vector index in the predetermined buffer, a non-planar vector index can be created based on the planar vector index of the predetermined buffer, and written to non-volatile memory in the form of a non-planar vector index. This avoids the frequent IO operations caused by writing to disk in the form of a planar vector index in large-scale vector data scenarios, and the non-planar vector index can improve the query speed.

[0054] As an example, this disclosure can optimize the vector index update process by dividing the vector index data into incremental segments and sealed segments. When new vector index data needs to be written, this new vector index data is first stored in a predetermined buffer, which can be regarded as incremental segment data. When the amount of data in the predetermined buffer reaches a preset threshold, a non-planar vector index can be created based on the planar vector index in the predetermined buffer, which is equivalent to sealing the planar vector index in the predetermined buffer. This can be regarded as sealed segment data. Then, the created non-planar vector index is written to local non-volatile memory.

[0055] As an example, the aforementioned planar vector index can be FlatIndex, and the aforementioned non-planar vector index can be an inverted index, a product quantization index, etc. The specific type of non-planar vector index to be used depends on the number of vectors under the planar vector index, and this disclosure does not limit this.

[0056] According to an exemplary embodiment of this disclosure, the above-described method of determining a first predetermined number of vectors with the highest similarity to a query vector using vector indexes in a predetermined buffer and non-volatile memory, and using these vectors as vector retrieval results, may include: determining the similarity between each vector under the vector index in the predetermined buffer and the query vector using vector indexes in the predetermined buffer; determining the similarity between each vector under the vector index in the non-volatile memory and the query vector using vector indexes in the non-volatile memory; merging the vectors under the vector indexes in the predetermined buffer and non-volatile memory and sorting them in descending order of similarity, and determining the first predetermined number of vectors with the highest similarity as vector retrieval results.

[0057] As an example, assuming the first predetermined quantity is 10, the similarity between all vectors under a predetermined buffer in memory and the query vector can be calculated using the vector index, and the similarity between all vectors under a predetermined buffer and the query vector can be calculated using the vector index in non-volatile memory. Then, the similarity corresponding to the predetermined buffer and the similarity corresponding to the non-volatile memory are combined and sorted, and the top 10 vectors are selected. This disclosure does not limit this process.

[0058] It should be noted that the aforementioned first predetermined quantity can be set as needed, and this disclosure does not limit it.

[0059] According to an exemplary embodiment of this disclosure, the creation of a planar vector index for an external document may include: in response to the external document being a plain text document, dividing the text content of the external document into blocks using a sliding window approach to obtain multiple text blocks; in response to the external document being a structured document, dividing the text content of the external document into blocks based on the hierarchical information of the external document to obtain multiple text blocks, wherein each text block contains hierarchical information of all levels to which it belongs in the external document; vectorizing each text block to obtain a vector for each text block; and creating a planar vector index based on the vector of each text block. Through this embodiment, a structured method is used for parsing, that is, the document is divided into blocks according to hierarchy based on the structural information in the document, and each block includes hierarchical information of all levels to which it belongs in the document. This structural information can provide more information to the text blocks during block division, making the retrieval results more accurate.

[0060] Specifically, in existing RAG technology, the index building process typically involves directly segmenting the text according to a preset block size or recursively segmenting based on punctuation marks. A drawback of this common segmentation method is its inability to capture global information. For example, if a piece of text exceeds the block size and is divided into multiple blocks, and the semantics of the search are related to the title of that text, but only the first block contains the paragraph title information while the others do not, this leads to reduced search accuracy. To address this issue, this disclosure considers that structured documents are the most common among various document formats, such as docx, xslx, and markdown. Therefore, this disclosure adopts a structured method for parsing, that is, segmenting the text according to its structure information, with each block including hierarchical information about all levels it belongs to within the document, such as multi-level headings in docx and table information in xslx tables. This structural information can provide more information to the text blocks during segmentation, making the search results more accurate.

[0061] As an example, let's assume a document in docx format contains second-level headings, which might look like this: title1 (Level 1 Heading) title2 (Second-level heading) Abcdefghijklmnopqrst... (Main text) The block segmentation results obtained using ordinary segmentation methods are generally as follows: Section 1: title1\ntitle2\nabcdefg... Block 2: hijklmnopqrst... The block-based results obtained after parsing using a structured method in this disclosure are as follows: Section 1: title1\ntitle2\nabcdefg... Block 2: title1\ntitle2\nhijklmnopqrst... As can be seen from the above example, when searching with "title1", the segmentation results obtained by the ordinary segmentation method cannot retrieve segment 2, while the segmentation results obtained by this disclosure can retrieve segment 2.

[0062] As an example, when writing to an external document, the index creation interface can be called. In response to this call, the external document to be written is received, and its text content is extracted using a parser corresponding to the external document. Then, the text content is converted to Markdown format. Next, the converted text content is cleaned to obtain cleaned text content. After obtaining the cleaned text content, the text segmentation operation described in the previous example is performed.

[0063] To better understand the text segmentation process in this implementation, the following will combine... Figure 2 A systematic explanation of the text segmentation process is provided.

[0064] Figure 2 This demonstrates a flowchart of document parsing and text segmentation, such as... Figure 2 As shown, firstly, the document type is identified by the file extension, then the corresponding parser is called to extract the text content, and then the extracted text content is converted into a unified intermediate representation in Markdown format. Next, text cleanup is performed, such as removing invalid characters, and then intelligent chunking can be performed.

[0065] If the external document is a plain text document, an overlapping sliding window approach can be used for segmentation. First, a default window size and stride are set. The sliding window ensures semantic continuity. If the external document is a structured document, the system segments the text according to semantic boundaries and identifies heading levels. Contextual information is added to each text block, meaning each text block contains hierarchical information about all levels it belongs to within the external document. Therefore, the final generated text blocks maintain the semantic integrity of the original document while meeting the input requirements of the vectorized model, improving the accuracy of semantic retrieval.

[0066] To better understand the index creation process, the following will combine... Figure 3 The index creation process is explained systematically.

[0067] Figure 3 This demonstrates a vector index creation process, such as Figure 3 As shown, after the index creation interface is called, the document is first parsed and segmented, and the corresponding interface is called for vectorization. Then, vector data and metadata are stored separately. The generated metadata can be directly stored in a local database (such as SQL) using Structured Query Language (SQL). It should be noted that metadata can include the document's text data and metadata such as document name and document path. After the vector data is generated, a vector caching mechanism can be used for data storage. Specifically, a new vector index, i.e., a flat vector index, is created based on the vector data and stored in a predetermined buffer in memory. It is determined whether the data volume in the predetermined buffer is greater than a threshold. If the result is no, it continues to be cached in the predetermined buffer. If the result is yes, an index write-to-disk signal is triggered. After the index write-to-disk signal is triggered, metadata write-to-disk operations and vector index write-to-disk operations can be performed, that is, the vector index in the predetermined buffer is saved to a vector index file in non-volatile memory. Then, the result of index construction is returned to the index creation interface, which receives the processing result.

[0068] As an example, when metadata (i.e., non-vector data) is stored in a local database, the database table structure can be as follows: Table 1 Collection of Database Tables

[0069] Table 2 embedding_metadata

[0070] Table 3 embedding_status table

[0071] Table 4 index_segment table

[0072] As an example, Figure 4 The relationship between the three tables is shown, such as Figure 4 As shown, Tables 2 and 3 are constructed based on Table 1. Table 2 shows the creation status of the vector index of the document, and Table 3 shows the information of the vector index.

[0073] According to an exemplary embodiment of this disclosure, a list of files to be deleted is received, wherein the list of files to be deleted includes an identifier of at least one document to be deleted; the index creation status of at least one document to be deleted is queried; for a document to be deleted whose index creation status is successful, a predetermined vector index is queried from memory based on the identifier of the document to be deleted; in response to the existence of the predetermined vector index in memory, the predetermined vector index is deleted from memory; in response to the absence of the predetermined vector index in memory, the predetermined vector index is queried from non-volatile memory, and the deletion flag corresponding to the predetermined vector index is set to 1, wherein the predetermined vector index is the vector index of the document to be deleted; for a document to be deleted whose index creation status is in the index creation process, the index creation process of the document to be deleted is stopped; for a document to be deleted whose index creation status is failed, the document to be deleted is skipped.

[0074] In this embodiment, a delete bit is maintained for each vector index. When deleting a vector index, physical deletion is not required; instead, the corresponding delete bit is set to 1. Subsequent vector retrieval processes will automatically skip these deleted vectors, allowing the deletion operation to update only one bit in the database. This avoids frequent disk I / O operations and index rebuilding, improving the performance of the deletion operation. Moreover, the index file does not need to be modified, maintaining the integrity of the index file and avoiding index fragmentation. Batch deletion and recovery operations are also supported. For example, accidental deletion can be recovered by resetting the bit, thereby ensuring retrieval accuracy while significantly improving the efficiency of the deletion operation and the overall stability of the system.

[0075] As an example, Figure 5 This demonstrates a vector index deletion process, such as Figure 5As shown, after the index deletion interface is called, the system receives a list of files containing information about the documents to be deleted, i.e., the list of files to be deleted as specified in the diagram. Deleting a vector index first requires determining the current processing status of the document (i.e., the index creation status mentioned above). Vector indexes of external documents in the "in progress" (i.e., within the index creation process) or "error" (i.e., index creation error) status do not need to be deleted, as neither has successfully created a vector index. For external documents in the index creation process, simply stop the ongoing process by adding the vector ID to the abort list. For external documents whose index creation failed, skip them. For external documents whose index creation was successful, first query the vector ID using the passed document name. Then, delete the vector index data in the predefined buffer and non-volatile memory using the vector ID. Specifically, if the vector index data corresponding to the vector ID is in the predefined buffer, delete it directly from the predefined buffer. If it is not in the predefined buffer, delete it from non-volatile memory. Deletion in non-volatile memory can be achieved by setting the deletion flag to 1. Finally, the deletion result can be returned to the index deletion interface.

[0076] According to an exemplary embodiment of this disclosure, determining the similarity between each vector under a vector index in non-volatile memory and a query vector using vector indexes in non-volatile memory may include: constructing a bitmap based on deletion flag bits of the vector indexes in non-volatile memory; and using the bitmap and the vector indexes in non-volatile memory to determine the similarity between each vector under a vector index in non-volatile memory where the deletion flag bit is 0 and the query vector. Through this embodiment, during vector retrieval, a bitmap is constructed based on the deletion flag bits, thereby automatically skipping deleted vectors during the vector search process.

[0077] As an example, during vector retrieval, the system constructs a bitmap (BitSet bitmap) based on the deleted bits of all vector indices. Since the positions of all deleted vector indices are marked as 1, these deleted vectors can be automatically skipped during the search process based on this BitSet bitmap.

[0078] It should be noted that the constructed BitSet bitmap can be stored in memory, and this disclosure does not limit this.

[0079] return Figure 1 In step 103, the text block corresponding to the vector retrieval result is obtained from the non-volatile memory.

[0080] In step 104, the text blocks corresponding to the vector retrieval results and the query questions are segmented into words.

[0081] In step 105, a keyword index is created based on the word segmentation results of the text blocks corresponding to the vector retrieval results.

[0082] In step 106, keyword retrieval is performed based on the word segmentation results and keyword index of the query question to obtain a second predetermined number of text blocks with the highest matching degree to the query question. The second predetermined number of text blocks are used to assist the local large language model in generating answers to the query question.

[0083] Specifically, the vector retrieval results mentioned above are obtained by calculating vector similarity, that is, by ranking the vectors in the vector space generated from the text information with the query vector, and selecting the vectors that rank higher. However, this method of obtaining vector retrieval results is highly dependent on the capability of the vectorization model. If the model capability is poor, it will lead to a decrease in retrieval accuracy. Therefore, in order to solve this problem, this embodiment proposes a keyword re-ranking algorithm.

[0084] As an example, during vector retrieval, a larger number of vectors are first acquired for backup, that is, twice the top K vectors (i.e., 2topK) are acquired as vector retrieval results. Then, based on the vector IDs of the vectors in the vector retrieval results, the corresponding source text information (i.e., the text blocks corresponding to the vector retrieval results) is retrieved from a pre-defined buffer and non-volatile memory. Next, a word segmentation algorithm is used to segment both the query question and the source text information. The word segmentation results of the source text information are used to create a keyword index. The word segmentation results of the query question are used to search in the keyword index. The search results can select the top K text blocks, i.e., the keyword search results, which serve as the retrieval results for the entire vector retrieval process. Finally, the required number of results can be returned based on user needs, i.e., selecting the required amount of text blocks from the top K results as the retrieval results for the query question, to assist the large language model in generating the answer to the query question.

[0085] It should be noted that the generated keyword index can be stored in memory, and this disclosure does not impose any restrictions on this.

[0086] To better understand the retrieval method of this disclosure, the following will be combined with... Figure 6 and Figure 7 Provide a systematic explanation.

[0087] Figure 6 This demonstrates a vector retrieval process, such as Figure 6 As shown, the vector retrieval process is a multi-layered and efficient semantic search system that employs a hybrid retrieval strategy combining vector similarity search and keyword re-ranking. The entire process is divided into five core stages: query preprocessing, vector search, result merging and sorting, metadata query, and keyword re-ranking.

[0088] First, a query request is received through a vector retrieval interface. This request can contain the user's query text (i.e., the query question mentioned above) and the expected number of results to be returned (e.g., top K). Second, the user's query text is vectorized. Third, a vector search is performed based on the vector of the query question. Specifically, a buffer is reserved in memory (e.g., ...). Figure 6 (caching in the middle) and non-volatile storage (such as...) Figure 6 The search is performed within a vector index on the disk (in the system's memory). The search process uses Euclidean distance to calculate similarity, selecting the top-similar vectors (e.g., the top 2K vectors). During the vector search, the system uses a BitSet mechanism to filter out deleted vectors, ensuring the validity of the search results. Next, the cached search results and the disk search results are merged to obtain the vector search results. Finally, the result ranking is optimized using a keyword re-sorting algorithm to provide more accurate semantic matching. Specifically, within a predetermined buffer (e.g., ... Figure 6 (in cache) and non-volatile memory (such as cache) Figure 6 The source text information of the vector retrieval results is obtained from the database in the database, and the keywords are rearranged based on the source text information.

[0089] Figure 7 A system architecture diagram of a local vector database is shown, such as Figure 7 As shown, the local vector database, as a system service, exposes vector index creation, vector retrieval, and index deletion interfaces. This means the local vector database can provide vector index construction, deletion, and retrieval functions through the interface service module. The data processing module implements document parsing, text segmentation, and text vectorization during vector index creation, as well as text vectorization and vector similarity calculation during vector retrieval. The data management module manages vector data and metadata, such as creating and deleting vector indexes, and storing and deleting metadata. Vector files can contain cached vector index data (i.e., the vector index data in the predetermined buffer in the above embodiment) and locally persisted vector files. The locally persisted vector files contain the vector index data already persisted to disk, while metadata management directly operates on the local database.

[0090] In summary, the cloud-based vector database solutions employed in this embodiment have limitations in terms of real-time performance, privacy protection, and cost control. For example, Milvus is a cloud-native, distributed vector database designed for near-nearest neighbor search of massive vectors. While Milvus, as a cloud-native distributed vector database, offers advantages such as horizontal scalability, high availability, and millisecond-level retrieval of hundreds of millions of vectors, making it suitable for large-scale online services, its disadvantages include deployment dependencies on components like etcd, Pulsar, and object storage, resulting in high resource consumption and complex maintenance in local development or edge scenarios. In contrast, the local vector database solution in this embodiment offers one-click startup, zero dependencies, and low resource consumption, making it suitable for single-machine deployment and overcoming the aforementioned drawbacks of cloud-based vector database solutions.

[0091] Specifically, this embodiment proposes a highly efficient and lightweight local vector database architecture design, aiming to address the technical limitations of existing cloud-based vector database solutions in terms of real-time performance, privacy protection, and cost control. Specifically, by employing a local storage system for vector storage and retrieval, it avoids network dependence on cloud solutions, eliminates the impact of network latency on real-time performance, and significantly improves real-time capabilities. Furthermore, the vector data storage and retrieval process in this embodiment is entirely performed locally, avoiding the security risks of transmitting sensitive data to the cloud. This ensures the complete privacy of users' personal data and sensitive information, reducing the risk of data leakage or privacy violations, and meeting stringent data privacy protection requirements. Therefore, this embodiment strengthens privacy protection and ensures data security, making it particularly suitable for fields with high privacy requirements such as healthcare and finance. Moreover, this localized solution avoids the storage fees and high costs of cloud services and pay-per-access billing, allowing users to avoid paying exorbitant fees for frequent database accesses. It also reduces reliance on external services, effectively lowering operational costs and improving resource utilization. This is particularly suitable for enterprises and individuals with limited budgets or who are cost-sensitive, enabling efficient vector data management and retrieval with limited resources.

[0092] Figure 8 This is a block diagram illustrating a retrieval device according to an exemplary embodiment. (Refer to...) Figure 8 The device includes a first acquisition unit 80, a first retrieval unit 82, a second acquisition unit 84, a word segmentation unit 86, a creation unit 88, and a second retrieval unit 810.

[0093] The first acquisition unit 80 is configured to acquire the query vector of the query question; the first retrieval unit 82 is configured to use the vector index in the local non-volatile memory to determine a first predetermined number of vectors with the highest similarity to the query vector, and use them as vector retrieval results, wherein the non-volatile memory is used to store the vector index of the external document and the text block corresponding to each vector under the vector index, and the text block is obtained by dividing the text content of the external document into blocks; the second acquisition unit 84 is configured to acquire the text block corresponding to the vector retrieval result from the non-volatile memory; the word segmentation unit 86 is configured to perform word segmentation on the text block corresponding to the vector retrieval result and the query question; the creation unit 88 is configured to create a keyword index based on the word segmentation result of the text block corresponding to the vector retrieval result; the second retrieval unit 810 is configured to perform keyword retrieval based on the word segmentation result of the query question and the keyword index, and obtain a second predetermined number of text blocks with the highest matching degree to the query question, wherein the second predetermined number of text blocks are used to assist the local large language model in generating an answer to the query question.

[0094] According to an exemplary embodiment of this disclosure, the first retrieval unit 82 is further configured to: create a planar vector index of an external document before determining a first predetermined number of vectors with high similarity to a query vector using a vector index in local non-volatile memory and using them as vector retrieval results; write the planar vector index into a predetermined buffer in local memory, wherein the predetermined buffer is a region in memory used to store the planar vector index of the external document; write the planar vector index in the predetermined buffer into non-volatile memory in response to the amount of data in the predetermined buffer reaching a preset threshold; and determine a first predetermined number of vectors with high similarity to a query vector using the vector index in the predetermined buffer and non-volatile memory and use them as vector retrieval results.

[0095] According to an exemplary embodiment of the present disclosure, the first retrieval unit 82 is further configured to, in response to the amount of data in the predetermined buffer reaching a preset threshold and the number of vectors under the planar vector index in the predetermined buffer being greater than a third predetermined number, create a non-planar vector index based on the planar vector index in the predetermined buffer; write the non-planar vector index to a non-volatile memory; and clear the predetermined buffer.

[0096] According to an exemplary embodiment of this disclosure, the first retrieval unit 82 is further configured to: respond to the external document being a plain text document, divide the text content of the external document into blocks using a sliding window method to obtain multiple text blocks; respond to the external document being a structured document, divide the text content of the external document into blocks based on the hierarchical information of the external document to obtain multiple text blocks, wherein each text block contains hierarchical information of all levels to which it belongs in the external document; vectorize each text block to obtain a vector for each text block; and create a planar vector index based on the vector of each text block.

[0097] According to an exemplary embodiment of this disclosure, the apparatus further includes a deletion unit configured to receive a list of deleted files, wherein the list of deleted files includes an identifier of at least one document to be deleted; query the index creation status of at least one document to be deleted; for a document to be deleted whose index creation status is successful, query a predetermined vector index from memory based on the identifier of the document to be deleted; in response to the existence of the predetermined vector index in memory, delete the predetermined vector index from memory; in response to the absence of the predetermined vector index in memory, query the predetermined vector index from non-volatile memory and set the deletion flag corresponding to the predetermined vector index to 1, wherein the predetermined vector index is the vector index of the document to be deleted; for a document to be deleted whose index creation status is in the index creation process, stop the index creation process of the document to be deleted; and for a document to be deleted whose index creation status is failed, skip the document to be deleted.

[0098] According to an exemplary embodiment of this disclosure, the first retrieval unit 82 is further configured to: utilize the vector index in the predetermined buffer of memory to determine the similarity between each vector under the vector index in the predetermined buffer and the query vector; utilize the vector index in non-volatile memory to determine the similarity between each vector under the vector index in non-volatile memory and the query vector; merge the vectors under the vector index in the predetermined buffer and non-volatile memory and sort them in descending order of similarity; and take the first predetermined number of vectors with the highest similarity as the vector retrieval result.

[0099] According to an exemplary embodiment of this disclosure, the first retrieval unit 82 is further configured to construct a bitmap based on the deletion flag bit of the vector index in the non-volatile memory; and to determine the similarity between each vector under the vector index in the non-volatile memory where the deletion flag bit is 0 and the query vector using the bitmap and the vector index in the non-volatile memory.

[0100] According to embodiments of this disclosure, an electronic device may be provided. Figure 9This is a block diagram of an electronic device 900 according to an embodiment of the present disclosure. The electronic device includes at least one memory 901 and at least one processor 902. The at least one memory stores a set of computer-executable instructions. When the set of computer-executable instructions is executed by the at least one processor, a retrieval method according to an embodiment of the present disclosure is performed.

[0101] As an example, electronic device 900 may be a PC, tablet, personal digital assistant, smartphone, or other device capable of executing the aforementioned set of instructions. Here, electronic device 900 is not necessarily a single electronic device, but may be any collection of devices or circuits capable of executing the aforementioned instructions (or instruction sets) individually or in combination. Electronic device 900 may also be part of an integrated control system or system manager, or may be configured to interconnect with a portable electronic device locally or remotely (e.g., via wireless transmission) through an interface.

[0102] In electronic device 900, processor 902 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, processor 902 may also include analog processors, digital processors, microprocessors, multi-core processors, processor arrays, network processors, etc.

[0103] The processor 902 can execute instructions or code stored in memory, wherein memory 901 can also store data. Instructions and data can also be sent and received via a network through a network interface device, wherein the network interface device can employ any known transmission protocol.

[0104] The memory 901 can be integrated with the processor 902, for example, by placing RAM or flash memory within an integrated circuit microprocessor. Alternatively, the memory 901 can include a separate device, such as an external disk drive, a storage array, or other storage device usable by any database system. The memory 901 and the processor 902 can be operatively coupled, or can communicate with each other, for example, via I / O ports, network connections, etc., enabling the processor 902 to read files stored in the memory 901.

[0105] In addition, the electronic device 900 may also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, mouse, touch input device, etc.). All components of the electronic device can be interconnected via a bus and / or network.

[0106] According to embodiments of this disclosure, a computer-readable storage medium may also be provided, wherein when instructions in the computer-readable storage medium are executed by at least one processor, the at least one processor causes to perform the retrieval method of embodiments of this disclosure. Examples of computer-readable storage media include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc storage, hard disk drive (HDD), solid-state drive (SSD), card storage (such as multimedia cards, secure digital (SD) cards, or ultra-fast digital (XD) cards), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid-state drive, and any other device configured to store a computer program and any associated data, data files, and data structures in a non-transitory manner and to provide the computer program and any associated data, data files, and data structures to a processor or computer so that the processor or computer can execute the computer program. The computer program in the aforementioned computer-readable storage medium can run in an environment deployed in computer devices such as clients, hosts, agent devices, servers, etc. Furthermore, in one example, the computer program and any associated data, data files, and data structures are distributed across a networked computer system, such that the computer program and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner through one or more processors or computers.

[0107] According to an embodiment of this disclosure, a computer program product is provided, including computer instructions, which, when executed by a processor, implement the retrieval method of this disclosure.

[0108] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the appended claims.

[0109] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A retrieval method, characterized in that, include: Obtain the query vector for the query question; Using the vector index in the local non-volatile memory, a first predetermined number of vectors with the highest similarity to the query vector are determined as vector retrieval results. The non-volatile memory is used to store the vector index of the external document and the text block corresponding to each vector under the vector index. The text block is obtained by dividing the text content of the external document into blocks. Retrieve the text block corresponding to the vector retrieval result from the non-volatile memory; The text blocks corresponding to the vector retrieval results and the query question are segmented into words; A keyword index is created based on the word segmentation results of the text blocks corresponding to the vector retrieval results; Based on the word segmentation results of the query question and the keyword index, keyword retrieval is performed to obtain a second predetermined number of text blocks with the highest matching degree to the query question. The second predetermined number of text blocks are used to assist the local large language model in generating an answer to the query question.

2. The retrieval method as described in claim 1, characterized in that, Before using the vector index in local non-volatile memory to determine a first predetermined number of vectors with high similarity to the query vector as vector retrieval results, the method further includes: Create a planar vector index for the external document; The planar vector index is written to a predetermined buffer in local memory, wherein the predetermined buffer is a region in the memory used to store the planar vector index of the external document; In response to the data volume in the predetermined buffer reaching a preset threshold, the plane vector index in the predetermined buffer is written into the non-volatile memory; The step of using the vector index in the local non-volatile memory to determine a first predetermined number of vectors with the highest similarity to the query vector as vector retrieval results includes: using the predetermined buffer and the vector index in the non-volatile memory to determine the first predetermined number of vectors with the highest similarity to the query vector as vector retrieval results.

3. The retrieval method as described in claim 2, characterized in that, The step of writing the plane vector index in the predetermined buffer to the non-volatile memory in response to the data volume in the predetermined buffer reaching a preset threshold includes: In response to the data volume in the predetermined buffer reaching a preset threshold and the number of vectors under the planar vector index in the predetermined buffer being greater than a third predetermined number, a non-planar vector index is created based on the planar vector index in the predetermined buffer; Write the non-planar vector index into the non-volatile memory and clear the predetermined buffer.

4. The retrieval method as described in claim 2, characterized in that, Creating the planar vector index of the external document includes: In response to the fact that the external document is a plain text document, the text content of the external document is divided into blocks using a sliding window method to obtain multiple text blocks; In response to the fact that the external document is a structured document, the text content of the external document is divided into blocks based on the hierarchical information of the external document to obtain multiple text blocks, wherein each text block contains hierarchical information of all levels to which it belongs in the external document; Vectorize each text block separately to obtain the vector of each text block; The planar vector index is created based on the vector of each text block.

5. The retrieval method as described in claim 2, characterized in that, Also includes: Receive a list of files to be deleted, wherein the list of files to be deleted includes the identifier of at least one document to be deleted; Query the index creation status of at least one document to be deleted; For a document to be deleted whose index creation status is "index creation successful", a predetermined vector index is queried from the memory based on the identifier of the document to be deleted; in response to the existence of the predetermined vector index in the memory, the predetermined vector index is deleted from the memory; in response to the absence of the predetermined vector index in the memory, the predetermined vector index is queried from the non-volatile memory, and the deletion flag corresponding to the predetermined vector index is set to 1, wherein the predetermined vector index is the vector index of the document to be deleted; For documents to be deleted whose index creation status is in the index creation process, stop the index creation process for the documents to be deleted; For documents to be deleted whose index creation status is "index creation failed", skip the documents to be deleted.

6. The retrieval method as described in claim 2, characterized in that, The step of using the predetermined buffer and the vector index in the non-volatile memory to determine a first predetermined number of vectors with high similarity to the query vector as vector retrieval results includes: Using the vector indices in the predetermined buffer of the memory, determine the similarity between each vector under the vector index in the predetermined buffer and the query vector; Using the vector index in the non-volatile memory, determine the similarity between each vector under the vector index in the non-volatile memory and the query vector; The vectors under the vector indices in the predetermined buffer and the non-volatile memory are merged and sorted in descending order of similarity. The first predetermined number of vectors with the highest similarity are determined as the vector retrieval results.

7. The retrieval method as described in claim 6, characterized in that, The step of determining the similarity between each vector under the vector index in the non-volatile memory and the query vector, using the vector index in the non-volatile memory, includes: A bitmap is constructed based on the deletion flag bits of the vector index in the non-volatile memory; Using the bitmap and the vector indexes in the non-volatile memory, the similarity between each vector under the vector index with the marker bit set to 0 in the non-volatile memory and the query vector is determined.

8. A retrieval device, characterized in that, include: The first acquisition unit is configured to acquire the query vector of the query question; The first retrieval unit is configured to use a vector index in a local non-volatile memory to determine a first predetermined number of vectors with the highest similarity to the query vector as vector retrieval results. The non-volatile memory is used to store the vector index of an external document and the text block corresponding to each vector under the vector index. The text block is obtained by dividing the text content of the external document into blocks. The second acquisition unit is configured to acquire the text block corresponding to the vector retrieval result from the non-volatile memory; The word segmentation unit is configured to segment the text block corresponding to the vector retrieval result and the query question into words. The creation unit is configured to create a keyword index based on the word segmentation results of the text blocks corresponding to the vector retrieval results; The second retrieval unit is configured to perform keyword retrieval based on the word segmentation results of the query question and the keyword index, and obtain a second predetermined number of text blocks with the highest matching degree to the query question. The second predetermined number of text blocks are used to assist the local large language model in generating an answer to the query question.

9. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the retrieval method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by at least one processor, the at least one processor causes the processor to perform the retrieval method as described in any one of claims 1 to 7.

11. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the retrieval method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Cloud storage based power full text retrieval method and system

    CN102156711A

  • Multi-document intelligent question and answer method and system based on large language model

    CN118394897A

  • Question and answer processing method and device, electronic equipment and storage medium

    CN119202151A

  • Large language model knowledge retrieval method and system based on semantic vectorization

    CN119336864A

  • Text block retrieval method, computer program product and equipment

    CN120470090A