Optimising vector embedding for natural language processing

The method optimizes vector embedding for natural language processing by incrementally building a vector database and dynamically incorporating new data, addressing inefficiencies and dynamic data handling challenges in current methods.

WO2025119443A1PCT designated stage expired Publication Date: 2025-06-12HUAWEI TECH CO LTD +1

Patent Information

Application Number
PCT/EP2023/084113
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-04
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Current vector embedding methods for natural language processing, particularly in RAG-related scenarios, are inefficient due to the need to scan the entire dataset for vector embeddings and inability to handle dynamically generated data.

Method used

The proposed method optimizes vector embedding by incrementally building a vector database, generating vector embeddings only for relevant data files as needed, and dynamically incorporating new data, thereby reducing computational overhead and supporting dynamic data generation.

Benefits of technology

This approach significantly reduces the computational burden of vector embedding, allows for efficient handling of dynamically generated data, and ensures a complete vector database by embedding relevant documents incrementally.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2023084113_12062025_PF_FP_ABST
    Figure EP2023084113_12062025_PF_FP_ABST
Patent Text Reader

Abstract

In some examples, a method of optimising vector embedding for natural language processing comprises receiving a query from a user, converting the received query into a set of search criteria, determining, using the search criteria, a set of data files from multiple data files, wherein the set of data files comprises a first data file, determining whether a vector embedding associated with the first data file exists, and, in response to determining that a vector embedding associated with the first data file does not exist, generating a vector embedding associated with the first data file and adding the vector embedding associated with the first data file to a vector database.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] OPTIMISING VECTOR EMBEDDING FOR NATURAL LANGUAGE PROCESSING

[0002] TECHNICAL FIELD

[0003] The present disclosure relates, in general, to optimising vector embedding for natural language processing. Aspects of the disclosure relate to optimising the initial step of vector embedding for RAG-related scenarios.

[0004] BACKGROUND

[0005] General-purpose language models can be fine-tuned to achieve several common tasks, such as sentiment analysis and named entity recognition. Generally, these tasks do not require the language model to acquire additional background knowledge.

[0006] For more complex and knowledge-intensive tasks, it is possible to build a language modelbased system that accesses external knowledge sources to complete tasks. This enables more factual consistency, improves reliability of the generated responses, and helps to mitigate the problem of "hallucination".

[0007] Artificial intelligence (Al) researchers introduced a method called Retrieval Augmented Generation (RAG) to address such knowledge-intensive tasks. RAG combines an information retrieval component with a text generator model. RAG can be fine-tuned and its internal knowledge can be modified in an efficient manner and without needing retraining of the entire model.

[0008] RAG takes an input and retrieves a set of relevant / supporting documents given a source (e.g., Wikipedia). The documents are concatenated as context with the original input prompt and fed to the text generator which produces the final output. This makes RAG adaptive for situations where facts could evolve over time. This is very useful as LLMs's parametric knowledge is static. RAG allows language models to bypass retraining, enabling access to the latest information for generating reliable outputs via retrieval-based generation.

[0009] On the first step, RAG leverages the vector search approach to find out the top-K documents. Vector search is an advanced approach to data retrieval used in machine learning and generative Al, that focuses on semantic meaning and similarity rather than specific keywords. By representing data as vectors in a high-dimensional space, it enables more accurate and intuitive search results. These set of vectors are created using a vector embedding algorithm and stored in a special vectors database.

[0010] There are several methods for creating the vector embeddings. The common factor for these methods is the long and heavy process required to scan the entire data set of documents / objects. That is, before running RAG, it is required to scan the entire data set, generate a vector embedding representation for each object and add the vector into a vectors database.

[0011] Another disadvantage of the current solutions is that they are not able to handle cases in which the relevant data (i.e., Storage Catalog) is generated dynamically and therefore the ingest stage could not be done prior to the generation query stage.

[0012] SUMMARY

[0013] An objective of the present disclosure is to optimise the step of vector embedding for natural language processing.

[0014] The foregoing and other objectives are achieved by the features of the independent claims.

[0015] Further implementation forms are apparent from the dependent claims, the description and the Figures.

[0016] A first aspect of the present disclosure provides a method of optimising vector embedding for natural language processing, the method comprising receiving a query from a user, converting the received query into a set of search criteria, determining, using the search criteria, a set of data files from multiple data files, wherein the set of data files comprises a first data file, determining whether a vector embedding associated with the first data file exists, and, in response to determining that a vector embedding associated with the first data file does not exist, generating a vector embedding associated with the first data file and adding the vector embedding associated with the first data file to a vector database.

[0017] Accordingly, an efficient and fast method of vector embeddings can be provided. By using this technique, the long and computational-heavy phase of vector embeddings for the entire dataset is broken into many small steps, such that the vectors database is built incrementally. In addition, the invention overcomes the potential problem of having an incomplete vectors database. Another benefit of the proposed solution is the ability to take into account data which is generated dynamically during the system operation (i.e., information about new backup file copies).

[0018] The method may further comprise marking the first data file as embedded, whereby to indicate that the vector embedding associated with the first data file exists.

[0019] The set of data files may further comprise a second data file, and the method may further comprise, in response to determining that the vector embedding associated with the first data file exists, determining whether a vector embedding associated with the second data file exists.

[0020] Converting the received query into the set of search criteria may comprise converting the received query into the set of search criteria using a large language model.

[0021] Determining, using the search criteria, the set of data files from the multiple data files may comprise querying a database comprising multiple data files, and selecting the set of data files based on a similarity between the set of search criteria and each of the multiple data files.

[0022] The method may further comprise determining whether similarity between the set of search criteria and the selected set of data files exceeds a first threshold value, and, in response to determining that the similarity between the set of search criteria and the selected set of data files is below the first threshold value, querying the database comprising multiple files to add additional data files of the multiple data files to the set of data files.

[0023] The method may further comprise generating an answer to the query received from the user based on the vector database.

[0024] The method may further comprise determining whether a similarity between the set of search criteria and the generated answer exceeds a second threshold value, and, in response to determining that the similarity between the set of search criteria and the generated is below the second threshold value, querying the database comprising multiple files to add additional data files of the multiple data files to the set of data files. A second aspect of the present disclosure provides a computer readable storage medium comprising computer program code, accessible by an apparatus comprising a processor, to provide instructions and / or data to the apparatus, the computer program code configured to, with the processor, cause the apparatus to receive a query from a user, convert the received query into a set of search criteria, determine, using the search criteria, a set of data files from multiple data files, wherein the set of data files comprises a first data file, determine whether a vector embedding associated with the first data file exists, and, in response to determining that a vector embedding associated with the first data file does not exist, generate a vector embedding associated with the first data file and add the vector embedding associated with the first data file to a vector database.

[0025] The computer readable storage medium may further comprise the computer program code configured to, with the processor, cause the apparatus to mark the first data file as embedded, whereby to indicate that the vector embedding associated with the first data file exists.

[0026] The set of data files may further comprise a second data file, and the computer readable storage medium may further comprise the computer program code configured to, with the processor, cause the apparatus to, in response to determining that the vector embedding associated with the first data file exists, determine whether a vector embedding associated with the second data file exists.

[0027] The computer code configured to, with the processor, cause the apparatus to determine, using the search criteria, a set of data files from the multiple data files may comprise computer code configured to, with the processor, cause the apparatus to query a database comprising multiple data files, and select the set of data files based on similarity between the set of search criteria and each of the multiple data files.

[0028] The computer readable storage medium may further comprise the computer program code configured to, with the processor, cause the apparatus to determine whether a similarity between the set of search criteria and the selected set of data files exceeds a first threshold value, and, in response to determining that the similarity between the set of search criteria and the selected set of data files is below the first threshold value, query the database comprising multiple files to add additional data files of the multiple data files to the set of data files. The computer readable storage medium may further comprise the computer program code configured to, with the processor, cause the apparatus to generate an answer to the query received from the user based on the vector database.

[0029] The computer readable storage medium may further comprise the computer program code configured to, with the processor, cause the apparatus to determine whether similarity between the set of search criteria and the generated answer exceeds a second threshold value, and, in response to determining that the similarity between the set of search criteria and the generated is below the second threshold value, query the database comprising multiple files to add additional data files of the multiple data files to the set of data files.

[0030] A third aspect of the present disclosure provides a computing system, comprising a processor, a memory coupled to the processor, configured to store program code executable by the processor, the program code comprising one or more instructions to cause the computing system to receive a query from a user, convert the received query into a set of search criteria, determine, using the search criteria, a set of data files from multiple data files, wherein the set of data files comprises a first data file, determine whether a vector embedding associated with the first data file exists, and, in response to determining that a vector embedding associated with the first data file does not exist, generate a vector embedding associated with the first data file and add the vector embedding associated with the first data file to a vector database.

[0031] These and other aspects of the invention will be apparent from the embodiment s) described below.

[0032] BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order that the present invention may be more readily understood, embodiments of the invention will now be described, by way of example, with reference to the accompanying drawings, in which:

[0034] Figure 1 is a schematic representation of a retrieval augmented generation (RAG) algorithm according to the prior art;

[0035] Figure 2 is a schematic representation of a retrieval augmented generation (RAG) algorithm according to an example;

[0036] Figure 3 is a flow chart of a method for optimising vector embedding for natural language processing according to an example; Figure 4 is a schematic representation of a catalogue pre-filter steps according to an example;

[0037] Figure 5 is a schematic representation of a catalogue pre-filter with a recursive process according to an example; and

[0038] Figure 6 is a schematic of a computing system according to an example.

[0039] DETAILED DESCRIPTION

[0040] Example embodiments are described below in sufficient detail to enable those of ordinary skill in the art to embody and implement the systems and processes herein described. It is important to understand that embodiments can be provided in many alternate forms and should not be construed as limited to the examples set forth herein.

[0041] Accordingly, while embodiments can be modified in various ways and take on various alternative forms, specific embodiments thereof are shown in the drawings and described in detail below as examples. There is no intent to limit to the particular forms disclosed. On the contrary, all modifications, equivalents, and alternatives falling within the scope of the appended claims should be included. Elements of the example embodiments are consistently denoted by the same reference numerals throughout the drawings and detailed description where appropriate.

[0042] The terminology used herein to describe embodiments is not intended to limit the scope. The articles “a,” “an,” and “the” are singular in that they have a single referent, however the use of the singular form in the present document should not preclude the presence of more than one referent. In other words, elements referred to in the singular can number one or more, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises,” “comprising,” “includes,” and / or “including,” when used herein, specify the presence of stated features, items, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, items, steps, operations, elements, components, and / or groups thereof.

[0043] Unless otherwise defined, all terms (including technical and scientific terms) used herein are to be interpreted as is customary in the art. It will be further understood that terms in common usage should also be interpreted as is customary in the relevant art and not in an idealized or overly formal sense unless expressly so defined herein. Seq2seq model is a family of machine learning approaches used for natural language processing (NLP). The seq2seq takes as an input a sequence of words (a sentence or multiple sentences) and generates an output sequence of words. It does so through the use of a recurrent neural network (RNN). A recurrent neural network is one of the two broad types of artificial neural networks, characterised by the direction of the flow of information between its layers. In contrast to a uni-direction feedforward neural network, it is a bi-directional artificial neural network, meaning that it allows the output from some nodes to affect the subsequent input to other nodes.

[0044] Figure 1 is a schematic representation of a retrieval augmented generation (RAG) algorithm according to the prior art. The RAG algorithm 100 combines a pre-trained retriever (a query encoder 101 and a documents index 102) with a pre-trained seq2seq model (generator 103). For a query (x) 104, the RAG algorithm uses a Maximum Inner Product Search (MIPS) 105 to find the top-K documents Zi, using the output 121 of the query encoder 101. For the final prediction (y) 106, the algorithm treats z as a latent variable and marginalises over the seq2seq predictions given different documents. In general, the RAG algorithm can be described as having two stages:

[0045] 1. Ingestion

[0046] - Load documents (using a document loader),

[0047] Split documents (using a text splitter),

[0048] Create embeddings for documents (using a text embedding model), Store documents.

[0049] 2. Generation

[0050] - Receive a user question,

[0051] - Look up the documents in the index relevant to the question,

[0052] Construct a PromptValue from the question and any relevant documents (using a PromptTemplate),

[0053] - Pass the PromptValue to a model,

[0054] Acquire the result and return the result to the user.

[0055] Unfortunately, RAG is associated with the previously described problems, namely the length and the complexity of the process of scanning the entire data set of documents / objects and the fact that it does not support cases in which the relevant data (i.e., the storage catalogue) is generated dynamically.

[0056] According to an example, there is provided a mechanism for optimisation of the initial step of vector embedding for RAG-related scenarios without the need to wait for the vector embeddings of the entire data-set. In particular, the invention is concerned with adding an incremental step into the query phase, called a pre-filter, in which the system scans a huge metadata index (e.g., the storage catalogue which contains all the backup files, including backup versions) and selects only the top-M relevant documents (M»K) for vector embeddings. Vector embeddings are numerical representations of words and / or sentences. These may take form of an array of numbers, or similar. In order to perform a semantic search, a vector database may compute a vector embedding for each data object as it is inserted and / or updated in the database. These embeddings may be placed into an index, so as to enable the database searches to be performed quickly.

[0057] The selected documents are embedded (if needed) and are added incrementally into the vectors database (if they have not been embedded yet). By using this technique, the long and computational-heavy phase of vector embeddings for the entire data-set is broken into many small steps, such that the vectors database is built incrementally. In addition, the invention overcomes the potential problem of having an incomplete vectors database. It does so through the use of a special pre-filter search in the storage catalogue in which the top-M retrieved documents are most likely to cover the user’ s question. Another benefit of the proposed solution is the ability to take into account data which is generated dynamically during the system operation (i.e., information about new backup file copies).

[0058] Examples in the present disclosure can be provided as methods, systems or machine-readable instructions, such as any combination of software, hardware, firmware or the like. Such machine-readable instructions may be included on a computer readable storage medium (including but not limited to disc storage, CD-ROM, optical storage, etc.) having computer readable program codes therein or thereon.

[0059] The present disclosure is described with reference to flow charts and / or block diagrams of the method, devices and systems according to examples of the present disclosure. Although the flow diagrams described above show a specific order of execution, the order of execution may differ from that which is depicted. Blocks described in relation to one flow chart may be combined with those of another flow chart. In some examples, some blocks of the flow diagrams may not be necessary and / or additional blocks may be added. It shall be understood that each flow and / or block in the flow charts and / or block diagrams, as well as combinations of the flows and / or diagrams in the flow charts and / or block diagrams can be realized by machine readable instructions.

[0060] The machine-readable instructions may, for example, be executed by a machine such as a general-purpose computer, user equipment such as a smart device, e.g., a smart phone, a special purpose computer, an embedded processor or processors of other programmable data processing devices to realize the functions described in the description and diagrams. In particular, a processor or processing apparatus may execute the machine-readable instructions. Thus, modules of apparatus (for example, a module implementing a comparator unit, or a firewall structure and so on) may be implemented by a processor executing machine readable instructions stored in a memory, or a processor operating in accordance with instructions embedded in logic circuitry. The term 'processor' is to be interpreted broadly to include a CPU, processing unit, ASIC, logic unit, or programmable gate set etc. The methods and modules may all be performed by a single processor or divided amongst several processors.

[0061] Such machine-readable instructions may also be stored in a computer readable storage that can guide the computer or other programmable data processing devices to operate in a specific mode. For example, the instructions may be provided on a non-transitory computer readable storage medium encoded with instructions, executable by a processor.

[0062] Figure 2 is a schematic representation of a retrieval augmented generation (RAG) algorithm according to an example. Compared to the RAG algorithm according to the prior art, depicted in Figure 1, in Figure 2 an additional stage is present. In particular, the RAG algorithm 200 - in addition to the query encoder 201, documents index 202, generator 203, and the retriever 205, corresponding to the elements of Figure 1 - additionally comprises a pre-filter 207. Similarly to the RAG algorithm 100, the RAG algorithm 200 takes as an input a user query 204 and returns an answer 206. To aid understanding of the pre-filter 307, reference is made to Figure 3, which is a flow chart of a method for optimising vector embedding for natural language processing according to an example. Figures 3 and 4 (and the accompanying description) provide more detail regarding the pre-filter algorithm. Figure 3 is a flow chart of a method for optimising vector embedding for natural language processing according to an example. In block 301, the method comprises receiving a query from a user. The query (which may also be referred to as a question) may comprise any text-based query, for example, “what is the weather like today in Boston?” or “what are the top-N candidate files for storage tiering?”.

[0063] In block 302, the received query is converted into a set of search criteria. Converting the received query into the set of search criteria may comprise converting the received query into the set of search criteria using a large language model (LLM). For example, the user’s query / question may be converted into a set of catalogue (e.g., metadata index and / or a database) search criteria. The search may comprise an SQL search or a No-SQL search.

[0064] In block 303, the method comprises determining, using the search criteria, a set of data files from multiple data files, wherein the set of data files comprises a first data file. Here, the term ‘first data file’ refers to a document of the set of documents (i.e., a document of a set of top-M documents) obtained using the search criteria from the database / catalogue. The set of data files may be selected based on a similarity between the set of search criteria (i.e., the converted user query) and each of the multiple data files. In other words, using the search criteria (obtained by converting the user’s query), a fast search may be executed in the catalogue to find the top-M documents. Importantly, M may comprise a number much larger than K (discussed below - M » K), in order to ensure that enough vectors are present for the next step of the method to be performed.

[0065] The method comprises, in block 304, determining whether a vector embedding associated with the first data file exists. In other words, at this stage, the method checks whether the first data file of the multiple data files (i.e., a document of the set of top-M documents) has already been embedded - that is, whether there already exists a vector embedding associated with this particular file. For each document in the results set (i.e., for each document of the multiple data files), if the document is already marked as embedded, the method may jump to the next document in the results.

[0066] In case the vectors database is already populated with a lot of data and most of the returned documents are already embedded, the method may comprise searching for the top-L documents (wherein L » M » K), such that the algorithm may always embed M new documents (i.e., the first M not embedded documents in the list of the top-L returned documents) into the vectors database.

[0067] In response to determining that a vector embedding associated with the first data file does not exist, in block 305, a vector embedding associated with the first data file is generated and added to a vector database. Furthermore, the first data file may be marked as embedded, in order to indicate that the vector embedding associated with the first data file already exists, so that the document can be skipped in the future.

[0068] After the completion of the steps detailed in blocks 301-305, the RAG algorithm may continue as normal.

[0069] Figure 4 is a schematic representation of a catalogue pre-filter steps according to an example. In the catalogue pre-filter 400, the user query may be encoded using a query encoder 401. The query may then be converted into a query in a database using a query converter 402. Then, using the LLM 403, the query may be converted into a set of catalogue (i.e., database) search criteria, in order to search the catalogue 404. This way, top-M documents 405 may be obtained. If any of the top-M documents 405 have not been embedded yet, vector embeddings may be applied to any of the documents, in order to obtain the vectors 406. The vector embeddings 406 may then be added to the vectors database 407, to be used by the RAG algorithm.

[0070] For some scenarios, it is possible that the top-M documents returned from the catalogue will not provide a good context for the next RAG step. In such case, the algorithm may repeat some of its steps (i.e., blocks 302 to block 305) in order to search for additional documents in the catalogue / database comprising multiple files. Figure 5 is a schematic representation of a catalogue pre-filter with a recursive process according to an example. Figure 5 is based on Figure 2 (with same elements functioning likewise). Additionally, in Figure 5, two feedback elements 510 and 512 are present. The recursive process realised by the presence of the two feedback elements 510 and 512 provides feedback, i.e., the feedback elements 510 and 512 check whether the results are good enough, as described in more detail below.

[0071] For example, the database (catalogue) comprising multiple files may initially be queried, and the set of data files may be selected based on a similarity between the set of search criteria and each of the multiple data files. Then, a check may be performed to determine whether similarity between the set of search criteria and the selected set of data files exceeds a first threshold value. In response to determining that the similarity between the set of search criteria and the selected set of data files is below the first threshold value, the database may be queried again in order to add additional data files of the multiple data files to the set of data files. The repeating step may end once the answer is determined to be good enough (i.e., similar enough, based on the similarity comparison), or after a predetermined amount of time / after a predetermined number of iterations. A similarity check may be performed either in the vectors database, and / or the result returned by the RAG algorithm.

[0072] Importantly, since the result of the conversion of the user’ s query into a search criteria using an LLM is probabilistic, the search criteria returned in this step are not deterministic. In other words, repeating the step described in block 302 multiple times will provide different results (and hence different top-M documents in block 303). Thus, the final result provided by the RAG algorithm can be improved.

[0073] Figure 6 is a schematic of a computing system according to an example. The computing system 600 comprises a processor 603, and a memory 605 coupled to the processor 603 and configured to store instructions or program code 607, executable by the processor 603. The computing system 600 can be, e.g., a computing system or apparatus, user equipment, a network device (physical or virtual), or part thereof. The computing system 600 comprises the program code 607 arranged to cause the computing system to perform the method of optimising vector embedding for natural language processing described above in relation to Figures 2-5.

[0074] According to an example, machine-readable instructions can be loaded onto a computer or other programmable data processing devices, so that the computer or other programmable data processing devices perform a series of operations to produce computer-implemented processing, thus the instructions executed on the computer or other programmable devices provide an operation for realizing functions specified by flow(s) in the flow charts and / or block(s) in the block diagrams.

[0075] Further, the teachings herein may be implemented in the form of a computer or software product, such as a non-transitory machine-readable storage medium, the computer software or product being stored in a storage medium and comprising a plurality of instructions, e.g., machine readable instructions, for making a computer device implement the methods recited in the examples of the present disclosure.

[0076] In some examples, some methods can be performed in a cloud-computing or network-based environment. Cloud-computing environments may provide various services and applications via the Internet. These cloud-based services (e.g., software as a service, platform as a service, infrastructure as a service, etc.) may be accessible through a web browser or other remote interface of the user equipment for example. Various functions described herein may be provided through a remote desktop environment or any other cloud-based computing environment.

[0077] While various embodiments have been described and / or illustrated herein in the context of fully functional computing systems, one or more of these exemplary embodiments may be distributed as a program product in a variety of forms, regardless of the particular type of computer- readable-storage media used to actually carry out the distribution. The embodiments disclosed herein may also be implemented using software modules that perform certain tasks. These software modules may include script, batch, or other executable files that may be stored on a computer-readable storage medium or in a computing system. In some embodiments, these software modules may configure a computing system to perform one or more of the exemplary embodiments disclosed herein. In addition, one or more of the modules described herein may transform data, physical devices, and / or representations of physical devices from one form to another.

[0078] The preceding description has been provided to enable others skilled in the art to best utilize various aspects of the exemplary embodiments disclosed herein. This exemplary description is not intended to be exhaustive or to be limited to any precise form disclosed. Many modifications and variations are possible without departing from the spirit and scope of the instant disclosure. The embodiments disclosed herein should be considered in all respects illustrative and not restrictive. Reference should be made to the appended claims and their equivalents in determining the scope of the instant disclosure.

Claims

CLAIMS1. A method of optimising vector embedding for natural language processing, the method comprising: receiving a query from a user (301); converting the received query into a set of search criteria (302); determining, using the search criteria, a set of data files from multiple data files (303), wherein the set of data files comprises a first data file; determining whether a vector embedding associated with the first data file exists (304); and in response to determining that a vector embedding associated with the first data file does not exist, generating a vector embedding associated with the first data file and adding the vector embedding associated with the first data file to a vector database (305).

2. The method of claim 1, further comprising: marking the first data file as embedded, whereby to indicate that the vector embedding associated with the first data file exists.

3. The method of claim 1 or 2, wherein the set of data files further comprises a second data file, the method further comprising: in response to determining that the vector embedding associated with the first data file exists, determining whether a vector embedding associated with the second data file exists.

4. The method of any one of claims 1, 2 or 3, wherein the converting the received query into the set of search criteria (302) comprises converting the received query into the set of search criteria using a large language model.

5. The method of any one of claims 1 to 4, wherein the determining, using the search criteria, the set of data files from the multiple data files (303) comprises: querying a database comprising multiple data files; and selecting the set of data files based on a similarity between the set of search criteria and each of the multiple data files.

6. The method of claim 5, further comprising: determining whether similarity between the set of search criteria and the selected set of data files exceeds a first threshold value; and in response to determining that the similarity between the set of search criteria and the selected set of data files is below the first threshold value, querying the database comprising multiple files to add additional data files of the multiple data files to the set of data files.

7. The method of any preceding claim, further comprising: generating an answer to the query received from the user based on the vector database.

8. The method of claim 7, further comprising: determining whether a similarity between the set of search criteria and the generated answer exceeds a second threshold value; and in response to determining that the similarity between the set of search criteria and the generated is below the second threshold value, querying the database comprising multiple files to add additional data files of the multiple data files to the set of data files.

9. A computer readable storage medium comprising computer program code, accessible by an apparatus comprising a processor, to provide instructions and / or data to the apparatus, the computer program code configured to, with the processor, cause the apparatus to: receive a query from a user; convert the received query into a set of search criteria; determine, using the search criteria, a set of data files from multiple data files, wherein the set of data files comprises a first data file; determine whether a vector embedding associated with the first data file exists; and in response to determining that a vector embedding associated with the first data file does not exist, generate a vector embedding associated with the first data file and add the vector embedding associated with the first data file to a vector database.

10. The computer readable storage medium of claim 9, further comprising the computer program code configured to, with the processor, cause the apparatus to: mark the first data file as embedded, whereby to indicate that the vector embedding associated with the first data file exists.

11. The computer readable storage medium of claim 9 or 10, wherein the set of data files further comprises a second data file, the computer readable storage medium further comprising the computer program code configured to, with the processor, cause the apparatus to: in response to determining that the vector embedding associated with the first data file exists, determine whether a vector embedding associated with the second data file exists.

12. The computer readable storage medium of claim 9, 10 or 11, wherein the computer code configured to, with the processor, cause the apparatus to determine, using the search criteria, a set of data files from the multiple data files comprises computer code configured to, with the processor, cause the apparatus to: query a database comprising multiple data files; and select the set of data files based on similarity between the set of search criteria and each of the multiple data files.

13. The computer readable storage medium of claim 12, further comprising the computer program code configured to, with the processor, cause the apparatus to: determine whether a similarity between the set of search criteria and the selected set of data files exceeds a first threshold value; and in response to determining that the similarity between the set of search criteria and the selected set of data files is below the first threshold value, query the database comprising multiple files to add additional data files of the multiple data files to the set of data files.

14. The computer readable storage medium of any one of claims 9 to 13, further comprising the computer program code configured to, with the processor, cause the apparatus to: generate an answer to the query received from the user based on the vector database.

15. The computer readable storage medium of claim 14, further comprising the computer program code configured to, with the processor, cause the apparatus to: determine whether similarity between the set of search criteria and the generated answer exceeds a second threshold value; and in response to determining that the similarity between the set of search criteria and the generated is below the second threshold value, query the database comprising multiple files to add additional data files of the multiple data files to the set of data files.

16. A computing system (600), comprising: a processor (603); a memory (605) coupled to the processor (603), configured to store: program code (607) executable by the processor (603), the program code (607) comprising one or more instructions to cause the computing system (600) to: receive a query from a user; convert the received query into a set of search criteria; determine, using the search criteria, a set of data files from multiple data files, wherein the set of data files comprises a first data file; determine whether a vector embedding associated with the first data file exists; and in response to determining that a vector embedding associated with the first data file does not exist, generate a vector embedding associated with the first data file and add the vector embedding associated with the first data file to a vector database.

Citation Information

Patent Citations

  • Artificial intelligence geospatial search

    US11809508B1

  • Deep Embedding for Natural Language Content Based on Semantic Dependencies

    US20180336183A1

Cited By

  • Optimized vector data storage

    US12705244B1

  • Semantic search for retrieval-augmented generation

    US20260178555A1

  • Optimized vector data storage

    US20260236475A1