Document search apparatus and method thereof

The document search technology addresses the inefficiencies in existing LLM-based document retrieval systems by preprocessing and vectorizing document data, clustering it using a preset embedding model, and deriving search results based on user queries, resulting in enhanced accuracy and efficiency for document retrieval in closed domains.

WO2025135602A1PCT designated stage expired Publication Date: 2025-06-26POSCO HLDG INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/019463
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-18
Filing Date
2024-12-02
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing document retrieval technologies using Large Language Models (LLM) face challenges in efficiently and accurately retrieving content from large datasets, especially in closed domains like companies, due to limitations in fine-tuning models for specific purposes and the inefficiency of retraining models for frequently changing data.

Method used

A document search technology that includes a preprocessing unit to convert document data into a vector format, a data management unit to split and cluster the data using a preset embedding model, and a search result derivation unit to calculate document corpus rankings and derive search results based on user queries, thereby enhancing the reliability and efficiency of document retrieval.

Benefits of technology

The proposed technology enables accurate and efficient retrieval of documents by utilizing a vector database to store preprocessed document data, allowing for quick and reliable search results based on user queries, thus improving data utilization and reducing work losses due to inefficient asset management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024019463_26062025_PF_FP_ABST
    Figure KR2024019463_26062025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to a document search technology, and provides a document search apparatus and method, the document search apparatus comprising: a preprocessing unit that preprocesses document data and generates preprocessed document data in order to store the document data in a vector database; a data management unit that splits the preprocessed document data, generates vector data by using a preset embedding model, and stores, in the vector database, a document corpus generated by clustering the vector data; and a search result derivation unit that, when a user query is input, calculates a document corpus ranking by using the user query and the document corpus, and derives a search result according to the document corpus ranking.
Need to check novelty before this filing date? Find Prior Art

Description

Document retrieval device and method thereof

[0001] The present disclosure relates to document retrieval technology. More specifically, it relates to technology for retrieving large amounts of documents and easily and accurately retrieving content within specific documents.

[0002] Recently, interest in LLMs (Large Language Models) has surged with the advent of the market. LLMs stand for large-scale language models. LLMs are used to perform natural language processing (NLP) tasks using deep learning algorithms and statistical modeling. These models can understand and generate sentence structure, grammar, and semantics by pre-training on large amounts of language data.

[0003] For example, in a problem of predicting the next word in a given context, LLM can generate the next word by identifying the similarity and context between words in a sentence. This task can be applied to various NLP tasks such as machine translation, text summarization, automatic writing, and question answering. LLM has various models, such as the Generative Pre-trained Transformer (GPT) and Bidirectional Encoder Representations from Transformers (BERT). Recently, attention has been focused on achieving more sophisticated language understanding and generation by utilizing large amounts of training data and large model architectures.

[0004] However, while pre-trained LLM models are typically used, fine-tuning them for specific purposes and fields requires fine-tuning them using the relevant data. However, fine-tuning them for each specific purpose is impractical. This is because retraining is not feasible for frequently changing data.

[0005] Thus, there are limitations in extracting reliable answers or search results due to the fact that retraining for LLM is practically impossible and the timing and type of training data are fixed when using a general-purpose model.

[0006] To overcome these limitations, in closed domains such as corporations, customized LLM models are being developed, or methods such as fine-tuning pre-trained models by additionally training them with domain-specific data are being attempted. However, this can result in high inefficiency when responding to frequent changes in search target data through model retraining, etc.

[0007] Therefore, there is a growing demand for technologies that can utilize the natural question-and-answer capabilities of LLM while increasing reliability based on data produced in a closed environment.

[0008] The present disclosure seeks to provide a document retrieval technology.

[0009] In one aspect, the present embodiments provide a document search device including a preprocessing unit that preprocesses document data to generate preprocessed document data in order to store the document data in a vector database, a data management unit that splits the preprocessed document data, generates vector data using a preset embedding model, clusters the vector data, and stores the generated document corpus in the vector database, and a search result derivation unit that, when a user query is input, calculates a document corpus ranking using a user query and a document corpus, and derives a search result according to the document corpus ranking.

[0010] In another aspect, the present embodiments provide a document retrieval method, including a preprocessing step of preprocessing document data to generate preprocessed document data in order to store the document data in a vector database, a data management step of splitting the preprocessed document data, generating vector data using a preset embedding model, and storing the generated document corpus by clustering the vector data in a vector database, and a search result derivation step of calculating a document corpus ranking using a user query and a document corpus when a user query is input, and deriving a search result according to the document corpus ranking.

[0011] According to the present disclosure, a document search technology can be provided.

[0012] Figure 1 is a drawing for explaining a query processing operation using a conventional LLM.

[0013] FIG. 2 is a drawing for explaining a document search system according to one embodiment.

[0014] FIG. 3 is a drawing for explaining the configuration of a document search device according to one embodiment.

[0015] FIG. 4 is a diagram for explaining an operation of preprocessing document data according to one embodiment.

[0016] FIG. 5 is a diagram for explaining an operation of converting spreadsheet data according to one embodiment.

[0017] FIG. 6 is a diagram for explaining an operation of converting spreadsheet data according to another embodiment.

[0018] FIG. 7 is a diagram illustrating an operation of converting spreadsheet data according to another embodiment.

[0019] FIG. 8 is a diagram illustrating the entire operation of processing a user's query using a document search device according to one embodiment.

[0020] FIG. 9 is a drawing for explaining a document search method according to one embodiment.

[0021] Hereinafter, some embodiments of the present disclosure will be described in detail with reference to exemplary drawings. When adding reference numerals to components in each drawing, identical components may have the same numerals as much as possible even if they are shown in different drawings. In addition, when describing the present embodiments, if it is determined that a detailed description of a related known configuration or function may obscure the gist of the technical idea of ​​the present invention, the detailed description may be omitted. When "includes," "has," "consists of," etc. are used in this specification, other parts may be added unless "only" is used. When a component is expressed in the singular, it may include a case in which the plural is included unless specifically stated otherwise.

[0022] Additionally, terms such as first, second, A, B, (a), (b), etc. may be used to describe components of the present disclosure. These terms are only intended to distinguish the components from other components, and the nature, order, sequence, or number of the components are not limited by the terms.

[0023] In a description of the positional relationship of components, when it is described that two or more components are "connected," "combined," or "connected," it should be understood that the two or more components may be directly "connected," "combined," or "connected," but that the two or more components may also be further "interposed" with another component to be "connected," "combined," or "connected." Here, the other component may be included in one or more of the two or more components that are "connected," "combined," or "connected" to each other.

[0024] In the description of the temporal flow relationship related to components, operation methods, or manufacturing methods, for example, when the temporal or flow relationship is described as “after”, “following”, “next to”, “before”, etc., it may also include cases where it is not continuous, unless “immediately” or “directly” is used.

[0025] Meanwhile, when numerical values ​​or corresponding information (e.g., levels, etc.) for components are mentioned, even without separate explicit description, the numerical values ​​or corresponding information may be interpreted as including an error range that may occur due to various factors (e.g., process factors, internal or external impact, noise, etc.).

[0026] The embodiments are described in detail with reference to the drawings below.

[0027] Figure 1 is a drawing for explaining a query processing operation using a conventional LLM.

[0028] Referring to Figure 1, the LLM can provide responses to user queries through pre-training. For example, when a user generates and inputs a query (110), the LLM (120) outputs a response (130) to the query (110) based on pre-trained content using a large text data set.

[0029] Large-Scale Language Models (LLMs, 120) are deep learning algorithms capable of performing various natural language processing (NLP) tasks. LLMs utilize Transformer models and are trained using massive data sets. They can recognize, translate, predict, or generate text or other content. LLMs (120) are also known as neural networks (NNs), which are computing systems inspired by the human brain. These neural networks operate using a network of layered nodes, much like neurons. Transformer models are the most common architecture for LLMs and consist of an encoder and decoder. Transformer models process data by tokenizing input information and then simultaneously executing mathematical equations to discover relationships between tokens. This allows computers to see patterns that humans would see when given the same query.

[0030] LLM (120) is a type of generative AI. Generative AI is a comprehensive term referring to artificial intelligence models capable of generating content. Generative AI can generate text, code, images, videos, music, and more. Examples of generative AI include Midjourney, DALL-E, and ChatGPT. Therefore, large-scale language models are a type of generative AI that are trained on text to generate textual content.

[0031] Such LLM (120) can be used in a wide range of applications, including language translation, sentence completion, sentiment analysis, and question answering. Furthermore, LLM (120) can be continuously improved by adding more data and parameters. When demonstrating in-context learning, large-scale language models have the advantage of learning quickly, as they do not require additional weights, resources, or parameters for training.

[0032] However, although LLM (120) may give the impression that it can understand the meaning of a query (110) and respond accurately, it also has various limitations such as hallucinations, security, and bias.

[0033] For example, when a query (110) is input, the LLM (120) can predict the next syntactically correct word or phrase and generate a response (130) even if there is no processing data for the query (110). This is called hallucination, and can produce results that do not match the user's intention or provide a response (130) with low accuracy. Bias refers to the appearance of a biased response (130) when the data used to train the language model is statistically biased. Security issues may also arise, such as parsing during the learning process or generating spam.

[0034] In particular, when users easily generate queries using internal corporate data and then use LLM to provide responses, responses that are not based on the internal corporate data may be generated, resulting in the illusion of results referencing completely different content. Furthermore, fine-tuning or developing LLMs using internal corporate data can lead to errors and a significant increase in development costs and time.

[0035] In this context, the present disclosure aims to provide technology that enables accurate responses to be derived from documents generated within a closed domain, such as within a company. Specifically, the disclosure aims to provide document retrieval technology that enables LLM responses to be generated based on internal company data by creating a separate database of internal company data that can be linked to LLM.

[0036] In particular, companies possess countless data related to their domains. Furthermore, various types of documents are created daily by employees and stored in internal assets for management purposes.

[0037] In this situation, as documents and various information accumulated daily, they become difficult to utilize properly, and it becomes difficult for authors to understand past work and information. This not only hinders the utilization of vast amounts of data, but can also lead to operational losses, such as having to redo previous research or documentation work. To address these operational inefficiencies and ensure proper access to these documents, new search technologies are needed.

[0038] When developing plans and strategies using existing assets, users often struggle to identify the source and details of relevant information. Furthermore, identifying large or complex documents can be time-consuming. To address this issue, the LLM-based document search system according to this embodiment allows users to identify summaries and sources of relevant information through natural language queries, and can even automatically pop-up source files when necessary. This eliminates the need to manually search for documents. Furthermore, the search engine leverages LLM to mitigate the workload associated with document creation and other repetitive tasks due to inadequate asset management.

[0039] FIG. 2 is a drawing for explaining a document search system according to one embodiment.

[0040] Referring to FIG. 2, when a user query (110) is input in the document retrieval system, the document retrieval device (200) searches for documents similar to the user query (110) in the database (210). The document retrieval device (200) transmits the user query (110) and the search results retrieved to the LLM (120). The LLM (120) uses the user query (110) and the search results retrieved from the database to generate a response (130) and provides it to the user.

[0041] That is, unlike Fig. 1, the document search system can prevent the provision of inaccurate answers, which is a problem of the conventional LLM system, by additionally configuring a document search device (200) and a database (210) for storing document data.

[0042] This can expand the capabilities of the large-scale language model, increasing the reliability of user queries and requests and providing information-rich answers. In other words, the document retrieval system according to this embodiment combines the search model with the generative model to operate as a single framework. This allows the system to generate and provide a response (130) containing accurate and rich information tailored to the context of the user query (110).

[0043] However, when using such a model, additional technology is required to configure a database (210) and perform appropriate search operations according to the input of a user query (110). In particular, documents produced by a company have various formats. In particular, in the case of spreadsheet data, it is difficult to accurately understand the meaning thereof, and additional preprocessing operations are required to use it in a database for searching. In addition, when a user query (110) is input, it takes time to search to extract documents or data with high similarity from the database (210). When a large number of documents produced by a company are databased, a problem may arise in which the time required for similarity-based search increases significantly.

[0044] To prevent such problems, the document search device (200) in this embodiment preprocesses and manages document data input into the aforementioned database (210), and establishes a search strategy according to the input of a query (110) to quickly derive search results.

[0045] Below, a document search device (200) that can operate on the aforementioned document search system is described.

[0046] FIG. 3 is a drawing for explaining the configuration of a document search device according to one embodiment.

[0047] Referring to FIG. 3, the document search device (200) may include a preprocessing unit (310) that preprocesses document data to generate preprocessed document data in order to store the document data in a vector database.

[0048] For example, a document search device (200) must store document data in a vector database to improve search speed and accuracy. The document data must be converted to vector format and stored. However, the document data produced comes in various forms, such as sentences, numbers, and images, and the document formats also vary. Therefore, preprocessing of the document data is necessary to store it in a vector database for easy future search. In particular, complex spreadsheet document formats, such as Excel, can cause problems such as hallucination.

[0049] Accordingly, the preprocessing unit (310) can preprocess document data through a preset process to generate preprocessed document data.

[0050] For example, if the document data is in the form of a spreadsheet file, the preprocessing unit (310) can convert the data of the spreadsheet file into a preset data format based on the reference row and reference column information.

[0051] For example, the reference row and reference column information may be set to a predefined row or column when the data in the spreadsheet file consists of rows or columns that are below the reference value. As another example, the reference row and reference column information may be set based on user input information when the data in the spreadsheet file consists of rows or columns that exceed the reference value.

[0052] For example, the preprocessing unit (310) can provide two types of preprocessing functions: automated and semi-automated. A key preprocessing step is to refine document data so that sentences can be generated based on accurate values ​​without hallucinations when generating the final results from LLM. Therefore, the preprocessing unit (310) transforms document data so that individual values ​​are mapped to a single row and column, thereby improving accuracy.

[0053] For example, when document data for low-complexity document data is input, the preprocessing unit (310) modifies and transforms the document data according to established criteria. For example, low-complexity document data may range from cases with single row and column values ​​to cases where rows or columns are merged, such that a single value does not match a row or column.

[0054] As another example, the preprocessing unit (310) may preprocess and convert an Excel file with multiple rows and columns. In this case, the preprocessing unit (310) may use user input for reference rows and / or reference columns, if present, to preprocess and convert the document data so that the rows and columns can be matched to a single value.

[0055] Additionally, the preprocessing unit (310) can release the merged cell if the reference row or reference column includes a merged cell and designate the data of the merged cell as the data of each released cell.

[0056] The preprocessing unit (310) can generate index information according to a specified order when there are multiple reference rows or reference columns, and convert the index to a preset data format by mapping the data of cells not included in the reference rows and reference columns. For example, the preset data format may be JSON, etc.

[0057] Meanwhile, the preprocessing unit (310) can prevent data loss by utilizing Python's openpyxl library. For example, the preprocessing unit (310) can preprocess document data so that the contents of one sheet can all be matched to one page of a PDF to be converted.

[0058] The document search device (200) may include a data management unit (320) that splits preprocessed document data, generates vector data using a preset embedding model, and stores the generated document corpus by clustering the vector data in a vector database.

[0059] Once preprocessed document data is generated, data management is required to store it in a vector database.

[0060] For example, the data management unit (320) can split preprocessed document data and generate vector data using a preset embedding model. For example, the data management unit (320) can segment preprocessed document data on which preprocessing has been performed. Various publicly available algorithms can be used as algorithms for separating preprocessed document data. Once the preprocessed document data is split, the data management unit (320) can convert the preprocessed document data into vector data using a preset embedding model for each split.

[0061] For example, the data management unit (320) can generate vector data by separating preprocessed document data into chunks of preset size based on the maximum token limit of a preset embedding model. The maximum token limit can be preset in various ways for each embedding model. For example, the embedding model may be an Ada2 embedding model.

[0062] Additionally, the data management unit (320) may perform topic-based clustering to increase search speed. For example, the data management unit (320) may cluster vector data into two or more document corpora using a density-based clustering model, and store the document corpora in a vector database by dividing them. A document corpus may refer to a group of vector data that are classified into the same cluster according to a clustering criterion. The density-based clustering model may utilize topic modeling techniques that create dense clusters that are easy to interpret while retaining important words in the topic description. Examples include, but are not limited to, a BERTopic model that utilizes BERT embeddings and class-based TF-IDF. The aforementioned clustering operation may also be performed using preprocessed document data.

[0063] Through this, the data management unit (320) can classify preprocessed document data by document corpus, create vector data through embedding, and store it in a vector database.

[0064] The document search device (200) may include a search result derivation unit (330) that, when a user query is input, calculates a document corpus ranking using the user query and the document corpus, and derives search results according to the document corpus ranking.

[0065] The search result derivation unit (330) can control the search to be performed quickly through two stages of sparse search and dense search when a user query is entered. For example, the search result derivation unit (330) can determine the similarity between the user query and the document corpus and calculate the document corpus ranking based on the similarity determination result. To determine the similarity with the document corpus, the data management unit (320) can also generate index information corresponding to the document corpus.

[0066] For example, the search result derivation unit (330) may calculate a score for each document corpus using the frequency with which words included in a user query appear within the document corpus, the amount of information provided by the words, and weighted information based on document length, and may then calculate a document corpus ranking based on the score. For example, the search result derivation unit (330) may determine similarity using the BM25 model. However, this model is merely exemplary and is not limited thereto.

[0067] The search result derivation unit (330) can set a search priority based on the document corpus ranking, and derive search results by measuring the cosine similarity of the user query using the highest-priority document corpus determined according to the priority. Through this, the search result derivation unit (330) can minimize the search time required to perform a similarity search between all vector data stored in the vector database and the user query. In other words, the optimal search result can be quickly derived by performing a sparse search by performing a ranking of the document corpus and performing a dense search that analyzes the similarity between the vector data included in a specific document corpus and the user query.

[0068] Meanwhile, the search result derivation unit (330) can change the document corpus according to the user's additional search input signal, measure the cosine similarity of the user's query, and provide additional search results. In this case, search results can be provided within the document corpus prioritized according to the second ranking.

[0069] The aforementioned similarity analysis can be applied in various ways, utilizing vector similarity analysis. For example, analysis techniques such as cosine similarity can be applied.

[0070] When a search result is derived, the document search device (200) can input a user query, search result, and document data into LLM to provide a response to the user.

[0071] Below, embodiments of each component operation of the aforementioned document search device (200) are described with reference to the drawings. The embodiments described below are for convenience of explanation and are not limited to the embodiments described below.

[0072] FIG. 4 is a diagram for explaining an operation of preprocessing document data according to one embodiment.

[0073] Referring to FIG. 4, the preprocessing unit (310) may perform preprocessing when document data (400) is input and transmit it to the data management unit (320). However, the preprocessing unit (310) may also perform preprocessing by selecting only data in a preset format. For example, unlike spreadsheets, PDF files or DOC files consisting only of text may not require a separate preprocessing operation. In this case, the document data (400) may be directly transmitted to the data management unit (320) and embedded.

[0074] For example, the preprocessing unit (310) may convert the data in the spreadsheet file into a preset data format based on reference row and reference column information when the document data is in the form of a spreadsheet file. A spreadsheet file contains a large amount of data, including data in a table format like Excel. In addition, the data in the spreadsheet file may be composed of rows and columns, may include an index, and may be configured in a complex form, such as merging cells.

[0075] The preprocessing unit (310) can perform automatic or semi-automatic preprocessing operations using reference row and reference column information included in document data in the form of a spreadsheet file.

[0076] For example, if the data in the spreadsheet file consists of rows or columns below a reference value, the preprocessing unit (310) can set a pre-specified row or column as a reference row or column and perform data conversion based on this. This is described as an automatic preprocessing operation.

[0077] The automatic preprocessing operation provides results from the user's file input for relatively less complex documents. These less complex cases range from cases with a single row and column to cases where rows or columns are merged, with no matching row or column value. In these cases, the preprocessing unit (310) can re-segment each row and column of the merged cells to match a single value for each cell.

[0078] FIG. 5 is a diagram for explaining an operation of converting spreadsheet data according to one embodiment.

[0079] Referring to FIG. 5, spreadsheet data (500) of document data may be composed of two rows and seven columns. The aforementioned reference value may be set to 2, and since the number of rows of the spreadsheet data (500) is set to 2, an automatic preprocessing operation may be performed.

[0080] When spreadsheet data (500) is input, the preprocessing unit (310) determines whether there are merged cells and cells containing formulas. In addition, the preprocessing unit (310) can delete all empty cells without values, leaving only cells containing values.

[0081] In cases where rows below the reference value are included, the preprocessing unit (310) can designate pre-designated rows or columns as reference rows and reference columns. The pre-designated rows or columns can be designated as the first remaining rows and columns excluding rows and columns with no values. In Fig. 5, the first row representing a region and the first column containing gasoline usage data can be designated as the reference row and reference column.

[0082] Once the reference row and reference column are determined, the preprocessing unit (310) can convert the spreadsheet data (500) into a preset data format (550) using the values ​​of the reference row and reference column as indexes. JSON is shown here as an example. For example, the preprocessing unit (310) can convert the value of each cell using the smaller number of rows and columns as an index. As another example, the preprocessing unit (310) can also set the row index as an index first. As shown in FIG. 5, the data includes two rows, and the preprocessing unit (310) can convert the format by matching the values ​​of each column and row using the gasoline usage value as an index.

[0083] Meanwhile, in the case of a file with multiple reference rows or columns, the preprocessing unit (310) cannot automatically set the reference rows and columns. That is, in the case where the data in the spreadsheet file consists of rows or columns exceeding the reference value, the reference row and reference column information can be set based on user input information. This is described as a semi-automatic preprocessing operation.

[0084] Semi-automatic preprocessing is suitable for spreadsheet files with multiple rows and columns. The preprocessing unit (310) can input user-defined reference row and reference column information, reconstruct the spreadsheet file using the input values, and convert the rows and columns to correspond to a single value.

[0085] For example, if there are multiple rows that should be an Index, the leftmost row becomes the largest category of the rows that will become the remaining Indexes due to the nature of the rows. In addition, there is a high possibility that the cells of the largest category will exist in the form of integrated cells. The preprocessing unit (310) can separate the integrated cells so that they can be matched to one row and fill in the existing values ​​in the separated empty spaces, thereby converting them so that individual values ​​can be mapped to each row. This operation is designed as a loop algorithm so that the rows that should become the remaining Indexes can be preprocessed in the same way. In addition, when the preprocessing unit (310) performs the task, if there is a left row of the reference row, the value of the left row is referenced to preprocess the value of the reference row so that an accurate value can be input.

[0086] FIG. 6 is a diagram for explaining an operation of converting spreadsheet data according to another embodiment.

[0087] Referring to FIG. 6, spreadsheet data (600) may include a number of rows and columns. When the reference value is set to 2, the preprocessing unit (310) may obtain reference row and reference column information based on user input information, as both the number of rows and columns exceeds 2.

[0088] As described above, the preprocessing unit (310) can unmerge the merged "by city / province," "gasoline," "passenger," and "combined" cells and add the data of the merged cells to each cell. Here, the reference row can be set to two rows, and the reference column can be set to three columns. The reference row and reference column can be set according to the user's input.

[0089] For example, if the row-based user input is entered as 1, 2, the row-based user input is set to the first row and the second row. Similarly, if the column-based user input is entered as 1, 3, the first to third columns can be set as the reference columns. The preprocessing unit (310) performs category confirmation work from the first to the third columns. Since the reference columns are set to multiple, the preprocessing unit (310) sets a new index such as gasoline_passenger_non-commercial, gasoline_passenger_commercial, or gasoline_passenger_non-commercial.

[0090] When the column-based index setting is completed, the preprocessing unit (310) can generate preprocessed document data (650) whose format has been converted by linking the row-based index value with the data of each cell.

[0091] FIG. 7 is a diagram illustrating an operation of converting spreadsheet data according to another embodiment.

[0092] Referring to FIG. 7, the preprocessing unit (310) can generate preprocessed document data even when there are no values ​​for some of the reference indices. For example, when spreadsheet data (700) is input, the preprocessing unit (310) can check for merged cells as described above, unmerge them, and generate data for each cell. In the case of 700, when the user input information is set to 1 and 4, columns 1 through 4 become reference columns. However, there may be cases where there are no values ​​in the reference columns.

[0093] In this case as well, the preprocessing unit (310) can create index information such as China_Heilongjiang_Yichun City, Germany_Salzkitter_, Germany_Göttingen_... That is, the reference column without a value is left blank and no other value is entered.

[0094] Through the above actions, files in the form of spreadsheets can also be converted to input into a vector database, thereby providing effective searches.

[0095] Meanwhile, as described above, the document retrieval device may include a data management unit. The data management unit may split preprocessed document data, generate vector data using a preset embedding model, and cluster the vector data to store the generated document corpus in a vector database.

[0096] For example, the data management unit performs vectorization on preprocessed document data using OpenAI's Ada2 embedding model. Because embedding a large number of documents has limitations in tokenization, the data management unit can input preprocessed document data through a loop algorithm for each document. The input preprocessed document data can be separated based on a preset chunk size. For example, the data management unit can separate based on chunk_size = 1500, chunk_overlap = 300, and the separator can be set to [" / n / n", " / n"," ",""]. ​​Therefore, if the size exceeds 1500, recursive splitting can be performed based on the value specified in the separator standard.

[0097] Meanwhile, performing a search across the entire document data set can present challenges in accurately retrieving information and generating results that address user queries. While accurate information retrieval is essential, dense retrieval alone can be challenging when retrieving individual chunks of complex, similar documents. Dense retrieval, a semantic search strategy, can be effective when the search target data is small and contains disparate content. However, repetitive and similar data require additional retrieval or filtering.

[0098] Therefore, the data management unit performs document categorization, categorizing by topic while separating search targets to improve search efficiency. For example, the data management unit can cluster preprocessed document data by topic using BERTopic. Clustering can be performed on vector data, or on preprocessed document data using BERT embeddings and class-based TF-IDF.

[0099] An example of a topic clustering criterion could be as follows.

[0100] Topic 1 - Secondary batteries, cathode materials, graphite, lithium ions…

[0101] Topic 2 - Hydrogen, Green, Water Electrolysis, Power Supply and Demand…

[0102] Topic 3 - Industry, Economy, Vietnam, Governance…

[0103] Additionally, the data management unit can configure individual retrieval engines by inputting document corpora categorized by topic into VectorDB.

[0104] The search result generation unit receives document corpus rankings and user queries as input and retrieves the final candidate chunks. For example, FAISS vectorDB can be used in the indexing area via VectorDB, and Euclidean distance can be used to measure the similarity or proximity between vectors. Similarly, Consine similarity can be used for similarity searches for user queries. However, the aforementioned uses are not limited to this.

[0105] For example, a document corpus categorized by topic is used not only as a vector index but also as input to the search result generation unit, which compares it with the user query to establish a ranking of the most similar document corpus. For example, the search result generation unit could use the BM25 model's bag-of-words concept to determine the frequency with which words in the user query appear in documents, thereby establishing a document corpus ranking.

[0106] For example, the search result generation unit can set the document corpus ranking by using TF-IDF (Term Frequency Inverse Document Frequency) to determine the importance of a word when the word frequently appears in a document but does not appear in other documents, and give it a higher score. Here, TF is a factor for the frequency count of the word appearance in the document. For example, the number of times each word appears can be counted, and the number of times can be normalized by dividing the count by the total number of words appearing in the document. IDF is a factor for the amount of information provided by a word. For example, IDF can be calculated by Document Frequency (DF): the number of documents in which the term t appears, and N: the total number of documents.

[0107] As another example, the search result generation unit can calculate scores based on the concept of TF-IDF, taking document length into account. In this case, a limit can be set for the TF value to ensure it remains within a certain range. Furthermore, if a word matches a document shorter than the average document length, weighting can be applied to that document.

[0108] Finally, the search result generation unit performs a search for the user query based on the document corpus order corresponding to the priority topic, resulting in search results based on the document category most similar to the user query. The search results are then input into the LLM to generate a response. After reviewing the generated response, the user can request searches and generation for other categories.

[0109] FIG. 8 is a diagram illustrating the entire operation of processing a user's query using a document search device according to one embodiment.

[0110]

[0111] *The operation of the overall system is described with reference to Figure 8. However, the contents described above are omitted or briefly explained.

[0112] When document data is input, the preprocessing unit (310) generates preprocessed document data through the aforementioned operations and transfers it to the data management unit (320). The data management unit (320) embeds the preprocessed document data, performs clustering, and stores vector data in a vector database. At this time, a document corpus can be generated through clustering, and a search engine can be configured for each document corpus.

[0113] When a user inputs a user query and a prompt through an input device (800), the user query is transmitted to a search result derivation unit (330). The search result derivation unit (330) uses the user query to derive a document corpus ranking through a vector index. Once the document corpus ranking is derived, the search result derivation unit (330) measures the similarity between the vector data included in the document corpus with the highest priority and the user query. The search result derived based on the similarity (including document data or vector data) is transmitted to the input device (800). The input device (800) provides a response to the user's query by inputting the prompt, the user query, and the enhanced context derived from the search result into the LLM (810) and receiving a generated text response.

[0114] In Fig. 8, a system including an input device (800) is exemplarily described, but the search result derivation unit (330) may also directly transmit the user query, search results, and prompts to the LLM (810).

[0115] Through the aforementioned actions, a large number of diverse document types can be stored in a vector database, and the document data stored in the vector database can be used to provide users with more accurate and faster search results in the form of generative text. Furthermore, by dividing the search process into two stages, search speed can be increased and computing load reduced.

[0116] Below, the operations performed by the aforementioned document retrieval device are described again, using a flowchart. The aforementioned operations can be performed at each step of the flowchart described below. Furthermore, each step described below can be split or combined. Furthermore, additional steps can be added, or their order can be changed.

[0117] FIG. 9 is a drawing for explaining a document search method according to one embodiment.

[0118] Referring to FIG. 9, the document search method may include a preprocessing step of preprocessing document data to generate preprocessed document data in order to store the document data in a vector database (S900).

[0119] For example, the preprocessing step can convert data in a spreadsheet file into a preset data format based on reference row and reference column information when the document data is in the form of a spreadsheet file.

[0120] For example, the reference row and reference column information may be set to a predefined row or column when the data in the spreadsheet file consists of rows or columns that are below the reference value. As another example, the reference row and reference column information may be set based on user input information when the data in the spreadsheet file consists of rows or columns that exceed the reference value.

[0121] For example, the preprocessing step can provide two types of preprocessing functions: automated and semi-automated. A crucial aspect of the preprocessing process is to refine the document data so that LLM can generate sentences based on accurate values ​​without hallucinations when generating the final results. Therefore, the preprocessing step improves accuracy by transforming the document data so that individual values ​​are mapped to a single row and column.

[0122] For example, the preprocessing step modifies and transforms low-complexity document data based on established criteria when it's entered. For example, low-complexity document data can range from having a single row and column of values ​​to cases where rows or columns are merged, resulting in a single value that doesn't match the row or column.

[0123] As another example, the preprocessing step can transform an Excel file with multiple rows and columns. In this case, the preprocessing step uses user input for the reference rows and / or columns, if any, to preprocess and transform the document data so that the rows and columns correspond to a single value.

[0124] Additionally, the preprocessing step can unmerge cells if the reference row or reference column contains merged cells and assign the data of the merged cells to the data of each unmerge cell.

[0125] The preprocessing step can generate index information in a specified order when there are multiple reference rows or columns, and map the index to data in cells not included in the reference rows or columns, thereby converting them into a preset data format. For example, the preset data format may be JSON, etc.

[0126] Meanwhile, the preprocessing step can prevent data loss. For example, the preprocessing step can preprocess document data so that the contents of a single sheet can all be mapped to a single converted PDF page.

[0127] The document search method may include a data management step of splitting preprocessed document data, generating vector data using a preset embedding model, and storing the generated document corpus by clustering the vector data in a vector database (S910).

[0128] Once preprocessed document data is generated, data management is required to store it in a vector database.

[0129] For example, the data management step can split preprocessed document data and generate vector data using a preset embedding model. For example, the data management step can segment preprocessed document data that has undergone preprocessing. Various publicly available algorithms can be used for splitting preprocessed document data. Once the preprocessed document data is split, the data management step can convert the preprocessed document data into vector data using a preset embedding model for each split.

[0130] For example, the data management step can generate vector data by dividing preprocessed document data into chunks of preset sizes based on the maximum token limit of a preset embedding model. The maximum token limit can be preset differently for each embedding model. For example, the embedding model can be an Ada2 embedding model.

[0131] Additionally, the data management step can perform topic-based clustering to increase search speed. For example, the data management step can cluster vector data into two or more document corpora using a density-based clustering model, and then classify the document corpora and store them in a vector database. A document corpus can refer to a group of vector data that are classified into the same cluster based on a clustering criterion. The density-based clustering model can utilize topic modeling techniques to create dense clusters that are easily interpretable while retaining important words in the topic description. Examples include, but are not limited to, the BERTopic model, which utilizes BERT embeddings and class-based TF-IDF. The aforementioned clustering operation can also be performed using preprocessed document data.

[0132] Through this, the data management step can classify preprocessed document data by document corpus, create vector data through embedding, and store it in a vector database.

[0133] The document search method may include a search result derivation step of, when a user query is input, calculating a document corpus ranking using the user query and the document corpus, and deriving search results according to the document corpus ranking (S920).

[0134] The search result derivation step can be controlled to quickly perform a search through two stages: sparse search and dense search when a user query is entered. For example, the search result derivation step can determine the similarity between the user query and the document corpus and calculate a document corpus ranking based on the similarity determination results. To determine similarity with the document corpus, the data management step can also generate index information corresponding to the document corpus.

[0135] For example, the search result derivation step can calculate a score for each document corpus using the frequency with which words included in the user query appear within the document corpus, the amount of information provided by the words, and weighted information based on document length. The search result derivation step can then calculate a document corpus ranking based on the score. For example, the search result derivation step can use the BM25 model to determine similarity. However, this model is merely exemplary and is not limited thereto.

[0136] The search result derivation step sets search priorities based on document corpus rankings, and uses the highest-priority document corpus determined based on these priorities to measure the cosine similarity of the user query to derive search results. This minimizes the search time required to perform a similarity search between all vector data stored in the vector database and the user query. Specifically, the search results derivation step performs sparse searches by ranking the document corpus, and then performs dense searches that analyze the similarity between the vector data contained within a specific document corpus and the user query, thereby quickly deriving optimal search results.

[0137] Meanwhile, the search result derivation step can provide additional search results by modifying the document corpus based on the user's additional search input signals and measuring the cosine similarity of the user's query. In this case, search results can be provided within the document corpus prioritized according to the second ranking.

[0138] The aforementioned similarity analysis can be applied in various ways, utilizing vector similarity analysis. For example, analysis techniques such as cosine similarity can be applied.

[0139] The above description is merely an illustrative example of the technical idea of ​​the present disclosure, and those skilled in the art to which the present disclosure pertains will appreciate that various modifications and variations can be made without departing from the essential characteristics of the technical idea of ​​the present disclosure. In addition, the present embodiments are not intended to limit the technical idea of ​​the present disclosure but rather to explain it, and therefore the scope of the technical idea of ​​the present disclosure is not limited by these embodiments. The scope of protection of the present disclosure should be interpreted by the claims below, and all technical ideas within a scope equivalent thereto should be interpreted as being included within the scope of the rights of the present disclosure.

[0140]

[0141] CROSS-REFERENCE TO RELATED APPLICATION

[0142] This patent application claims priority under 35 USC § 119(a) to Korean Patent Application No. 10-2023-0184032, filed December 18, 2023, the entire contents of which are incorporated herein by reference. Furthermore, this patent application claims priority in countries other than the United States for the same reasons, the entire contents of which are incorporated herein by reference.

Claims

1. In a document search device, A preprocessing unit that preprocesses the document data to create preprocessed document data in order to store the document data in a vector database; A data management unit that splits the above preprocessed document data, generates vector data using a preset embedding model, and stores the document corpus generated by clustering the vector data in the vector database; and A document search device including a search result derivation unit that, when a user query is input, calculates a document corpus ranking using the user query and the document corpus, and derives search results according to the document corpus ranking.

2. In paragraph 1, The above preprocessing unit, A document search device that converts data of the spreadsheet file into a preset data format based on reference row and reference column information when the document data is in the form of a spreadsheet file.

3. In paragraph 2, The above reference row and reference column information is, If the data in the above spreadsheet file consists of rows or columns below the standard value, it is set to a pre-specified row or column, A document retrieval device set based on user input information when the data in the above spreadsheet file consists of rows or columns exceeding the above criteria.

4. In paragraph 2, The above preprocessing unit, If the above reference row or reference column includes a merged cell, the merged cell is unmerged and the data of the merged cell is designated as the data of each unmerged cell, A document search device that generates index information according to a specified order when there are multiple reference rows or reference columns, and maps the index to data of cells not included in the reference rows and reference columns and converts it into the preset data format.

5. In paragraph 1, The above data management department, A document search device that generates the vector data by dividing the preprocessed document data based on a preset chunk size according to the maximum token limit of the preset embedding model.

6. In paragraph 1, The above data management department, A document search device that clusters the above vector data into two or more document corpora using a density-based clustering model, and stores the document corpora by dividing them in the vector database.

7. In paragraph 1, The above search result derivation section is, A document search device that determines the similarity between the user query and the document corpus, and calculates a document corpus ranking based on the result of the similarity determination.

8. In paragraph 7, The above search result derivation section is, A document search device that calculates a score for each document corpus by using the frequency with which words included in the user query appear in the document corpus, the amount of information provided by the words, and weight information according to the document length, and calculates a document corpus ranking based on the score.

9. In paragraph 1, The above search result derivation section is, A document search device that sets a search priority based on the above document corpus ranking, and derives the search result by measuring the cosine similarity of the user query using the highest priority document corpus determined according to the above priority.

10. In paragraph 9, The above search result derivation section is, A document search device that changes the document corpus according to an additional search input signal from a user, measures the cosine similarity of the user query, and provides additional search results.

11. In terms of document search methods, A preprocessing step of preprocessing the document data to create preprocessed document data in order to store the document data in a vector database; A data management step of splitting the above preprocessed document data, generating vector data using a preset embedding model, and storing the document corpus generated by clustering the vector data in the vector database; and A document search method including a search result derivation step of, when a user query is input, calculating a document corpus ranking using the user query and the document corpus, and deriving search results according to the document corpus ranking.

12. In paragraph 11, The above preprocessing step is, A document search method for converting data of the spreadsheet file into a preset data format based on reference row and reference column information when the above document data is in the form of a spreadsheet file.

13. In paragraph 12, The above reference row and reference column information is, If the data in the above spreadsheet file consists of rows or columns below the standard value, it is set to a pre-specified row or column, A document search method set based on user input information when the data in the above spreadsheet file consists of rows or columns exceeding the above criteria.

14. In paragraph 12, The above preprocessing step is, If the above reference row or reference column includes a merged cell, the merged cell is unmerged and the data of the merged cell is designated as the data of each unmerged cell, A document search method for generating index information according to a specified order when there are multiple reference rows or reference columns, and converting the index to data of cells not included in the reference rows and reference columns into the preset data format.

15. In paragraph 11, The above data management steps are: A document search method for generating vector data by dividing the preprocessed document data based on a preset chunk size according to a maximum token limit of the preset embedding model.

16. In paragraph 11, The above data management steps are: A document search method comprising: performing clustering of the above vector data into two or more document corpora using a density-based clustering model, and storing the document corpora in the vector database by dividing them.

17. In paragraph 11, The above search result derivation step is, A document search method for determining the similarity between the user query and the document corpus, and calculating the document corpus ranking based on the result of the similarity determination.

18. In paragraph 17, The above search result derivation step is, A document search method for calculating a score for each document corpus by using the frequency with which words included in the user query appear in the document corpus, the amount of information provided by the words, and weight information according to the document length, and calculating a document corpus ranking based on the score.

19. In paragraph 11, The above search result derivation step is, A document search method for setting a search priority based on the above document corpus ranking, and deriving the search result by measuring the cosine similarity of the user query using the highest priority document corpus determined according to the above priority.

20. In paragraph 19, The above search result derivation step is, A document search method for providing additional search results by changing the document corpus according to an additional search input signal of the user and measuring the cosine similarity of the user query.

Citation Information

Patent Citations

  • Gender for receiving and charging wireless low power and configuration method thereof

    KR1020200116642A

  • Method of processing for rice

    KR1020210033707A

  • System and method for the analysis of large amount of data

    KR102029942B1

  • Deep learning solution providing method performed by deep learning platform providing device for providing deep learning solution platform

    WO2022092448A1

  • KR20210023452A