Information providing method and system using full-text index in which graph data structure and vector data are integrated

By integrating a graph data structure and vector data into a full-text index, the system addresses the limitations of existing search methods by providing richer search results that incorporate both semantic and structural data perspectives.

WO2025135415A1PCT designated stage expired Publication Date: 2025-06-26INTELLECTUS CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/014830
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-20
Filing Date
2024-09-30
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing search methods using indexes are limited to providing data similar to the search term, failing to offer richer search results that incorporate both semantic and structural perspectives of the data.

Method used

The method and system utilize a full-text index integrating a graph data structure and vector data, allowing for the determination of entity nodes based on natural language queries, and providing data associated with these nodes, along with connected data, to offer richer search results.

Benefits of technology

This approach enables the provision of information from both semantic and structural perspectives, enhancing search results by including data similar to the query and data connected through a logical data model on the graph.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024014830_26062025_PF_FP_ABST
    Figure KR2024014830_26062025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides an information providing method using a full-text index, the method being performed by at least one processor. The method comprises the steps of: receiving a natural language query; determining, on the basis of the natural language query, a first entity node stored in a first index database formed in a graph structure; extracting data associated with the first entity node from a source database; and outputting the result for the received natural language query by using data associated with a first entity.
Need to check novelty before this filing date? Find Prior Art

Description

Information provision method and system using a full-text index integrating graph data structures and vector data

[0001] The present disclosure relates to a method and system for providing information, and more particularly, to a method and system for determining a first entity node stored in a first index database formed in a graph structure based on a received natural language query, and providing data associated with the first entity node as a result of the natural language query.

[0002] As network and information processing technologies advance, a variety of data can be collected and utilized to provide users with diverse services. To utilize this data, it's crucial to effectively provide users with the data they seek.

[0003] Meanwhile, an index is one way to improve search efficiency by pre-sorting and storing data corresponding to key items in accumulated data. However, searches using such an index have the limitation of simply providing data similar to the search term.

[0004] The present disclosure provides a method and system (device) for providing information using a full text index that integrates a graph data structure and vector data to solve the above-described problems.

[0005] The present disclosure can be implemented in various ways, including as a method, a device (system), or a computer program stored on a readable storage medium.

[0006] An information providing method according to one embodiment of the present disclosure may include a step of receiving a natural language query, a step of determining a first entity node stored in a first index database formed in a graph structure based on the natural language query, a step of extracting data associated with the first entity node from a source database, and a step of outputting a result for the received natural language query using the data associated with the first entity.

[0007] According to one embodiment of the present disclosure, the step of determining the first entity node may include the step of extracting an entity from a natural language query using a machine learning model, and the step of determining the first entity node based on a similarity between the extracted entity and an entity of an entity node included in a first index database.

[0008] According to one embodiment of the present disclosure, the first index database includes a data node associated with source data stored in a source database and at least one entity node generated from the source data using a machine learning model, and the data node and the entity node can be connected by edges using a logical data model.

[0009] According to one embodiment of the present disclosure, a data node includes an identifier associated with source data and text embedding vector information associated with the source data, an entity node includes entity and graph embedding vector information, and a text embedding vector transformed from the source data and a graph embedding vector transformed from the data node and the entity node can be stored in a second index database.

[0010] According to one embodiment of the present disclosure, the step of extracting data associated with a first entity node may include the step of determining a first data node associated with the first entity node and the step of extracting data associated with the first data node from a source database.

[0011] According to one embodiment of the present disclosure, the step of extracting data associated with a first entity node may include the step of determining a second entity node based on graph embedding vector information of the first entity node, the step of determining a second data node associated with the second entity node, and the step of extracting data associated with the second data node from a source database.

[0012] According to one embodiment of the present disclosure, the step of extracting data associated with a first entity node may include the step of determining a first data node associated with the first entity node, the step of determining a second data node based on text embedding vector information of the first data node, and the step of extracting data associated with the second data node from a source database.

[0013] According to one embodiment of the present disclosure, an entity node is connected by an edge to at least one of a plurality of data nodes or a plurality of entity nodes stored in a first index database by a logical data model, and the step of extracting data associated with the first entity node may include the step of determining a second entity node connected to the first entity node, the step of determining a first data node connected to the second entity node, and the step of extracting data associated with the first data node from a source database.

[0014] A computer-readable non-transitory recording medium having recorded thereon instructions for executing a method according to one embodiment of the present disclosure on a computer may be provided.

[0015] A system according to one embodiment of the present disclosure comprises a communication module, a memory, and at least one processor connected to the memory and configured to execute at least one computer-readable program contained in the memory, wherein the at least one program may include instructions for receiving a natural language query, determining a first entity node stored in a first index database formed in a graph structure based on the natural language query, extracting data associated with the first entity node from a source database, and outputting a result for the received natural language query using the data associated with the first entity.

[0016] According to some embodiments of the present disclosure, entity nodes associated with a received natural language query can be determined as an index. These entity nodes can provide data similar to the natural language query, as well as data connected by a logical data model on the graph. Consequently, users can receive richer search results.

[0017] According to some embodiments of the present disclosure, not only text embedding vectors but also node and graph embedding vectors within a graph can be generated as indices for full text data. Accordingly, input text data can be systematically stored from both a semantic and structural perspective.

[0018] According to some embodiments of the present disclosure, two types of indexes can provide information on natural language queries from both a semantic and structural perspective. Consequently, users can receive richer search results.

[0019] The effects of the present disclosure are not limited to the effects mentioned above, and other effects not mentioned can be clearly understood by a person having ordinary skill in the art to which the present disclosure belongs (referred to as “one skilled in the art”) from the description of the claims.

[0020] Embodiments of the present disclosure will be described below with reference to the accompanying drawings, wherein like reference numerals represent similar elements, but are not limited thereto.

[0021] FIG. 1 illustrates an example of an information providing method according to one embodiment of the present disclosure.

[0022] FIG. 2 is a schematic diagram showing a configuration in which an information processing system is connected to enable communication with a plurality of user terminals to provide information according to one embodiment of the present disclosure.

[0023] FIG. 3 is a block diagram showing the internal configuration of a user terminal and an information processing system according to one embodiment of the present disclosure.

[0024] FIG. 4 is a diagram illustrating an example of a method for generating an index according to one embodiment of the present disclosure.

[0025] FIG. 5 is a diagram illustrating an example of a method for providing query results for a natural language query according to one embodiment of the present disclosure.

[0026] FIG. 6 is a diagram illustrating an example of data stored in a database according to one embodiment of the present disclosure.

[0027] FIG. 7 is a diagram illustrating an example of a data configuration stored in an index database of a graph structure according to one embodiment of the present disclosure.

[0028] FIG. 8 is a flowchart illustrating an example of a method according to one embodiment of the present disclosure.

[0029] Hereinafter, specific details for implementing the present disclosure will be described in detail with reference to the attached drawings. However, in the following description, specific descriptions of widely known functions or configurations will be omitted if they may unnecessarily obscure the gist of the present disclosure.

[0030] In the attached drawings, identical or corresponding components are assigned the same reference numerals. Furthermore, in the description of the embodiments below, duplicate descriptions of identical or corresponding components may be omitted. However, even if a description of a component is omitted, it is not intended that such component is not included in any embodiment.

[0031] The advantages and features of the disclosed embodiments, and methods for achieving them, will become clearer with reference to the embodiments described below, along with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below and may be implemented in various different forms. These embodiments are provided solely to ensure the completeness of the disclosure and to fully inform those skilled in the art of the scope of the invention.

[0032] The terms used in this specification will be briefly explained, followed by a detailed description of the disclosed embodiments. The terms used in this specification have been selected from widely used, current terms, taking into account the functions of the present disclosure. However, these terms may vary depending on the intentions of engineers working in the relevant field, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, terms may be arbitrarily selected by the applicant, and in such cases, their meanings will be described in detail in the relevant description of the invention. Therefore, the terms used in this disclosure should not be defined simply as names of terms, but rather based on their meanings and the overall content of the present disclosure.

[0033] In this specification, singular expressions include plural expressions unless the context clearly indicates otherwise. Furthermore, plural expressions include singular expressions unless the context clearly indicates otherwise. When a part of the specification is said to include a component, this does not exclude other components, but rather implies that other components may be included, unless otherwise specifically stated.

[0034] Also, the term 'module' or 'part' used in the specification means a software or hardware component, and the 'module' or 'part' performs certain roles. However, the 'module' or 'part' is not limited to software or hardware. The 'module' or 'part' may be configured to reside on an addressable storage medium and may be configured to execute one or more processors. Thus, as an example, the 'module' or 'part' may include at least one of components such as software components, object-oriented software components, class components, and task components, processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuitry, data, databases, data structures, tables, arrays, or variables. The functionality provided within the components and 'modules' or 'parts' may be combined into a smaller number of components and 'modules' or 'parts', or further separated into additional components and 'modules' or 'parts'.

[0035] According to one embodiment of the present disclosure, a 'module' or 'unit' may be implemented as a processor and a memory. 'Processor' should be broadly construed to include a general-purpose processor, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a controller, a microcontroller, a state machine, and the like. In some circumstances, a 'processor' may also refer to an application-specific integrated circuit (ASIC), a programmable logic device (PLD), a field-programmable gate array (FPGA), and the like. A 'processor' may also refer to a combination of processing devices, such as, for example, a combination of a DSP and a microprocessor, a combination of multiple microprocessors, a combination of one or more microprocessors in conjunction with a DSP core, or any other such combination of configurations. In addition, 'memory' should be broadly construed to include any electronic component capable of storing electronic information. 'Memory' may refer to various types of processor-readable media, such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, magnetic or optical data storage, registers, etc. Memory is said to be in electronic communication with the processor if the processor can read information from, and / or write information to, the memory. Memory integrated in a processor is in electronic communication with the processor.

[0036] In the present disclosure, the "system" may include, but is not limited to, at least one of a server device and a cloud device. For example, the system may be comprised of one or more server devices. As another example, the system may be comprised of one or more cloud devices. As yet another example, the system may be configured and operated by a combination of a server device and a cloud device.

[0037] In the present disclosure, 'display' may refer to any display device associated with a computing device, for example, any display device capable of displaying any information / data controlled by or provided from the computing device.

[0038] In the present disclosure, 'each of the plurality of As' or 'each of the plurality of As' may refer to each of all components included in the plurality of As, or may refer to each of some components included in the plurality of As.

[0039] In this disclosure, "data" can refer to anything that can be expressed in the form of information, such as facts, numbers, letters, images, or sounds. Furthermore, data may exist on the system in raw or processed form. In this disclosure, the terms "data" and "data item" may be used interchangeably to refer to the same or similar information. Furthermore, in this disclosure, "data" may include multiple data items.

[0040] In this disclosure, a "machine learning model" may include any model used to infer an answer to a given input. In one embodiment, the machine learning model may include an artificial neural network model including an input layer, multiple hidden layers, and an output layer. Here, each layer may include multiple nodes. In this disclosure, each of the multiple machine learning models is described as a separate machine learning model, but this is not limited thereto, and some or all of the multiple machine learning models may be implemented as a single machine learning model. Furthermore, a single machine learning model may include multiple machine learning models. In this disclosure, the terms "machine learning model" and "artificial neural network model" may be used interchangeably to refer to the same or similar models.

[0041] In the present disclosure, a "large language model (LLM)" may refer to a language model capable of inference without fine-tuning using a method such as few-shot learning, and may have more than 10 times as many parameters (e.g., more than 100 billion parameters) as a conventional general language model. In the present disclosure, a "language model" may include a large language model.

[0042] FIG. 1 illustrates an example of an information provision method according to one embodiment of the present disclosure. In one embodiment, a processor (e.g., at least one processor of an information processing system) may receive a natural language query (110). Here, the natural language query (110) may refer to full text composed of natural language input by a user to search for specific information.

[0043] In one embodiment, the processor may determine entity nodes stored in an index database (120) formed in a graph structure based on a natural language query (110). Specifically, the processor may extract entities from the natural language query (110) using a machine learning model. In this case, the processor may search the index database (120) and determine entity nodes containing entities similar to the extracted entities as an index.

[0044] In one embodiment, the index database (120) may include data nodes (124) associated with source data stored in the source database and at least one entity node (122) generated from the source data using a machine learning model. Furthermore, the data nodes (124) and entity nodes (122) may be connected by edges (126) using a logical data model. An example of a graph in which nodes are connected by edges is described in detail below with reference to FIG. 7 .

[0045] In one embodiment, the processor may extract data associated with an entity node from a source database. Furthermore, the processor may output a query result (130) for a natural language query (110) using the data associated with the entity node. Here, the data associated with the entity node may be determined based on nodes included in the graph, vectors associated with the nodes, and the like. An example of extracting data associated with an entity node is described in detail below with reference to FIG. 7.

[0046] This configuration allows entity nodes associated with received natural language queries to be determined as indexes. These entity nodes can provide data similar to the natural language query, as well as data linked by a logical data model within the graph. Consequently, users can receive richer search results.

[0047] FIG. 2 is a schematic diagram illustrating a configuration in which an information processing system (230) is connected to a plurality of user terminals (210_1, 210_2, 210_3) to enable communication with each other in order to provide information according to one embodiment of the present disclosure. As illustrated, the plurality of user terminals (210_1, 210_2, 210_3) may be connected to an information processing system (230) capable of providing an information provision service via a network (220). Here, the plurality of user terminals (210_1, 210_2, 210_3) may include terminals of users receiving the information provision service.

[0048] In one embodiment, the information processing system (230) may include one or more server devices and / or databases capable of storing, providing, and executing computer executable programs (e.g., downloadable applications) and data associated with providing information provision services, or one or more distributed computing devices and / or distributed databases based on cloud computing services.

[0049] The information provision service provided by the information processing system (230) can be provided to users through information provision service applications, web browsers, web browser extension programs, etc. installed on each of a plurality of user terminals (210_1, 210_2, 210_3). For example, the information processing system (230) can provide information corresponding to a search request for a natural language query received from a user terminal (210_1, 210_2, 210_3) or perform corresponding processing through the information provision service application, etc.

[0050] A plurality of user terminals (210_1, 210_2, 210_3) can communicate with an information processing system (230) via a network (220). The network (220) can be configured to enable communication between the plurality of user terminals (210_1, 210_2, 210_3) and the information processing system (230). Depending on the installation environment, the network (220) can be configured as a wired network such as Ethernet, a wired home network (Power Line Communication), a telephone line communication device, and RS-serial communication, a wireless network such as a mobile communication network, WLAN (Wireless LAN), Wi-Fi, Bluetooth, and ZigBee, or a combination thereof. The communication method is not limited, and may include not only a communication method utilizing a communication network (e.g., a mobile communication network, wired Internet, wireless Internet, broadcasting network, satellite network, etc.) that the network (220) may include, but also short-range wireless communication between user terminals (210_1, 210_2, 210_3).

[0051] In FIG. 2, a mobile phone terminal (210_1), a tablet terminal (210_2), and a PC terminal (210_3) are illustrated as examples of user terminals, but are not limited thereto, and the user terminals (210_1, 210_2, 210_3) may be any computing device capable of wired and / or wireless communication and capable of installing and executing an information provision service application or a web browser. For example, the user terminal may include an AI speaker, a smartphone, a mobile phone, a navigation device, a computer, a laptop, a digital broadcasting terminal, a PDA (Personal Digital Assistants), a PMP (Portable Multimedia Player), a tablet PC, a game console, a wearable device, an IoT (Internet of Things) device, a VR (virtual reality) device, an AR (augmented reality) device, a set-top box, and the like. In addition, although FIG. 2 illustrates three user terminals (210_1, 210_2, 210_3) communicating with the information processing system (230) via the network (220), this is not limited thereto, and a different number of user terminals may be configured to communicate with the information processing system (230) via the network (220).

[0052] In FIG. 2, a configuration in which a user's request is transmitted to an information processing system (230) through a user terminal (210_1, 210_2, 210_3) is exemplarily illustrated, but the present invention is not limited thereto. The user's request may be provided to the information processing system (230) through an input device associated with the information processing system (230) without passing through the user terminal (210_1, 210_2, 210_3), and the result of processing the user's request may be provided to the user through an output device (e.g., a display, etc.) associated with the information processing system (230).

[0053] Although FIG. 2 illustrates that user terminals (210_1, 210_2, 210_3) receive information provision services from the information processing system (230), the present invention is not limited thereto. For example, information provision services may be provided through information provision programs / applications installed on user terminals (210_1, 210_2, 210_3) without communication with the information processing system (230). Furthermore, although the information processing system (230) is illustrated as a single device, the present invention is not limited thereto, and the information processing system (230) may be composed of multiple devices.

[0054] FIG. 3 is a block diagram showing the internal configuration of a user terminal (210) and an information processing system (230) according to one embodiment of the present disclosure. The user terminal (210) may refer to any computing device capable of executing applications, web browsers, etc., and capable of wired / wireless communication, and may include, for example, a mobile phone terminal (210_1), a tablet terminal (210_2), a PC terminal (210_3), etc. of FIG. 2. As illustrated, the user terminal (210) may include a memory (312), a processor (314), a communication module (316), and an input / output interface (318). Similarly, the information processing system (230) may include a memory (332), a processor (334), a communication module (336), and an input / output interface (338). As illustrated in FIG. 3, the user terminal (210) and the information processing system (230) may be configured to communicate information and / or data via a network (220) using respective communication modules (316, 336). In addition, the input / output device (320) may be configured to input information and / or data to the user terminal (210) or output information and / or data generated from the user terminal (210) via the input / output interface (318).

[0055] The memory (312, 332) may include any non-transitory computer-readable recording medium. According to one embodiment, the memory (312, 332) may include a permanent mass storage device such as a read-only memory (ROM), a disk drive, a solid-state drive (SSD), or flash memory. As another example, a permanent mass storage device such as a ROM, an SSD, a flash memory, or a disk drive may be included in the user terminal (210) or the information processing system (230) as a separate permanent storage device distinct from the memory. In addition, an operating system and at least one program code may be stored in the memory (312, 332).

[0056] These software components may be loaded from a computer-readable recording medium separate from the memory (312, 332). This separate computer-readable recording medium may include a recording medium directly connectable to the user terminal (210) and the information processing system (230), and may include, for example, a computer-readable recording medium such as a floppy drive, a disk, a tape, a DVD / CD-ROM drive, a memory card, etc. As another example, the software components may be loaded into the memory (312, 332) through a communication module (316, 336) other than a computer-readable recording medium. For example, at least one program may be loaded into the memory (312, 332) based on a computer program that is installed by files provided by developers or a file distribution system that distributes installation files of applications through a network (220).

[0057] The processor (314, 334) may be configured to process instructions of a computer program by performing basic arithmetic, logic, and input / output operations. Instructions may be provided to the processor (314, 334) by a memory (312, 332) or a communication module (316, 336). For example, the processor (314, 334) may be configured to execute instructions received according to program code stored in a storage device such as the memory (312, 332).

[0058] The communication module (316, 336) may provide a configuration or function for the user terminal (210) and the information processing system (230) to communicate with each other via the network (220), and may provide a configuration or function for the user terminal (210) and / or the information processing system (230) to communicate with another user terminal or another system (e.g., a separate cloud system, etc.). For example, a request or data (e.g., a search request for a natural language query, etc.) generated by the processor (314) of the user terminal (210) according to a program code stored in a recording device such as a memory (312) may be transmitted to the information processing system (230) via the network (220) under the control of the communication module (316). Conversely, a control signal or command provided under the control of the processor (334) of the information processing system (230) can be received by the user terminal (210) through the communication module (316) of the user terminal (210) via the communication module (336) and the network (220).

[0059] The input / output interface (318) may be a means for interfacing with an input / output device (320). As an example, the input device may include a device such as a camera, keyboard, microphone, mouse, etc., including an audio sensor and / or an image sensor, and the output device may include a device such as a display, a speaker, a haptic feedback device, etc. As another example, the input / output interface (318) may be a means for interfacing with a device that has a configuration or function integrated into one for performing input and output, such as a touch screen. For example, when the processor (314) of the user terminal (210) processes a command of a computer program loaded into the memory (312), a service screen configured using information and / or data provided by the information processing system (230) or another user terminal may be displayed on the display through the input / output interface (318). In FIG. 3, the input / output device (320) is illustrated as not being included in the user terminal (210), but is not limited thereto, and may be configured as a single device with the user terminal (210). In addition, the input / output interface (338) of the information processing system (230) may be a means for interfacing with a device (not shown) for input or output that is connected to the information processing system (230) or that the information processing system (230) may include. In FIG. 3, the input / output interfaces (318, 338) are illustrated as elements configured separately from the processors (314, 334), but are not limited thereto, and the input / output interfaces (318, 338) may be configured to be included in the processors (314, 334).

[0060] The user terminal (210) and the information processing system (230) may include more components than those shown in FIG. 3. However, there is no need to explicitly illustrate most of the conventional components. In one embodiment, the user terminal (210) may be implemented to include at least some of the input / output devices (320) described above. In addition, the user terminal (210) may further include other components, such as a transceiver, a Global Positioning System (GPS) module, a camera, various sensors, a database, and the like. For example, if the user terminal (210) is a smartphone, it may include components that a smartphone generally includes, and various components, such as an acceleration sensor, a gyro sensor, a microphone module, a camera module, various physical buttons, buttons using a touch panel, input / output ports, and a vibrator for vibration, may be implemented to be further included in the user terminal (210).

[0061] While a program or application for providing information services, etc. is in operation, the processor (314) can receive text, images, videos, voices and / or actions, etc. input or selected through input devices such as a camera, microphone, including a touch screen, keyboard, audio sensor and / or image sensor connected to an input / output interface (318), and can store the received text, images, videos, voices and / or actions, etc. in a memory (312) or provide them to an information processing system (230) through a communication module (316) and a network (220).

[0062] The processor (314) of the user terminal (210) may be configured to manage, process, and / or store information and / or data received from an input / output device (320), another user terminal, an information processing system (230), and / or multiple external systems. The information and / or data processed by the processor (314) may be provided to the information processing system (230) via a communication module (316) and a network (220). The processor (314) of the user terminal (210) may transmit the information and / or data to the input / output device (320) via an input / output interface (318) and output the information and / or data. For example, the processor (314) may output or display the received information and / or data on a screen associated with the user terminal (210).

[0063] The processor (334) of the information processing system (230) may be configured to manage, process, and / or store information and / or data received from multiple user terminals (210) and / or multiple external systems. Information and / or data processed by the processor (334) may be provided to the user terminal (210) via a communication module (336) and a network (220).

[0064] FIG. 4 is a diagram illustrating an example of a method for generating an index according to one embodiment of the present disclosure. In one embodiment, the processor (334) may generate a full-text index. Here, the full-text index is index information for unstructured text (or natural language) data, and may be stored in the form of nodes and / or vectors. In addition, the processor (334) may include a data change management unit (430), a graph transformation unit (440), and a vector transformation unit (450).

[0065] The data change management unit (430) can detect changes in data within the source database (420) and request index creation. For example, when new text data (410) is stored in the source database (420), the data change management unit (430) can request the graph transformation unit (440) and the vector transformation unit (450) to create an index for the text data (410) stored in the source database (420).

[0066] The graph conversion unit (440) can convert text data (410) into a graph structure. Specifically, the graph conversion unit (440) can generate data nodes associated with the text data (410). In addition, the graph conversion unit (440) can generate entity nodes by extracting entities from the text data (410) using a pre-trained machine learning model (e.g., a very large language model, etc.). Additionally, the graph conversion unit (440) can connect the generated data nodes and entity nodes with edges using a logical data model. Here, the logical data model can refer to a model that expresses relationships between nodes according to a predefined schema (e.g., an inclusion relationship, a synonym relationship, a connection relationship, a dependency relationship, etc.). For example, the graph conversion unit (440) can connect the generated data nodes and entity nodes with edges by indicating an "inclusion relationship."

[0067] In one embodiment, the graph transformation unit (440) may store the generated data nodes and entity nodes in a first index database (442). Here, the first index database (442) may be formed in a graph structure. In this case, the graph transformation unit (440) may store the data nodes, entity nodes, and edges connecting the data nodes and entity nodes together in the first index database (442).

[0068] In one embodiment, the graph transformation unit (440) may merge the previously stored graph of the first index database (442) with the newly created data nodes and entity nodes. In this case, the graph transformation unit (440) may use a logical data model to connect the newly stored entity nodes with the data nodes and / or entity nodes in the previously stored graph as edges. Similarly, the graph transformation unit (440) may use a logical data model to connect the newly stored data nodes with the entity nodes in the previously stored graph as edges.

[0069] In one embodiment, a data node may include an identifier of the data node, an identifier (or primary key) associated with text data (410), text embedding vector information associated with the text data (410), etc. Here, the identifier associated with the text data (410) may mean an identifier for accessing the text data (410) stored in the source database (420). In addition, the text embedding vector information associated with the text data (410) may include a pointer indicating a text embedding vector into which the text data (410) is converted, stored in the second index database (452).

[0070] In one embodiment, an entity node may include an identifier of the entity node, an entity extracted from text data (410), graph embedding vector information, etc. Here, the entity may mean an entity extracted from the text data (410) using a pre-trained machine learning model (e.g., a very large language model, etc.). In addition, the graph embedding vector information may include a pointer indicating a graph embedding vector converted from an entity node stored in the second index database (452).

[0071] The vector conversion unit (450) can convert text data (410) into a text embedding vector. Specifically, the vector conversion unit (450) can convert text data (410) into a vector using a pre-trained text embedding model. Here, the pre-trained text embedding model can refer to any machine learning model trained to convert input text into an embedding vector.

[0072] In one embodiment, the vector conversion unit (450) may store the converted text embedding vector in the second index database (452). In addition, the vector conversion unit (450) may sort the converted text embedding vector based on similarity. In this case, the vector conversion unit (450) may calculate the similarity between vectors through Euclidean distance, cosine similarity, etc., but is not limited thereto.

[0073] In one embodiment, the vector conversion unit (450) can convert the data nodes and entity nodes generated by the graph conversion unit (440) into graph embedding vectors. Specifically, the vector conversion unit (450) can convert the data nodes and entity nodes into vectors using a pre-trained graph embedding model. Here, the pre-trained graph embedding model can refer to any machine learning model trained to convert an input graph into an embedding vector. The "Node2Vec" algorithm can be used to convert the graph into an embedding vector, but is not limited thereto. Accordingly, the vector conversion unit (450) can sort the converted graph embedding vectors based on similarity. That is, nodes connected on the graph can be mapped to close locations in the vector space.

[0074] In one embodiment, when specific text data is deleted from the source database (420), the data change management unit (430) can synchronize the source database (420), the first index database (442), and the second index database (452). For example, the data change management unit (430) can access the first index database (442) and remove the data node corresponding to the deleted text data and the edge connected to the data node. In addition, the data change management unit (430) can access the second index database (452) and delete the graph embedding vector of the deleted data node and the text embedding vector converted from the deleted text data.

[0075] In one embodiment, the data change management unit (430) can traverse the entity nodes in the first index database (442) to remove specific entity nodes. Specifically, the data change management unit (430) can access the entity nodes of the first index database (442) by referencing the graph embedding vector of the second index database (452). Alternatively, the data change management unit (430) can directly access the entity nodes of the first index database (442). In this case, the data change management unit (430) can remove at least one data node and a specific entity node that is not connected to at least one entity node by an edge among the entity nodes of the second index database (452). Additionally, the data change management unit (430) can access the second index database (452) to delete the graph embedding vector into which the removed entity node is transformed. Removal of these entity nodes can be performed synchronously / asynchronously or periodically.

[0076] This configuration allows for the generation of not only text embedding vectors as indices for full text data, but also node and graph embedding vectors within the graph. Consequently, input text data can be systematically stored from both a semantic and structural perspective.

[0077] FIG. 5 is a diagram illustrating an example of a method for providing a query result (540) for a natural language query (510) according to one embodiment of the present disclosure. In one embodiment, the processor (334) may provide a query result (540) for a natural language query (510). Here, the natural language query (510) may refer to text in a natural language form (e.g., in the form of a word, sentence, query, etc.) that a user searches for to obtain specific information. In addition, the processor (334) may include an index search unit (520) and an information provision unit (530).

[0078] When a natural language query (510) is received, the index search unit (520) can extract entities from the natural language query (510) using a pre-trained machine learning model (e.g., a very large language model, etc.). In addition, the index search unit (520) can determine a first entity node that includes an entity similar to the extracted entity through a graph traversal, etc. in the first index database (522). The index search unit (520) can transmit an index associated with the extracted first entity node to the information provision unit (530).

[0079] In one embodiment, the index search unit (520) may determine a first data node associated with a first entity node in the first index database (522). In this case, the index search unit (520) may transmit an index associated with the first data node to the information provider (530).

[0080] In one embodiment, the index search unit (520) may determine a second entity node based on graph embedding vector information of the first entity node. Specifically, the index search unit (520) may determine a second graph embedding vector similar to the first graph embedding vector of the first entity node from the second index database (524). In this case, the index search unit (520) may determine a second entity node associated with the determined second graph embedding vector. In addition, the index search unit (520) may determine a second data node associated with the second entity node from the first index database (522). Accordingly, the index search unit (520) may transmit an index associated with the second data node to the information provider (530).

[0081] In one embodiment, the index search unit (520) may determine a first data node connected to a first entity node from a first index database (522). Furthermore, the index search unit (520) may determine a second data node based on text embedding vector information of the first data node. Specifically, the index search unit (520) may determine a second text embedding vector similar to the first text embedding vector of the first data node from a second index database (524). In this case, the index search unit (520) may determine a second data node associated with the determined second text embedding vector. Accordingly, the index search unit (520) may transmit an index associated with the second data node to the information provider (530).

[0082] In one embodiment, the index search unit (520) may determine a second entity node connected to a first entity node in the first index database (522). Furthermore, the index search unit (520) may determine a first data node connected to the second entity node in the first index database (522). Accordingly, the index search unit (520) may transmit an index associated with the first data node to the information provider (530).

[0083] The information providing unit (530) can provide a query result (540) for a natural language query (510). Specifically, the information providing unit (530) can receive an index from the index search unit (520). In this case, the information providing unit (530) can extract data associated with the index from the source database (532). Accordingly, the information providing unit (530) can output the extracted data as a query result (540).

[0084] This configuration allows two types of indexes to provide information on natural language queries from both a semantic and structural perspective. Consequently, users can receive richer search results.

[0085] FIG. 6 is a diagram illustrating examples of data stored in a database according to one embodiment of the present disclosure. A first data example (610) illustrates an example of data stored in a source database. In one embodiment, the source database may include source data (614) and an identifier (612) of the source data (614). For example, the identifier (612) may be "111-aaa," and the source data (614) may be "This product can machine-process natural language."

[0086] A second data example (630) illustrates an example of data stored in a first index database (e.g., 442 of FIG. 4). In one embodiment, the first index database may include a data node (632) associated with source data (614) and an entity node (634) comprising entities extracted from the source data (614) using a machine learning model. Additionally, the first index database may include an edge (636) connecting the data node (632) and the entity node (634) using a logical data model.

[0087] A third data example (620) illustrates an example of data stored in a second index database (e.g., 452 of FIG. 4). In one embodiment, the second index database may include a text embedding vector (622) and a graph embedding vector (624). Here, the text embedding vector (622) may be generated and stored by transforming the original data (614) using a pre-trained text embedding model. For example, the text embedding vector (622) converted from the original data (614) may be an n-dimensional vector, such as "[0.234, 0.255, 쪋]". In addition, the graph embedding vector (624) may be generated and stored by transforming data nodes and / or entity nodes stored in the first index database using a pre-trained graph embedding model. For example, the graph embedding vector (624) converted from the entity node (634) may be an m-dimensional vector, such as "[0.123, 0.124, 쪋]". The text embedding vector may represent a semantic viewpoint of the source data (614) (e.g., similarity between data), and the graph embedding vector may represent a structural viewpoint (e.g., structure and relationship between data).

[0088] In one embodiment, the data node (632) may include a node identifier, an identifier associated with the source data (614), and text embedding vector information. For example, the node identifier (or Node ID) of the data node (632) may be “abc,” and the identifier (or Data ID) associated with the source data (614) may be “111-aaa” as the identifier (612) of the source data (614). Additionally, the text embedding vector information (or txt embedding) may be a pointer indicating “[0.234, 0.255, 쪋],” which is a converted text embedding vector of the source data (614) stored in the second index database.

[0089] In one embodiment, the entity node (634) may include a node identifier, entity information extracted from the source data (614), and graph embedding vector information. For example, the node identifier (or Node ID) of the entity node (634) may be “abcd,” and the entity information (or Name) extracted from the source data (614) may be “natural language.” Here, the entity information may include words, phrases, etc. associated with the source data (614). In addition, the graph embedding vector information (or graph embedding) may be a pointer indicating "[0.123, 0.124, 쪋]," which is a converted graph embedding vector of the entity node (634) stored in the second index database.

[0090] Although FIG. 6 illustrates that the data node (632) does not include index information associated with the graph embedding vector, this is not limited thereto. For example, the data node (632) may also include graph embedding vector information, like the entity node (634).

[0091] FIG. 7 is a diagram illustrating an example of a data structure stored in an index database having a graph structure according to one embodiment of the present disclosure. In one embodiment, the index database having a graph structure (e.g., 522 of FIG. 5) may include a plurality of data nodes (710, 720, 730) and a plurality of entity nodes (717, 716, 724, 734). Here, each of the plurality of data nodes (710, 720, 730) may be connected to a plurality of source data (712, 722, 732) through an identifier associated with the source data. For example, a first data node (710) may be connected to a first source data (712), a second data node (720) may be connected to a second source data (722), and a third data node (730) may be connected to a third source data (732). In this way, one data node can be connected one-to-one with one source data.

[0092] In one embodiment, each of the plurality of data nodes (710, 720, 730) may be connected to at least one entity node. For example, a first data node (710) may be edge-connected to a first entity node (714) and a second entity node (716) by a logical data model, a second data node (720) may be edge-connected to a second entity node (716) and a third entity node (724), and a third data node (730) may be edge-connected to a fourth entity node (734).

[0093] In one embodiment, two nodes may be connected by edges through a logical data model that expresses relationships between nodes according to a predefined schema (e.g., an inclusion relationship, a synonym relationship, a connection relationship, a dependency relationship, etc.). For example, a first entity node (714) and a second entity node (716) may be connected by edges to a first data node (710) as an "inclusion relationship." In another example, a second entity node (716) may be connected by edges to a fourth entity node (734) as an "synonym relationship."

[0094] In one embodiment, an index may be determined in an index database to provide information based on a natural language query entered by a user. For example, if the second entity node (716) includes entities most similar to the entities extracted from the natural language query, the first data node (710) and the second data node (720) associated with the second entity node (716) may be determined as indexes. Accordingly, the first source data (712) and the second source data (722) associated with the first data node (710) and the second data node (720) may be provided as query results.

[0095] As another example, if the first entity node (714) includes an entity most similar to an entity extracted from a natural language query, a second entity node (716) including a graph embedding vector similar to the graph embedding vector of the first entity node (714) may be selected. In this case, the first data node (710) and the second data node (720) connected to the second entity node (716) may be determined as indices. Accordingly, the first source data (712) and the second source data (722) connected to the first data node (710) and the second data node (720) may be provided as a query result.

[0096] As another example, if the first entity node (714) includes an entity most similar to an entity extracted from a natural language query, the first data node (710) connected to the first entity node (714) may be selected. In this case, the second data node (720) including a text embedding vector similar to the text embedding vector of the first data node (710) may be determined as an index. Accordingly, the first source data (712) and / or the second source data (722) connected to the first data node (710) and the second data node (720) may be provided as a query result.

[0097] As another example, if the second entity node (716) includes an entity most similar to the entity extracted from the natural language query, the fourth entity node (734) connected to the second entity node (716) may be selected. In this case, the third data node (730) connected to the fourth entity node (734) may be determined as an index. Accordingly, the third source data (732) connected to the third data node (730) may be provided as a query result.

[0098] In one embodiment, the provided query results may be determined based on the order of nodes reached first through graph traversal. For example, if the second data node (720) is determined to be indexed before the third data node (730) through graph traversal, the second source data (722) may be provided as a query result before the third source data (732).

[0099] In one embodiment, entity nodes not connected by edges in an index database formed in a graph structure may be removed. Specifically, when source data is deleted, data nodes corresponding to the deleted source data and edges connected to the data nodes may be removed. Furthermore, entity nodes not connected by edges to at least one data node and at least one entity node within the graph may be removed through entity node traversal. This entity node removal may be performed synchronously, asynchronously, or periodically.

[0100] In Fig. 7, it is described that data nodes and entity nodes are connected by edges by indicating an “inclusion relationship” or a “synonym relationship,” but this is not limited to this, and they may be connected by edges by indicating various relationships such as a “dependency relationship.”

[0101] FIG. 8 is a flowchart illustrating an example of a method (800) according to one embodiment of the present disclosure. In one embodiment, the method (800) may be performed by at least one processor. The method (800) may begin with the processor receiving a natural language query (S810).

[0102] Thereafter, the processor may determine a first entity node stored in a first index database formed in a graph structure based on the natural language query (S820). Specifically, the processor may extract entities from the natural language query using a machine learning model. Furthermore, the processor may determine the first entity node based on the similarity between the extracted entities and entities included in the entity nodes of the first index database.

[0103] Thereafter, the processor can extract data associated with the first entity node from the source database (S830). Additionally, the processor can output a result for the received natural language query using the data associated with the first entity (S840).

[0104] In one embodiment, the first index database may include data nodes associated with source data stored in the source database, and at least one entity node generated from the source data using a machine learning model. Here, the data nodes and entity nodes may be connected by edges using a logical data model. Furthermore, the data nodes may include identifiers associated with the source data and text embedding vector information associated with the source data, and the entity nodes may include entity and graph embedding vector information. In this case, the text embedding vectors obtained by transforming the source data and the graph embedding vectors obtained by transforming the data nodes and entity nodes may be stored in the second index database.

[0105] In one embodiment, the processor may determine a first data node associated with a first entity node. Additionally, the processor may extract data associated with the first data node from a source database.

[0106] In one embodiment, the processor may determine a second entity node based on the graph embedding vector information of the first entity node. Furthermore, the processor may determine a second data node associated with the second entity node. Additionally, the processor may extract data associated with the second data node from the source database.

[0107] In one embodiment, the processor may determine a first data node associated with a first entity node. Furthermore, the processor may determine a second data node based on the text embedding vector information of the first data node. Additionally, the processor may extract data associated with the second data node from the source database.

[0108] In one embodiment, an entity node may be edge-connected to at least one of a plurality of data nodes or a plurality of entity nodes stored in a first index database, based on a logical data model. In this case, the processor may determine a second entity node connected to the first entity node. Furthermore, the processor may determine a first data node connected to the second entity node. Additionally, the processor may extract data associated with the first data node from the source database.

[0109] The above-described method may be provided as a computer program stored on a computer-readable recording medium for execution on a computer. The medium may be one that continuously stores a computer-executable program or one that temporarily stores it for execution or download. In addition, the medium may be various recording means or storage means in the form of a single or multiple hardware combinations, and is not limited to a medium directly connected to a computer system, but may also be distributed over a network. Examples of the medium may include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and those configured to store program instructions, including ROM, RAM, and flash memory. In addition, examples of other media may include recording or storage media managed by app stores that distribute applications, sites that supply or distribute various software, servers, etc.

[0110] The methods, operations, or techniques of the present disclosure may be implemented by various means. For example, these techniques may be implemented in hardware, firmware, software, or a combination thereof. Those skilled in the art will appreciate that the various exemplary logical blocks, modules, circuits, and algorithm steps described in connection with the disclosure herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various exemplary components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software will depend on the particular application and the design requirements imposed on the overall system. Those skilled in the art may implement the described functionality in various ways for each particular application, but such implementations should not be construed as departing from the scope of the present disclosure.

[0111] In a hardware implementation, the processing units used to perform the techniques may be implemented within one or more ASICs, DSPs, GPUs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, electronic devices, other electronic units designed to perform the functions described herein, a computer, or a combination thereof.

[0112] Accordingly, the various exemplary logical blocks, modules, and circuits described in connection with the present disclosure may be implemented or performed by any combination of a general-purpose processor, a DSP, an ASIC, an FPGA or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or those designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.

[0113] In a firmware and / or software implementation, the techniques may be implemented as instructions stored on a computer-readable medium, such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, a compact disc (CD), a magnetic or optical data storage device, etc. The instructions may be executable by one or more processors and may cause the processor(s) to perform certain aspects of the functionality described herein.

[0114] When implemented in software, the techniques may be stored on or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media includes both computer storage media and communication media, including any medium that facilitates transfer of a computer program from one place to another. Storage media may be any available media that can be accessed by a computer. By way of example, and not limitation, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer. In addition, any connection is suitably made to a computer-readable medium.

[0115] For example, if the software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, digital subscriber line, or wireless technologies such as infrared, radio, and microwave are included within the definition of media. Disk and disc, as used herein, includes compact discs, laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where disks usually reproduce data magnetically, whereas discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0116] A software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium may be coupled to the processor such that the processor can read information from, and write information to, the storage medium. Alternatively, the storage medium may be integral to the processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. Alternatively, the processor and the storage medium may reside as discrete components in the user terminal.

[0117] While the embodiments described above have been described as utilizing aspects of the presently disclosed subject matter in one or more standalone computer systems, the present disclosure is not limited thereto and may be implemented in conjunction with any computing environment, such as a network or distributed computing environment. Furthermore, aspects of the present disclosure may be implemented in multiple processing chips or devices, and storage may be similarly affected across multiple devices. Such devices may include personal computers, network servers, and portable devices.

[0118] While the present disclosure has been described in connection with certain embodiments herein, various modifications and variations may be made without departing from the scope of the present disclosure, which would be apparent to those skilled in the art. Furthermore, such modifications and variations are intended to fall within the scope of the claims appended to this specification.

Claims

1. A method for providing information using a full-text index, performed by at least one processor, Step of receiving a natural language query; A step of determining a first entity node stored in a first index database formed in a graph structure based on the natural language query; A step of extracting data associated with the first entity node from a source database; and A step of outputting a result for the received natural language query using data associated with the first entity. A method of providing information, including:

2. In paragraph 1, The step of determining the first entity node comprises: A step of extracting entities from the natural language query using a machine learning model; and A step of determining the first entity node based on the similarity between the extracted entity and the entity node included in the first index database. A method of providing information, including:

3. In paragraph 1, The first index database includes a data node associated with the source data stored in the source database and at least one entity node generated from the source data using a machine learning model, An information providing method in which the above data nodes and the above entity nodes are connected to edges using a logical data model.

4. In paragraph 3, The above data node includes an identifier associated with the source data and text embedding vector information associated with the source data, The above entity node contains entity and graph embedding vector information, A method for providing information, wherein a text embedding vector converted from the above-mentioned source data and a graph embedding vector converted from the above-mentioned data node and the above-mentioned entity node are stored in a second index database.

5. In paragraph 3, The step of extracting data associated with the first entity node comprises: A step of determining a first data node connected to the first entity node; and A step of extracting data associated with the first data node from the source database. A method of providing information, including:

6. In paragraph 4, The step of extracting data associated with the first entity node comprises: A step of determining a second entity node based on the graph embedding vector information of the first entity node; A step of determining a second data node connected to the second entity node; and A step of extracting data associated with the second data node from the source database; A method of providing information, including:

7. In paragraph 4, The step of extracting data associated with the first entity node comprises: A step of determining a first data node connected to the first entity node; A step of determining a second data node based on the text embedding vector information of the first data node; and A step of extracting data associated with the second data node from the source database. A method of providing information, including:

8. In paragraph 3, The above entity node is connected by an edge to at least one of a plurality of data nodes or a plurality of entity nodes stored in the first index database by the above logical data model, The step of extracting data associated with the first entity node comprises: A step of determining a second entity node connected to the first entity node; A step of determining a first data node connected to the second entity node; and A step of extracting data associated with the first data node from the source database. A method of providing information, including:

9. A computer-readable, non-transitory recording medium recording commands for executing the method according to Article 1 on a computer.

10. As a system, Communication module; memory; and At least one processor coupled to said memory and configured to execute at least one computer-readable program contained in said memory, At least one of the above programs, Receive natural language queries, Based on the above natural language query, a first entity node stored in a first index database formed in a graph structure is determined, Extracting data associated with the first entity node from the source database, A system comprising commands for outputting a result for the received natural language query using data associated with the first entity.

Citation Information

Patent Citations

  • Identifying entities using a deep-learning model

    KR1020180099812A

  • Graft copolymer, method for preparing the copoymer and resin composition comprising the copolymer

    KR1020210141332A

  • Method for Alkyation of Norbornene Alcohol Compounds

    KR1020230067142A

  • Method and device for connecting map application to the process for affiliation authentication of user account

    KR1020240163276A

  • Method and system for generating knowledge graphs automatically

    KR102603767B1