Data search method and electronic device

The data retrieval method addresses inefficiencies in distributed data environments by using keyword and context-based node identification with search indexes, enhancing search efficiency and access right management.

WO2026095278A1PCT designated stage Publication Date: 2026-05-07INTELLECTUS CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
INTELLECTUS CORP
Filing Date
2025-08-12
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

In distributed data environments, existing data search technologies face challenges in efficiently retrieving data while managing access rights, as search indexes may not fully address data usage rights and storage placement considerations.

Method used

A data retrieval method that generates keywords and context information, identifies nodes for data retrieval using first- and second-level search indexes with graph data structures, and transmits this information to nodes for efficient data retrieval, while controlling access rights through access permission information.

Benefits of technology

Enables fast and efficient data retrieval by identifying appropriate nodes and managing access rights, improving search efficiency and data usage control in distributed environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025012240_07052026_PF_FP_ABST
    Figure KR2025012240_07052026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed is a data search method performed by at least one processor, the data search method comprising the steps of: generating a keyword and context information on the basis of a search word that is input; identifying, on the basis of the keyword, the context information and a first layered search index, a node to undergo a data search from among a plurality of nodes, each node including a data repository and a second layered search index for searching for data stored in the data repository; transmitting the keyword and the context information to the identified node; and receiving, from the identified node, a data search result including data retrieved on the basis of the keyword, the context information and the second layered search index stored in the identified node.
Need to check novelty before this filing date? Find Prior Art

Description

Data retrieval methods and electronic devices

[0001] The present disclosure relates to a data retrieval method and an electronic device.

[0002] Data is utilized as a key asset in most business sectors, and as data continues to accumulate and diversify in form, the importance of integrated management is emerging. In particular, because there is a possibility that data may be misused for purposes other than those authorized to service providers, it is becoming crucial to clearly identify data owners and control data usage rights.

[0003] Meanwhile, data integration platforms are evolving into forms such as databases, data warehouses, data lakes, and data fabrics in response to these demands. In particular, data fabrics can accommodate various existing requirements in a distributed environment without integrating data into physical storage, while also offering advantages such as protecting the rights of data owners. However, due to the nature of distributed environments, consideration of the storage placement of search indexes and the scope of the data included may be necessary when performing searches on the entire dataset within a data integration platform. For instance, since search indexes contain only a portion of the original data, it may be difficult to completely resolve issues regarding data usage rights in such distributed environments. Accordingly, there is a demand for the development of data search technologies that support fast and efficient data retrieval while controlling access rights in these distributed environments.

[0004] The present disclosure provides a data retrieval method and an electronic device for solving the above-mentioned problems.

[0005] The present disclosure may be implemented in various ways, including a computer-readable non-transient recording medium that records instructions for execution in a method, device (system), and / or computer.

[0006] According to one embodiment of the present disclosure, a data retrieval method performed by at least one processor may include: generating keywords and context information based on input search terms; identifying a node among a plurality of nodes to be a data retrieval target based on keywords, context information and a first-level search index, wherein each of the plurality of nodes includes a data storage and a second-level search index for retrieving data stored in the data storage; transmitting keywords and context information to the identified node; and receiving a data retrieval result from the identified node, the result including data retrieved based on keywords, context information and a second-level search index stored in the identified node.

[0007] According to one embodiment, the data retrieval method further includes the step of receiving access permission information for at least one of data stored in each of the plurality of nodes or a second layer search index stored in each of the plurality of nodes from each of the plurality of nodes, and the step of identifying a node to be a data retrieval target may include the step of identifying a node to be a data retrieval target among the plurality of nodes based on the access permission information received from each of the plurality of nodes.

[0008] According to one embodiment, the data retrieval method further includes the step of transmitting a designated program to each of a plurality of nodes, and access rights information can be generated through a designated program installed on each of the plurality of nodes.

[0009] According to one embodiment, a first layer search index and a second layer search index stored in each of a plurality of nodes have a graph data structure, and the second layer search index stored in each of the plurality of nodes can be connected to the first layer search index through the edges of the graph data included in the first layer search index.

[0010] According to one embodiment, a second layer search index stored in each of a plurality of nodes may include a first embedding vector value for data stored in each of the plurality of nodes.

[0011] According to one embodiment, the first layer search index may include at least one of information for each of a plurality of nodes, metadata for data stored in each of a plurality of nodes, or a first embedding vector value included in a second layer search index stored in each of a plurality of nodes.

[0012] According to one embodiment, the step of identifying a node to be a data search target may include: generating a second embedding vector value based on at least one of keyword or context information; obtaining a first array by applying a filter to the second embedding vector value, wherein each element of the first array includes information regarding the number of times a first hash value is obtained by inputting the second embedding vector value into each of a plurality of hash functions included in the filter; identifying an association between data stored in each of a plurality of nodes and a search term based on a result value obtained by comparing the first array with the second array obtained by applying a filter to the first embedding vector value included in a second layer search index stored in each of a plurality of nodes, wherein each element of the second array includes information regarding the number of times a second hash value is obtained by inputting the first embedding vector value into each of a plurality of hash functions; and identifying a node to be a data search target among a plurality of nodes based on the identified association.

[0013] According to one embodiment, the step of identifying an association between data stored in each of a plurality of nodes and a search term may include: identifying a first element among the elements of a first array such that the number of times a first hash value is acquired is greater than or equal to a specified number; identifying a second element among the elements of a second array such that it corresponds to the first element; and identifying that there is an association between data stored in each of the plurality of nodes and a search term if the value of the second element is greater than or equal to a specified value.

[0014] A computer-readable, non-transient recording medium may be provided that records instructions for executing a data retrieval method according to one embodiment of the present disclosure on a computer.

[0015] According to one embodiment of the present disclosure, an electronic device comprises a memory and at least one processor connected to the memory and configured to execute at least one computer-readable program included in the memory, and the at least one program may include instructions for generating keywords and context information based on input search terms, identifying a node among a plurality of nodes to be a data search target based on keywords, context information and a first-level search index - each of the plurality of nodes includes a data storage and a second-level search index for searching data stored in the data storage - transmitting keywords and context information to the identified node, and receiving a data search result from the identified node including data searched based on keywords, context information and a second-level search index stored in the identified node.

[0016] According to some embodiments of the present disclosure, data can be searched more quickly and efficiently by identifying a data search target and searching for data based on a first-level search index stored in an electronic device and a second-level search index stored in each of a plurality of nodes including a data store.

[0017] In addition, according to some embodiments of the present disclosure, access rights to data can be efficiently controlled by identifying a data search target based on access rights information for a second layer search index.

[0018] The effects of the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by a person skilled in the art to which the present disclosure pertains (referred to as "person skilled in the art") from the description in the claims.

[0019] Embodiments of the present disclosure will be described with reference to the accompanying drawings described below, wherein similar reference numerals indicate similar elements, but are not limited thereto.

[0020] FIG. 1 is a diagram illustrating an exemplary system for data retrieval according to one embodiment of the present disclosure.

[0021] FIG. 2 is a schematic diagram showing a configuration in which an information processing system is connected to communicate with a plurality of user terminals in relation to data processing according to one embodiment of the present disclosure.

[0022] FIG. 3 is a block diagram showing the internal configuration of a user terminal and an information processing system according to one embodiment of the present disclosure.

[0023] FIG. 4 is a drawing for explaining the configuration of an electronic device for data retrieval according to one embodiment of the present disclosure.

[0024] FIG. 5 is a diagram illustrating the configuration of a node for data retrieval according to one embodiment of the present disclosure.

[0025] FIG. 6 is a diagram illustrating a method for determining whether to search a hierarchical search index using a filter according to one embodiment of the present disclosure.

[0026] FIG. 7 is a diagram illustrating a method for generating a hierarchical search index according to one embodiment of the present disclosure.

[0027] FIG. 8 is a drawing for explaining a data retrieval method according to one embodiment of the present disclosure.

[0028] FIG. 9 is a drawing for explaining a method for identifying a data search target according to one embodiment of the present disclosure.

[0029] Hereinafter, specific details for implementing the present disclosure will be described in detail with reference to the attached drawings. However, in the following description, specific descriptions regarding widely known functions or configurations will be omitted if there is a risk that the gist of the present disclosure may be unnecessarily obscured.

[0030] In the attached drawings, identical or corresponding components are assigned the same reference numerals. Additionally, in the description of the following embodiments, the description of identical or corresponding components may be omitted. However, even if a description of a component is omitted, it is not intended that such component is not included in any embodiment.

[0031] The advantages and features of the disclosed embodiments and the methods for achieving them will become clear by referring to the embodiments described below in conjunction with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below but may be implemented in various different forms, and the embodiments provided are merely to make the present disclosure complete and to fully inform those skilled in the art of the scope of the invention.

[0032] The terms used in this specification will be briefly explained, and the disclosed embodiments will be described in detail. The terms used in this specification have been selected to be as generally used as possible, taking into account their functions in this disclosure; however, these terms may vary depending on the intent of those skilled in the art, case law, the emergence of new technologies, etc. Additionally, in specific cases, terms may be arbitrarily selected by the applicant, and in such cases, their meanings will be described in detail in the relevant description of the invention. Therefore, the terms used in this disclosure should be defined not merely by their names, but based on their meanings and the content throughout this disclosure.

[0033] In this specification, singular expressions include plural expressions unless the context clearly specifies them as singular. Additionally, plural expressions include singular expressions unless the context clearly specifies them as plural. Throughout the specification, when a part is described as including a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components.

[0034] Additionally, the terms 'module' or 'part' as used in the specification refer to software or hardware components, and the 'module' or 'part' performs certain roles. However, the meaning of 'module' or 'part' is not limited to software or hardware. The 'module' or 'part' may be configured to reside in an addressable storage medium or configured to run on one or more processors. Thus, as an example, the 'module' or 'part' may include components such as software components, object-oriented software components, class components, and task components, and at least one of processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, or variables. The components and the functions provided within the 'module' or 'part' may be combined into a smaller number of components and 'modules' or 'parts', or further separated into additional components and 'modules' or 'parts'.

[0035] According to one embodiment of the present disclosure, a ‘module’ or ‘part’ may be implemented as a processor and memory. The term ‘processor’ should be broadly interpreted to include a general-purpose processor, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a controller, a microcontroller, a state machine, etc. In some environments, the term ‘processor’ may refer to an application-specific integrated circuit (ASIC), a programmable logic device (PLD), a field programmable gate array (FPGA), etc. The term ‘processor’ may also refer to a combination of processing devices, such as, for example, a combination of a DSP and a microprocessor, a combination of multiple microprocessors, a combination of one or more microprocessors combined with a DSP core, or any other combination of such configurations. Additionally, the term ‘memory’ should be broadly interpreted to include any electronic component capable of storing electronic information. 'Memory' may refer to various types of processor-readable media, such as Random Access Memory (RAM), Read-Only Memory (ROM), Non-Volatile Random Access Memory (NVRAM), Programmable Read-Only Memory (PROM), Erasable-Programmable Read-Only Memory (EPROM), Electrically Erasable PROM (EEPROM), Flash Memory, Magnetic or Marked Data Storage Devices, Registers, etc. If a processor can read information from memory and / or write information to memory, the memory is said to be in an electronic communication state with the processor. Memory integrated into a processor is in an electronic communication state with the processor.

[0036] In addition, terms such as first, second, A, B, (a), (b), etc. used in the following embodiments are used merely to distinguish one component from another, and the essence, order, or sequence of the said component is not limited by such terms.

[0037] Additionally, in the following embodiments, where it is stated that one component is 'connected', 'coupled', or 'joined' to another component, it should be understood that the component may be directly connected or joined to the other component, but that another component may also be 'connected', 'coupled', or 'joined' between each component.

[0038] Additionally, as used in the following embodiments, 'comprises' and / or 'comprising' do not exclude the presence or addition of one or more other components, steps, actions, and / or elements to the mentioned components, steps, actions, and / or elements.

[0039] Hereinafter, various embodiments of the present disclosure will be described in detail with reference to the attached drawings.

[0040] FIG. 1 is a diagram illustrating an exemplary system for data retrieval according to an embodiment of the present disclosure. Referring to FIG. 1, the system for data retrieval may include an electronic device (100) and a plurality of nodes (120). Here, the system may be referred to as a data integration platform, a distributed data platform, etc., the electronic device (100) may perform the role of a core in the system, and each of the plurality of nodes (120) may perform the role of a data storage. At this time, the electronic device (100) may be connected to each of the plurality of nodes (120) through a communication circuit (or communication module). Each of the plurality of nodes (120) may be any network device that includes the ability to process and store data (122a, 124a), such as, for example, a server, a router, a network PC, a peer device, etc. Each of the plurality of nodes (120) may be a peer that is interconnected with one another through a LAN or WAN, such as an intranet or the Internet, for example.

[0041] When examining the data search process, the electronic device (100) can generate keywords and context information based on the input search term (102) when a search term (102) is input. Keywords represent key words or phrases in the search term (102). The electronic device (100) can extract frequently occurring words as keywords by calculating the frequency with which a specific word appears in the search term (102), or select keywords by identifying the presence of the corresponding keyword in the search term (102) using a predefined keyword list. Additionally, context information represents text surrounding the keyword or information related to the overall topic of the search term (102), and can be used to understand the meaning of the keyword more clearly. This context information can be generated through topic modeling or by converting the search term (102) into an embedding vector. Here, the embedding vector may be a method of representing data (e.g., the search term (102)) as a fixed-dimensional real vector or a result value obtained through this method. For example, when there is high semantic similarity between data, the embedding vectors corresponding to the data can be located adjacent to each other in the representation space. That is, semantic similarity regarding context information can be identified through the embedding vectors.

[0042] Then, the electronic device (100) can identify a node among a plurality of nodes (120) that is the target of data search based on keywords, context information, and a first-level search index (110). Here, each of the plurality of nodes (120) may include a data storage and a second-level search index (122b, 124b) for searching data (122a, 124a) stored in the data storage. At this time, the second-level search index (122b, 124b) stored in each of the plurality of nodes (120) may include an embedding vector value for the data (122a, 124a) stored in each of the plurality of nodes (120). For example, a second layer search index (122b) stored in a first node (122) may include an embedding vector value for data (122a) stored in the first node (122), and a second layer search index (124b) stored in an nth node (124) (where n is a natural number greater than or equal to 2) may include an embedding vector value for data (124a) stored in the nth node (124). Additionally, the first layer search index (110) stored in the electronic device (100) may include at least one of information for each of the plurality of nodes (120), metadata for data (122a, 124a) stored in each of the plurality of nodes (120), embedding vector values ​​included in the second layer search index (122b, 124b) stored in each of the plurality of nodes (120), or filter values ​​using embedding vector values ​​included in the second layer search index (122b, 124b) stored in each of the plurality of nodes (120). Here, the information for each of the plurality of nodes (120) may include, for example, identification information capable of identifying each of the plurality of nodes (120).Additionally, metadata for data (122a, 124a) stored in each of the multiple nodes (120) may include, for example, at least one of the title, summary, description, or basic information that the owner of the data (122a, 124a) inputs when linking the data (122a, 124a). As described above, the system for data retrieval may have a structure in which the owner of the data (e.g., each of the multiple nodes (120)) can set an index (e.g., a second-level search index (122b, 124b)). Furthermore, the system for data retrieval may improve the overall performance of the search by increasing the efficiency of the search process by supporting a rapid determination of whether to search for the second-level search index (122b, 124b) stored in each of the multiple nodes (120) using the first-level search index (110).

[0043] According to one embodiment, the first layer search index (110) and the second layer search index (122b, 124b) stored in each of the plurality of nodes (120) may have a graph data structure. Here, the graph data structure may represent a data structure that structures information and mutual relationships between information by expressing entities and relations between entities as nodes and edges. At this time, the second layer search index (122b, 124b) stored in each of the plurality of nodes (120) may be connected to the first layer search index (110) through the edges of the graph data included in the first layer search index (110). For example, the first layer search index (110) stored in the electronic device (100) and the second layer search index (122b, 124b) stored in each of the plurality of nodes (120) may be logically connected in the form of a graph data structure.

[0044] According to one embodiment, an electronic device (100) may receive access rights information for at least one of data (122a, 124a) stored in each of the plurality of nodes (120) or a second layer search index (122b, 124b) stored in each of the plurality of nodes (120) from each of the plurality of nodes (120). At this time, the electronic device (100) may identify a node among the plurality of nodes (120) that is the target of data search based on the access rights information received from each of the plurality of nodes (120).

[0045] Then, the electronic device (100) can transmit keyword and context information to a node identified as a data search target. For example, the electronic device (100) can transmit keyword and context information generated based on an input search term (102) to a node identified as a data search target, and request the node to search for data (122a, 124a) stored in the node's data storage. At this time, the node identified as a data search target can respond to the request and search for data (122a, 124a) based on the received keyword, the received context information, and the second layer search index (122b, 124b) stored in the node.

[0046] Then, the electronic device (100) can receive a data search result (104) including data (122a, 124a) searched based on keywords, context information, and a second layer search index (122b, 124b) stored in the node identified as a data search target. Accordingly, the electronic device (100) can provide the search result (104) for the input search term (102) to the user.

[0047] FIG. 2 is a schematic diagram showing a configuration in which an information processing system (230) is connected to communicate with a plurality of user terminals (210_1, 210_2, 210_3) in relation to data processing according to one embodiment of the present disclosure. The information processing system (230) may include system(s) (e.g., the system of FIG. 1) capable of providing a data processing service (e.g., a data search-based service). In one embodiment, the information processing system (230) may include one or more server devices and / or databases capable of storing, providing, and executing computer-executable programs (e.g., downloadable applications) and data related to the data processing service, or one or more distributed computing devices and / or distributed databases based on cloud computing services. For example, the information processing system (230) may include separate systems (e.g., servers) for the data processing service.

[0048] Data processing services, etc. provided by the information processing system (230) can be provided to the user through a data processing application, a web browser application, etc. installed on each of the multiple user terminals (210_1, 210_2, 210_3).

[0049] Multiple user terminals (210_1, 210_2, 210_3) can communicate with an information processing system (230) through a network (220). The network (220) can be configured to enable communication between the multiple user terminals (210_1, 210_2, 210_3) and the information processing system (230). Depending on the installation environment, the network (220) may be configured as a wired network such as Ethernet, Power Line Communication, telephone line communication device and RS-serial communication, a mobile communication network, a Wireless LAN (WLAN), Wi-Fi, Bluetooth and ZigBee, or a combination thereof. The communication method is not limited and may include not only communication methods utilizing communication networks that the network (220) may include (e.g., mobile communication network, wired internet, wireless internet, broadcasting network, satellite network, etc.) but also short-range wireless communication between user terminals (210_1, 210_2, 210_3).

[0050] For example, multiple user terminals (210_1, 210_2, 210_3) can transmit a data processing request and a command associated with a user request for data processing to an information processing system (230) through a network (220), and the information processing system (230) can receive this.

[0051] In FIG. 2, a mobile phone terminal (210_1), a tablet terminal (210_2), and a PC terminal (210_3) are shown as examples of user terminals, but are not limited thereto, and the user terminals (210_1, 210_2, 210_3) may be any computing device capable of wired and / or wireless communication and capable of installing and running data processing applications, etc. For example, user terminals may include smartphones, mobile phones, navigation systems, computers, laptops, digital broadcasting terminals, PDAs (Personal Digital Assistants), PMPs (Portable Multimedia Players), tablet PCs, game consoles, wearable devices, IoT (Internet of Things) devices, VR (Virtual Reality) devices, AR (Augmented Reality) devices, etc. Additionally, FIG. 2 illustrates three user terminals (210_1, 210_2, 210_3) communicating with an information processing system (230) through a network (220), but is not limited thereto, and may be configured so that a different number of user terminals communicate with an information processing system (230) through a network (220).

[0052] FIG. 3 is a block diagram showing the internal configuration of a user terminal (210) and an information processing system (230) according to one embodiment of the present disclosure. The user terminal (210) may refer to any computing device capable of executing data processing applications, etc., and capable of wired / wireless communication, and may include, for example, the mobile phone terminal (210_1), tablet terminal (210_2), PC terminal (210_3) of FIG. 2. As illustrated, the user terminal (210) may include a memory (312), a processor (314), a communication module (316), and an input / output interface (318). Similarly, the information processing system (230) may include a memory (332), a processor (334), a communication module (336), and an input / output interface (338). As illustrated in FIG. 3, the user terminal (210) and the information processing system (230) may be configured to communicate information and / or data through the network (220) using their respective communication modules (316, 336). Additionally, the input / output device (320) may be configured to input information and / or data to the user terminal (210) or output information and / or data generated from the user terminal (210) through the input / output interface (318).

[0053] The memory (312, 332) may include any non-transient computer-readable recording medium. According to one embodiment, the memory (312, 332) may include a non-perishable permanent mass storage device such as ROM (read-only memory), a disk drive, an SSD (solid-state drive), or a flash memory. As another example, a non-perishable permanent mass storage device such as ROM, an SSD, a flash memory, or a disk drive may be included in the user terminal (210) or the information processing system (230) as a separate permanent storage device distinct from the memory. Additionally, the memory (312, 332) may store an operating system and at least one program code (e.g., code for an application associated with a data processing service).

[0054] These software components may be loaded from a computer-readable recording medium separate from memory (312, 332). This separate computer-readable recording medium may include a recording medium that can be directly connected to the user terminal (210) and the information processing system (230), for example, a computer-readable recording medium such as a floppy drive, disk, tape, DVD / CD-ROM drive, or memory card. As another example, the software components may be loaded into memory (312, 332) via a communication module (316, 336) rather than a computer-readable recording medium. For example, at least one program may be loaded into memory (312, 332) based on a computer program (e.g., an application associated with a data processing service, etc.) that is installed by files provided through a network (220) by developers or a file distribution system that distributes installation files for the application.

[0055] The processor (314, 334) may be configured to process instructions of a computer program by performing basic arithmetic, logic, and input / output operations. Instructions may be provided to the processor (314, 334) by memory (312, 332) or a communication module (316, 336). For example, the processor (314, 334) may be configured to execute instructions received according to program code stored in a recording device such as memory (312, 332).

[0056] The communication module (316, 336) may provide a configuration or function for the user terminal (210) and the information processing system (230) to communicate with each other via the network (220), and may provide a configuration or function for the user terminal (210) and / or the information processing system (230) to communicate with another user terminal or another system (e.g., a separate cloud system). For example, a request or data (e.g., a data processing request or data, etc.) generated by the processor (314) of the user terminal (210) according to program code stored in a recording device such as memory (312) may be transmitted to the information processing system (230) via the network (220) under the control of the communication module (316). Conversely, a control signal or command provided under the control of the processor (334) of the information processing system (230) can be received by the user terminal (210) through the communication module (336) and the network (220) via the communication module (316) of the user terminal (210).

[0057] The input / output interface (318) may be a means for interfacing with an input / output device (320). As an example, the input device may include a device such as a camera including an audio sensor and / or an image sensor, a keyboard, a microphone, or a mouse, and the output device may include a device such as a display, a speaker, or a haptic feedback device. As another example, the input / output interface (318) may be a means for interfacing with a device in which the configuration or function for performing input and output is integrated into one, such as a touchscreen. Although the input / output device (320) is depicted in FIG. 3 as not being included in the user terminal (210), it is not limited thereto and may be configured as a single device with the user terminal (210). Additionally, the input / output interface (338) of the information processing system (230) may be a means for interfacing with a device (not shown) for input or output that is connected to the information processing system (230) or that the information processing system (230) may include. In FIG. 3, the input / output interface (318, 338) is shown as an element configured separately from the processor (314, 334), but is not limited thereto, and the input / output interface (318, 338) may be configured to be included in the processor (314, 334).

[0058] The user terminal (210) and the information processing system (230) may include more components than those of FIG. 3. However, it is not necessary to clearly illustrate most of the conventional technical components. In one embodiment, the user terminal (210) may be implemented to include at least some of the input / output devices (320) described above. Additionally, the user terminal (210) may further include other components such as a transceiver, a GPS (Global Positioning System) module, a camera, various sensors, a database, etc. For example, if the user terminal (210) is a smartphone, it may include components that are generally included in a smartphone, and may be implemented to include various components such as an accelerometer, a gyroscope, a microphone module, a camera module, various physical buttons, buttons using a touch panel, input / output ports, and a vibrator for vibration.

[0059] According to one embodiment, the processor (314) of the user terminal (210) may be configured to operate a data processing application or a web browser application that provides data processing services. At this time, program code associated with the application may be loaded into the memory (312) of the user terminal (210). While the application is operating, the processor (314) of the user terminal (210) may receive information and / or data provided from an input / output device (320) through an input / output interface (318) or receive information and / or data from an information processing system (230) through a communication module (316), and may process the received information and / or data and store it in the memory (312). Additionally, such information and / or data may be provided to the information processing system (230) through the communication module (316).

[0060] While the data processing application is in operation, the processor (314) may receive voice data, text, images, videos, etc., that are input or selected through an input device such as a touch screen, keyboard, audio sensor and / or image sensor, camera, microphone, etc., connected to the input / output interface (318), and may store the received voice data, text, images and / or videos, etc. in memory (312) or provide them to an information processing system (230) through a communication module (316) and a network (220). In one embodiment, the processor (314) may receive user input input through an input device and provide data / requests corresponding to the received user input to an information processing system (230) through a network (220) and a communication module (316).

[0061] The processor (314) of the user terminal (210) can transmit information and / or data to an input / output device (320) through an input / output interface (318) and output it. For example, the processor (314) of the user terminal (210) can output the processed information and / or data through an output device (320), such as a display output device (e.g., touch screen, display, etc.) or a voice output device (e.g., speaker).

[0062] The processor (334) of the information processing system (230) may be configured to manage, process, and / or store information and / or data received from a plurality of user terminals (210) and / or a plurality of external systems. The information and / or data processed by the processor (334) may be provided to the user terminals (210) through a communication module (336) and a network (220).

[0063] FIG. 4 is a diagram illustrating the configuration of an electronic device for data retrieval according to one embodiment of the present disclosure. Referring to FIG. 4, the electronic device (400) for data retrieval (e.g., the electronic device (100) of FIG. 1) may include a processor (410) and a memory (420). However, the configuration of the electronic device (400) is not limited thereto. According to various embodiments, the electronic device (400) may further include at least one other component in addition to the components described above. For example, the electronic device (400) may further include a communication circuit (or communication module) for receiving various data from an external electronic device (e.g., the node (120) of FIG. 1).

[0064] The processor (410) is connected to memory (420) and may be configured to execute at least one computer-readable program (e.g., a management program (422)) contained in memory (420). For example, the processor (410) may execute software (or a program) to control at least one other component (e.g., a hardware or software component) of an electronic device (400) connected to the processor (410) and may perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (410) may load instructions or data received from other components (e.g., a communication circuit) into volatile memory, process the instructions or data stored in volatile memory, and store the resulting data in non-volatile memory.

[0065] The memory (420) may store various data used by at least one component of the electronic device (400) (e.g., processor (410)). The data may include, for example, input data or output data for software (or programs) and related instructions. The memory (420) may include volatile memory or non-volatile memory. According to one embodiment, the memory (420) may store a management program (422) and a first layer search index (424) (e.g., the first layer search index (110) of FIG. 1).

[0066] At least one program executed by the processor (410) may include instructions related to data retrieval. In the following description, the processor (410) is described as performing a certain function, but this is for convenience of explanation, and the function performed by the processor (410) can be understood as the processor (410) executing instructions included in at least one program stored in memory (420).

[0067] The processor (410) generates keywords and context information based on an input search term, and can identify a node that is the target of data search among a plurality of nodes (e.g., node (120) of FIG. 1) included in a system for data search (e.g., system of FIG. 1). At this time, each of the plurality of nodes may include a data storage and a second layer search index (e.g., second layer search index (122b, 124b) of FIG. 1) for searching data stored in the data storage (e.g., data (122a, 124a) of FIG. 1). Additionally, the second layer search index may include an embedding vector value for the data stored in the corresponding node.

[0068] The first layer search index (424) may include at least one of information for each of a plurality of nodes, metadata for data stored in each of a plurality of nodes, an embedding vector value included in the second layer search index stored in each of a plurality of nodes, or a filter value using the embedding vector value included in the second layer search index stored in each of a plurality of nodes. In this way, the first layer search index (424) can support determining whether to search the second layer search index stored in each of a plurality of nodes during the process of identifying the data search target for the input search term. For example, the processor (410) can identify the node among the plurality of nodes that is the data search target by comparing the embedding vector value stored in the first layer search index (424) (hereinafter referred to as the first embedding vector value) and the embedding vector value generated from the input search term (hereinafter referred to as the second embedding vector value). Alternatively, the processor (410) can identify a node among a plurality of nodes that is the target of data search by using an embedding vector value generated from a filter value stored in the first layer search index (424) and an input search term. A method for identifying a node that is the target of data search using an embedding vector will be explained with reference to FIG. 6.

[0069] Information regarding each of the plurality of nodes included in the first layer search index (424) may include identification information capable of identifying each of the plurality of nodes. Additionally, metadata regarding data stored in each of the plurality of nodes included in the first layer search index (424) may include at least one of the data title, summary, description, or basic information entered by the data owner when linking the data.

[0070] According to one embodiment, the processor (410) may receive access permission information regarding at least one of the data stored in each of the multiple nodes or the second layer search index stored in each of the multiple nodes from each of the multiple nodes. At this time, the processor (410) may identify the node among the multiple nodes that is the target of data search based on the access permission information received from each of the multiple nodes. For example, even if data suitable for an input search term is stored in any of the multiple nodes, the processor (410) may exclude the node from the data search target if the user does not have access permission to the data stored in that node or at least one of the second layer search index. According to one embodiment, the access permission information regarding the data stored in each of the multiple nodes or at least one of the second layer search index stored in each of the multiple nodes may be stored in the first layer search index (424). For example, the first layer search index (424) may include information for each of a plurality of nodes, metadata for data stored in each of a plurality of nodes, an embedding vector value included in a second layer search index stored in each of a plurality of nodes, or a filter value using an embedding vector value included in a second layer search index stored in each of a plurality of nodes, along with access right information for at least one of the data stored in each of a plurality of nodes or the second layer search index stored in each of a plurality of nodes. In this way, the system for data retrieval can support the data owner (e.g., each of a plurality of nodes) in setting access rights for at least one of the data or the index (e.g., the second layer search index), and by identifying the node to be searched based on the access right information set directly by the data owner, it can support fast and efficient data retrieval while controlling access rights to the data.

[0071] The processor (410) can transmit keyword and context information to a node identified as a data search target. For example, the processor (410) can transmit keyword and context information generated based on an input search term to a node identified as a data search target, thereby requesting the node to search for data stored in the node's data storage.

[0072] The processor (410) can receive data search results including data searched based on keywords, context information, and a second layer search index stored in the node identified as a data search target. Then, the processor (410) can provide the received data search results to the user.

[0073] According to one embodiment, if there are multiple nodes among a plurality of nodes identified as data search targets, the processor (410) may set a scheduling for performing data search operations for the multiple nodes identified as data search targets in parallel. Additionally, the processor (410) may process the results of the data search operations performed on the multiple nodes and provide them to the user.

[0074] According to one embodiment, the function of the processor (410) associated with the above-described data search can be configured by a management program (422). Here, the management program (422) may be referred to as a search index management program, etc. The management program (422) may receive a search term from a user and provide an interface capable of providing search results for the entered search term to the user. Additionally, the management program (422) may be connected via a network and communicate with a program (e.g., the agent program (510) of FIG. 5) capable of managing data stored in each of a plurality of nodes and a second-layer search index.

[0075] FIG. 5 is a diagram illustrating the configuration of a node for data retrieval according to an embodiment of the present disclosure. The node (500) described in FIG. 5 may represent any one of a plurality of nodes (e.g., node (120) of FIG. 1) included in a system for data retrieval (e.g., system of FIG. 1).

[0076] Referring to FIG. 5, a node (500) for data retrieval may include a data storage (e.g., memory) for storing data (520) (e.g., data (122a, 124a) of FIG. 1) and a second layer search index (530) (e.g., second layer search index (122b, 124b) of FIG. 1) for retrieving data (520) stored in the data storage. Although not illustrated, the node (500) may further include a control circuit (or control module) (e.g., processor) as an electronic device that can be connected via a network with the electronic device (400) of FIG. 4. In the following description, the function performed by the node (500) in connection with data retrieval is described as the control circuit performing the corresponding function, but this is for convenience of explanation and the function performed by the control circuit can be understood as the control circuit executing instructions included in at least one program (e.g., agent program (510)) stored in the node (500).

[0077] The control circuit may generate a second-level search index (530) by scanning metadata for the data (520) and at least a portion of the data (520). According to one embodiment, the control circuit may generate an embedding vector value by embedding the data (520). Then, the control circuit may generate a second-level search index (530) using the embedding vector value for the data (520). For example, the second-level search index (530) may include the embedding vector value for the data (520). Additionally, the control circuit may update the second-level search index (530) when the data (520) changes.

[0078] The control circuit may transmit at least one of information about the node (500), metadata about the data (520) stored in the node (500), or an embedding vector value included in the second layer search index (530) stored in the node (500) to the core of the system for data retrieval (e.g., the electronic device (100) of FIG. 1 or the electronic device (400) of FIG. 4). Here, the information about the node (500) may include identification information capable of identifying the node (500). Additionally, the metadata about the data (520) stored in the node (500) may include at least one of the title, summary, description of the data (520), or basic information entered by the owner of the data (520) when linking the data (520). At this time, the core for data retrieval can generate (or set) a first-level search index (e.g., the first-level search index (110) of FIG. 1 or the first-level search index (424) of FIG. 4) based on at least one of information about the node (500) received from the node (500), metadata about the data (520) stored in the node (500), or an embedding vector value included in the second-level search index (530) stored in the node (500). For example, the first-level search index may include at least one of information about the node (500), metadata about the data (520) stored in the node (500), an embedding vector value included in the second-level search index (530) stored in the node (500), or a filter value using the embedding vector value included in the second-level search index (530) stored in the node (500). Here, a method for generating or updating a filter value using an embedding vector value included in a second layer search index (530) stored in a node (500) will be explained with reference to FIG. 6.

[0079] According to one embodiment, the control circuit may transmit access permission information for at least one of the data (520) stored in the node (500) or the second layer search index (530) stored in the node (500) to the core of the system for data search. At this time, the core of the system for data search may identify the node among a plurality of nodes that is the target of data search based on the access permission information received from the node (500). For example, even if data suitable for the input search term is stored in the node (500), the core of the system for data search may exclude the node (500) from the data search target if the user does not have access permission for at least one of the data (520) or the second layer search index (530) stored in the node (500). According to one embodiment, the access permission information for at least one of the data (520) stored in the node (500) or the second layer search index (530) stored in the node (500) may be stored in the first layer search index. In this way, the system for data retrieval can support the owner of the data (520) (e.g., node (500)) in setting access rights to at least one of the data (520) or the second layer search index (530), and by identifying the node to be searched based on the access rights information set directly by the owner of the data (520), it can support the rapid and efficient retrieval of the data (520) while controlling access rights to the data (520).

[0080] When a search request for data (520) is received, the control circuit can search for data (520) based on a second layer search index (530). According to one embodiment, when the control circuit receives keywords and context information generated based on a search term input from the core of the system for data search, it can search for data (520) based on the received keywords, the received context information, and the second layer search index (530) stored in the node (500). Then, the control circuit can transmit the data search results to the core of the system for data search.

[0081] According to one embodiment, the function of the control circuit associated with the above-described data search can be set by an agent program (510). Here, the agent program (510) may be referred to as a search index agent program, etc. The agent program (510) may provide an interface that can set a second-level search index (530) for searching data (520). Additionally, the agent program (510) may provide an interface that can set access rights information for at least one of the data (520) or the second-level search index (530). For example, access rights information for at least one of the data (520) or the second-level search index (530) may be generated through the agent program (510) installed on the node (500).

[0082] According to one embodiment, the core of the system for data retrieval can distribute an agent program (510) to a network that can access a node (500). For example, the core of the system for data retrieval can transmit the agent program (510) to the node (500). At this time, the control circuit of the node (500) can perform a data retrieval function by installing the received agent program (510) on the node (500) and executing commands included in the agent program (510). For example, the control circuit can communicate via a network with a management program (e.g., the management program (422) of FIG. 4) installed in the core of the system for data retrieval through an agent program (510), and can report (or transmit) to the core of the system for data retrieval access rights information for at least one of the data (520) stored in the node (500) or the second-layer search index (530) stored in the node (500), along with at least one of information about the node (500), metadata about the data (520) stored in the node (500), or an embedding vector value included in the second-layer search index (530) stored in the node (500).

[0083] FIG. 6 is a diagram illustrating a method for determining whether a hierarchical search index is searched using a filter according to an embodiment of the present disclosure. Referring to FIG. 6, a processor (e.g., processor (410) of FIG. 4) of an electronic device for data search (e.g., electronic device (100) of FIG. 1 or electronic device (400) of FIG. 4) can identify a node that is the target of data search among a plurality of nodes (e.g., node (120) of FIG. 1) included in a system for data search (e.g., system of FIG. 1) by determining whether a hierarchical search index is searched through a comparison of embedding vectors (602) using a filter (600). For example, the electronic device can identify a node that is the target of data search among a plurality of nodes by comparing a first embedding vector value stored in a first hierarchical search index (e.g., first hierarchical search index (110) of FIG. 1 or first hierarchical search index (424) of FIG. 4) with a second embedding vector value generated from an input search term. Here, the first embedding vector value corresponds to an embedding vector value included in a second layer search index (e.g., the second layer search index (122b, 124b) of FIG. 1 or the second layer search index (530) of FIG. 5) stored in each of the multiple nodes, and may be an embedding vector value for data stored in the corresponding node (e.g., data (122a, 124a) of FIG. 1 or data (520) of FIG. 5). Alternatively, the processor may identify a node among the multiple nodes that is the target of data search by using a filter value included in the first layer search index and a second embedding vector value generated from an input search term. Here, the filter value included in the first layer search index may be generated or updated when scanning data stored in each of the multiple nodes and generating a second layer search index therefor, or it may be generated when generating the first layer search index.

[0084] According to one embodiment, the filter (600) may include a bloom filter. Here, the bloom filter may represent an algorithm that uses a hash function (610, 620) to represent data (e.g., an embedding vector (602)) in bits and can quickly check the existence of data through a bit operation. For example, a first-layer search index may include bloom filter information for a second-layer search index of each of a plurality of nodes, and may support a quick determination of whether to actually search the second-layer search index of each of the plurality of nodes for the search result search.

[0085] When looking at the process of generating a filter value for an embedding vector (602) using a filter (600), the processor can obtain an array (630) by applying the filter (600) to the embedding vector value (602). At this time, each element of the obtained array (630) (e.g., elements (632, 634)) may contain information regarding the number of times the hash value (612, 622) obtained by inputting the embedding vector value (602) to each of the multiple hash functions (610, 620) included in the filter (600) is obtained. For example, the array (630) may be a counter array that counts the number of times each of the multiple hash values ​​(612, 622) calculated from each of the multiple hash functions (610, 620) included in the filter (600) is obtained by inputting the embedding vector value (602). Here, the obtained array (630) can be a filter value for the embedding vector (602). Additionally, the filter value obtained through the process described above can be included in the first layer search index. As described above, the processor does not store the result in a bit array of a specific length using a hash function (610, 620) for the target data, but rather uses a counter array to represent the entire data stored in the node and can use this to identify the association between the data and the search term.

[0086] The processor can generate an embedding vector (e.g., a second embedding vector) based on at least one of the keywords or context information generated based on the input search term. Then, the processor can identify the association between the data stored in each of the multiple nodes and the search term by using the filter value (e.g., array (630)) included in the first layer search index and the embedding vector value generated based on the input search term. According to one embodiment, the processor can identify the association between the data stored in each of the multiple nodes and the search term by performing a binary operation using the filter value included in the first layer search index and the embedding vector value generated based on the input search term. For example, the processor may determine that there is no association between the data stored in the node and the search term when the element corresponding to the embedding vector value generated based on the input search term among the elements of the filter value included in the first layer search index, i.e., array (630), is a value of '0'. As another example, the processor may determine that there is an association between the data stored in the node and the search term when the element corresponding to the embedding vector value generated based on the input search term among the elements of the array (630) containing the filter value in the first layer search index is not a '0' value.

[0087] Then, the processor can identify the node among the multiple nodes that is the target of data retrieval based on the identified association between the data stored in each of the multiple nodes and the search term. For example, the processor can identify the node among the multiple nodes that is determined to have an association between the stored data and the search term as the node that is the target of data retrieval.

[0088] According to one embodiment, when a processor identifies the value of an element corresponding to an embedding vector value generated based on an input search term among the elements of an array (630), that is, a filter value included in a first-level search index, the processor can adjust the accuracy and performance of the filter (600) by setting a threshold for the value of the element. For example, if the value of an element corresponding to an embedding vector value generated based on an input search term among the elements of the array (630) is greater than or equal to a specified value (or threshold), the processor can identify that there is an association between the data stored in each of the multiple nodes and the search term, and if it is less than the specified value, the processor can identify that there is no association between the data stored in each of the multiple nodes and the search term.

[0089] FIG. 7 is a diagram illustrating a method for generating a hierarchical search index according to an embodiment of the present disclosure. Referring to FIG. 7, a control circuit of a node (e.g., node (120) of FIG. 1 or node (500) of FIG. 5) included in a system for data retrieval (e.g., system of FIG. 1) can, in step 710 (S710), generate an embedding vector value for data stored in the node (e.g., data (122a, 124a) of FIG. 1 or data (520) of FIG. 5). For example, the control circuit can generate an embedding vector value by embedding the data.

[0090] In step 720 (S720), the control circuit may generate a second-layer search index (e.g., the second-layer search index (122b, 124b) of FIG. 1 or the second-layer search index (530) of FIG. 5) based on the generated embedding vector values. For example, the control circuit may include embedding vector values ​​for data in the second-layer search index. According to one embodiment, the control circuit may update the second-layer search index when data stored in the node changes.

[0091] In step 730 (S730), the control circuit may transmit at least one of information about the node, metadata about the data stored in the node, or an embedding vector value included in the second layer search index to the core of the system for data retrieval (e.g., the electronic device (100) of FIG. 1 or the electronic device (400) of FIG. 4). Here, the information about the node may include identification information capable of identifying the node. Additionally, the metadata about the data stored in the node may include at least one of the title, summary, description, or basic information entered by the owner of the data when linking the data.

[0092] At this time, the processor of the core for data retrieval (e.g., the processor (410) of FIG. 4) may generate (or set) a first-layer search index (e.g., the first-layer search index (110) of FIG. 1 or the first-layer search index (424) of FIG. 4) based on at least one of information about the node received from the node, metadata about the data stored in the node, or an embedding vector value included in the second-layer search index stored in the node. For example, the processor may include at least one of information about the node, metadata about the data stored in the node, an embedding vector value included in the second-layer search index stored in the node, or a filter value using the embedding vector value included in the second-layer search index stored in the node in the first-layer search index. Here, the filter value may be generated or updated by the processor of the core for data retrieval using the embedding vector value included in the second-layer search index stored in the node.

[0093] According to one embodiment, a control circuit may transmit access right information for at least one of data stored in a node or a second-layer search index stored in a node to the core of a system for data retrieval. At this time, the processor of the core of the system for data retrieval may identify a node among a plurality of nodes that is the target of data retrieval based on the access right information received from the node. For example, even if data suitable for an input search term is stored in a node, the processor of the core of the system for data retrieval may exclude the node from the data retrieval target if the user does not have access right to at least one of the data stored in the node or the second-layer search index. According to one embodiment, the access right information for at least one of the data stored in a node or the second-layer search index stored in a node may be stored in a first-layer search index.

[0094] FIG. 8 is a diagram illustrating a data retrieval method according to an embodiment of the present disclosure. Referring to FIG. 8, a processor (e.g., processor (410) of FIG. 4) of an electronic device (or core) (e.g., electronic device (100) of FIG. 1 or electronic device (400) of FIG. 4) included in a system for data retrieval (e.g., system of FIG. 1) can generate keyword and context information based on an input search term in step 810 (S810).

[0095] In step 820 (S820), the processor can identify a node among a plurality of nodes (e.g., node (120) of FIG. 1 or node (500) of FIG. 5) that is the target of data search based on keywords, context information, and a first layer search index (e.g., first layer search index (110) of FIG. 1 or first layer search index (424) of FIG. 4). At this time, each of the plurality of nodes may include a data store and a second layer search index (e.g., second layer search index (122b, 124b) of FIG. 1 or second layer search index (530) of FIG. 5) for searching for data stored in the data store (e.g., data (122a, 124a) of FIG. 1 or data (520) of FIG. 5). Here, the second layer search index may include an embedding vector value for data stored in the corresponding node, and the first layer search index may include at least one of information for each of the plurality of nodes, metadata for data stored in each of the plurality of nodes, an embedding vector value included in the second layer search index stored in each of the plurality of nodes, or a filter value using the embedding vector value included in the second layer search index stored in each of the plurality of nodes. Accordingly, the processor can identify the node among the plurality of nodes that is the target of data search by comparing the embedding vector value stored in the first layer search index with the embedding vector value generated from the input search term. Alternatively, the processor can identify the node among the plurality of nodes that is the target of data search by using the filter value stored in the first layer search index and the embedding vector value generated from the input search term.

[0096] According to one embodiment, a processor may receive access permission information from each of a plurality of nodes regarding at least one of data stored in each of the plurality of nodes or a second-layer search index stored in each of the plurality of nodes. At this time, the processor may identify a node among the plurality of nodes that is a target for data search based on the access permission information received from each of the plurality of nodes. For example, even if data suitable for an input search term is stored in any of the plurality of nodes, the processor may exclude a node from the data search target if the user does not have access permission to at least one of the data stored in that node or a second-layer search index. According to one embodiment, access permission information regarding at least one of the data stored in each of the plurality of nodes or a second-layer search index stored in each of the plurality of nodes may be stored in a first-layer search index.

[0097] In step 830 (S830), the processor may transmit keyword and context information to a node identified as a data search target among a plurality of nodes. For example, the processor may transmit keyword and context information generated based on an input search term to a node identified as a data search target, and request the node to search for data stored in the node's data storage.

[0098] In step 840 (S840), the processor may receive data search results including data searched from an identified node based on keywords, context information, and a second-layer search index. Then, the processor may provide the received data search results to a user.

[0099] According to one embodiment, if there are multiple nodes among a plurality of nodes identified as data search targets, the processor may set a scheduling for performing data search operations for the multiple nodes identified as data search targets in parallel. Additionally, the processor may process the results of the data search operations performed on the multiple nodes and provide them to a user.

[0100] FIG. 9 is a diagram illustrating a method for identifying a data search target according to an embodiment of the present disclosure. Referring to FIG. 9, a processor (e.g., processor (410) of FIG. 4) of an electronic device (or core) (e.g., electronic device (100) of FIG. 1 or electronic device (400) of FIG. 4) included in a system for data search (e.g., system of FIG. 1) may, in step 910 (S910), generate an embedding vector (e.g., embedding vector (602) of FIG. 6) based on an input search term. For example, the processor may generate an embedding vector value (hereinafter referred to as a second embedding vector value) based on at least one of the keywords or context information generated based on the input search term.

[0101] In step 920 (S920), the processor may obtain an array by applying a filter to the embedding vector. For example, the processor may obtain an array (e.g., the array (630) of FIG. 6) (hereinafter referred to as the first array) by applying a filter (e.g., the filter (600) of FIG. 6) to the second embedding vector value. At this time, each element of the first array obtained (e.g., elements (632, 634) of FIG. 6) may include information regarding the number of times the second embedding vector value is input to each of the multiple hash functions (e.g., hash functions (610, 620) of FIG. 6) included in the filter and the hash value (e.g., hash values ​​(612, 622) of FIG. 6) is obtained. For example, the first array may be a counter array that counts the number of times each of the multiple hash values ​​calculated from each of the multiple hash functions included in the filter is obtained by inputting the second embedding vector value.

[0102] In step 930 (S930), the processor can identify the association between data and a search term by comparing it with an array (e.g., array (630) of FIG. 6) (hereinafter referred to as the second array) obtained by applying a filter (e.g., filter (600) of FIG. 6) to an embedding vector (hereinafter referred to as the first embedding vector) included in a second layer search index (e.g., node (120) of FIG. 1 or node (500) of FIG. 5) stored in each of a plurality of nodes (e.g., node (120) of FIG. 1 or node (500) of FIG. 5). For example, the processor may obtain a second array by applying a filter to a first embedding vector value included in a second layer search index stored in each of a plurality of nodes (i.e., stored in a first layer search index (e.g., the first layer search index (110) of FIG. 1 or the first layer search index (424) of FIG. 4). At this time, each element of the second array obtained (e.g., elements (632, 634) of FIG. 6) may include information regarding the number of times a hash value (e.g., hash value (612, 622) of FIG. 6) obtained by inputting the first embedding vector value into each of a plurality of hash functions (e.g., hash function (610, 620) of FIG. 6) included in the filter is obtained. For example, the second array may be a counter array that counts the number of times each of a plurality of hash values ​​calculated from each of a plurality of hash functions included in the filter is obtained by inputting the first embedding vector value. Then, the processor can identify the association between the data stored in each of the multiple nodes and the search term based on the result of comparing the first array and the second array. For example, the processor can compare each element of the first array with the element of the second array corresponding to that element, and if there is no matching part, determine that there is no association, and if there is a matching part, determine that there is an association.For example, the processor may determine that there is no association between the data stored in the node and the search term when either the n-th element of the first array or the n-th element of the second array has a value of '0'. As another example, the processor may determine that there is an association between the data stored in the node and the search term when neither the n-th element of the first array nor the n-th element of the second array has a value of '0'.

[0103] According to one embodiment, when a processor identifies a matching part by comparing each element of a first array with an element of a second array corresponding to that element, it can adjust the accuracy and performance of the filter by setting a threshold for the matching part. For example, the processor can identify a first element among the elements of the first array in which the number of times a first hash value is obtained is greater than or equal to a specified number (e.g., 1 time), and can identify a second element among the elements of the second array that corresponds to the first element. Then, if the value of the second element is greater than or equal to a specified value (or threshold), the processor can identify that there is an association between the data stored in each of the multiple nodes and the search term, and if the value of the second element is less than the specified value, the processor can identify that there is no association between the data stored in each of the multiple nodes and the search term.

[0104] In step 940 (S940), the processor can identify a node among the multiple nodes that is the target of data retrieval based on the identified association between the data stored in each of the multiple nodes and the search term. For example, the processor can identify a node among the multiple nodes that is determined to have an association between the stored data and the search term as the node that is the target of data retrieval.

[0105] According to one embodiment, the processor may omit step 920 and, in step 930, identify the association between the data stored in each of the multiple nodes and the search term by using the filter value included in the first layer search index and the second embedding vector value generated from the input search term. At this time, the filter value included in the first layer search index may be generated or updated when scanning the data stored in each of the multiple nodes and generating the second layer search index therefor, or it may be generated when generating the first layer search index. For example, as described with reference to FIG. 6, the processor may obtain an array by applying a filter to the first embedding vector value. Here, the obtained array may become the filter value for the first embedding vector, and the processor may include the obtained filter value in the first layer search index. Then, the processor may identify the association between the data stored in each of the multiple nodes and the search term by performing a binary operation using the filter value included in the first layer search index and the second embedding vector value generated based on the input search term. For example, the processor may determine that there is no association between the data stored in the node and the search term when the filter value included in the first-level search index—that is, the element among the array elements corresponding to the embedding vector value generated based on the input search term—is a value of '0'. As another example, the processor may determine that there is an association between the data stored in the node and the search term when the filter value included in the first-level search index—that is, the element among the array elements corresponding to the embedding vector value generated based on the input search term—is not a value of '0'.

[0106] The above-described flowchart and description are merely examples and may be implemented differently in some embodiments. For instance, in some embodiments, the order of each step may be changed, some steps may be repeated, some steps may be omitted, or some steps may be added.

[0107] The method described above may be provided as a computer program stored on a computer-readable recording medium for execution on a computer. The medium may continuously store a computer-executable program, or temporarily store it for execution or download. Additionally, the medium may be various recording or storage means in the form of a single or multiple hardware components, and may not be limited to a medium directly connected to a computer system but may exist distributed over a network. Examples of media may include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and media configured to store program instructions, including ROM, RAM, and flash memory. Furthermore, other examples of media may include recording or storage media managed by app stores that distribute applications or sites and servers that supply or distribute various other software.

[0108] The methods, operations, or techniques of the present disclosure may be implemented by various means. For example, these techniques may be implemented in hardware, firmware, software, or a combination thereof. Those skilled in the art will understand that the various exemplary logical blocks, modules, circuits, and algorithmic steps described in connection with the disclosure herein may be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate such interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and steps have been generally described above in terms of their functional aspects. Whether such functions are implemented in hardware or in software depends on the design requirements imposed on the specific application and the overall system. Those skilled in the art may implement the functions described in various ways for each specific application, but such implementations should not be construed as departing from the scope of the present disclosure.

[0109] In a hardware implementation, the processing units used to perform the techniques may be implemented in one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, electronic devices, other electronic units designed to perform the functions described in this disclosure, computers, or a combination thereof.

[0110] Accordingly, the various exemplary logic blocks, modules, and circuits described in connection with the present disclosure may be implemented or performed by any combination of general-purpose processors, DSPs, ASICs, FPGAs or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or those designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, for example, a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors coupled with a DSP core, or any other combination of configurations.

[0111] In firmware and / or software implementations, techniques may be implemented as instructions stored on a computer-readable medium such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, compact disc (CD), magnetic or marked data storage device, etc. The instructions may be executable by one or more processors, and may cause the processor(s) to perform specific aspects of the functions described in this disclosure.

[0112] When implemented in software, the techniques described above may be stored on a computer-readable medium as one or more instructions or code, or transmitted through a computer-readable medium. Computer-readable media include both computer storage media and communication media, including any medium that facilitates the transmission of a computer program from one place to another. Storage media may be any available media accessible by a computer. As a non-limiting example, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium accessible by a computer that can be used to transfer or store desired program code in the form of instructions or data structures. Additionally, any connection is appropriately referred to as a computer-readable medium.

[0113] For example, if software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair cable, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, coaxial cable, fiber optic cable, twisted pair cable, digital subscriber line, or wireless technologies such as infrared, radio, and microwave are included within the definition of a medium. As used herein, disks and discs include CDs, laser discs, optical discs, DVDs (digital versatile discs), floppy disks, and Blu-ray discs, wherein disks usually play data magnetically, whereas discs play data optically using a laser. The above combinations should also be included within the scope of computer-readable media.

[0114] The software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other known form of storage medium. An exemplary storage medium may be connected to a processor so that the processor can read information from the storage medium or write information to the storage medium. Alternatively, the storage medium may be integrated into the processor. The processor and the storage medium may exist within an ASIC. The ASIC may exist within a user terminal. Alternatively, the processor and the storage medium may exist as separate components within the user terminal.

[0115] Although the embodiments described above have been described as utilizing aspects of the subject matter disclosed herein in one or more standalone computer systems, the present disclosure is not limited thereto and may be implemented in conjunction with any computing environment, such as a network or a distributed computing environment. Furthermore, aspects of the subject matter in the present disclosure may be implemented in a plurality of processing chips or devices, and storage may be similarly affected across a plurality of devices. Such devices may include PCs, network servers, and portable devices.

[0116] Although the present disclosure has been described in relation to some embodiments, various modifications and changes may be made without departing from the scope of the present disclosure as understood by a person skilled in the art to which the invention of the present disclosure pertains. Furthermore, such modifications and changes should be considered to fall within the scope of the claims appended to this specification.

Claims

1. A data retrieval method performed by at least one processor, A step of generating keyword and context information based on the input search term; A step of identifying a node among a plurality of nodes that is a data search target based on the above keyword, the above context information, and the first layer search index - each of the plurality of nodes includes a data storage and a second layer search index for searching data stored in the data storage -; A step of transmitting the keyword and the context information to the identified node; and A step of receiving a data search result including data searched from the identified node based on the keyword, the context information, and the second layer search index stored in the identified node. A data retrieval method including 2. In Paragraph 1, A step of receiving access permission information for at least one of the data stored in each of the plurality of nodes or the second layer search index stored in each of the plurality of nodes from each of the plurality of nodes. Includes more, The step of identifying the node to be the target of the above data search is, A step of identifying a node among the plurality of nodes that is the target of data search based on the access permission information received from each of the plurality of nodes. A data retrieval method including 3. In Paragraph 2, A step of transmitting a program designated to each of the plurality of nodes mentioned above. Includes more, A data retrieval method in which the above access rights information is generated through the above-mentioned designated program installed on each of the above-mentioned multiple nodes.

4. In Paragraph 1, The first layer search index and the second layer search index stored in each of the plurality of nodes have a graph data structure, A data search method in which the second layer search index stored in each of the plurality of nodes is connected to the first layer search index through the edges of the graph data included in the first layer search index.

5. In Paragraph 1, The second layer search index stored in each of the plurality of nodes is, A data retrieval method comprising a first embedding vector value for data stored in each of the plurality of nodes.

6. In Paragraph 5, The above-mentioned first-level search index is, A data retrieval method comprising at least one of information for each of the plurality of nodes, metadata for data stored in each of the plurality of nodes, a first embedding vector value included in the second layer search index stored in each of the plurality of nodes, or a filter value using the first embedding vector value included in the second layer search index stored in each of the plurality of nodes.

7. In Paragraph 6, The step of identifying the node to be the target of the above data search is, A step of generating a second embedding vector value based on at least one of the above keywords or the above context information; A step of identifying the association between the data stored in each of the plurality of nodes and the search term using a filter value and a second embedding vector value using the first embedding vector value included in the second layer search index stored in each of the plurality of nodes - wherein the filter value corresponds to an array obtained by applying a filter to the first embedding vector value, and each element of the array includes information regarding the number of times a hash value obtained by inputting the first embedding vector value into each of the plurality of hash functions included in the filter -; and A step of identifying the node among the plurality of nodes that is the target of the data search based on the aforementioned identified association. A data retrieval method including 8. In Paragraph 7, The step of identifying the association between the data stored in each of the plurality of nodes and the search term is: A step of identifying that there is an association between the data stored in each of the plurality of nodes and the search term if the value of the element corresponding to the second embedding vector value among the elements of the array is greater than or equal to a specified value. A data retrieval method including 9. A computer-readable, non-transient recording medium recording instructions for executing the method according to paragraph 1 on a computer.

10. In an electronic device, Memory; and It includes at least one processor connected to the memory and configured to execute at least one computer-readable program contained in the memory, and The above at least one program is, Based on the entered search term, generate keyword and context information, and Based on the above keyword, the above context information, and the first layer search index, a node among a plurality of nodes that is the target of data search is identified - each of the plurality of nodes includes a data repository and a second layer search index for searching data stored in the data repository -, Transmit the keyword and context information to the identified node, and An electronic device comprising instructions for receiving a data search result including data searched from the identified node based on the keyword, the context information, and the second layer search index stored in the identified node.