A multi-source private domain data retrieval system for data space
By collaborating with a global retrieval agent and federated private domain data sources, and employing decentralized data indexing and retrieval/access decoupling technology, the problem of cross-domain retrieval and data rights protection for multi-source heterogeneous private domain data is solved, enabling cross-domain retrieval while protecting data privacy and rights.
Patent Information
- Application Number
- CN202510973326.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-07-15
AI Technical Summary
In the data space, it is difficult to achieve cross-domain retrieval of multi-source heterogeneous private domain data, and traditional data discovery solutions are difficult to protect data rights, especially the security of privacy information.
It adopts a decentralized data indexing and discovery paradigm, and achieves cross-domain retrieval by collaborating with private domain data sources in a global retrieval proxy. It utilizes the idea of decoupling retrieval and access, and only obtains data identifiers without directly accessing private domain data.
While meeting the needs of data rights protection, it enables cross-domain retrieval of multi-source private domain data, protecting the privacy and rights of data owners.
Smart Images

Figure CN120822241B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, in particular to a multi-source private domain data retrieval system for data space. BACKGROUND
[0002] Data space is a complex ecosystem containing massive multi-source heterogeneous private domain data. As an important part of this space, private domain data carries key information in different fields, scenarios and organizations, covers multiple dimensions such as business operation, social governance and scientific research, and contains huge economic potential and social significance. It is urgent to discover and utilize it through efficient technical means to fully release its value.
[0003] However, due to data rights considerations, data owners usually do not want data to leave their control domain, making it difficult to establish centralized indexes for scattered multi-source private domain data to achieve cross-domain retrieval using traditional data discovery solutions. In traditional data discovery solutions, the retrieval agent usually takes charge of retrieval and returns original data, which makes private domain data containing privacy information likely to be leaked to the retrieval agent, so that data rights (i.e. privacy and rights of data owners) are difficult to be guaranteed. SUMMARY
[0004] The purpose of the embodiments of the present application is to provide a multi-source private domain data retrieval scheme for data space, which can realize cross-domain retrieval of multi-source private domain data on the premise of meeting the demand of data rights guarantee.
[0005] In a first aspect, the embodiments of the present application provide a multi-source private domain data retrieval system for data space, which comprises a client, a global retrieval agent and a plurality of federated private domain data sources in the data space which are in communication connection with the global retrieval agent, wherein:
[0006] Each of the private domain data sources is configured to store private domain data of a data owner and establish an index to provide local retrieval service and local access service of the private domain data;
[0007] The client is configured to receive a user query string, request a retrieval result list from the global agent according to the user query string, request corresponding private domain data from the local access service of one or more private domain data sources to which the user query string is directed according to the retrieval result list, and generate a query result corresponding to the user query string according to the obtained private domain data, wherein the retrieval result list comprises data identifiers of private domain data in one or more private domain data sources to which the user query string is directed and which are semantically related to the user query string;
[0008] The global search agent requests a search result list from a local search service of one or more private domain data sources to which the user query string is directed, integrates the obtained search result list, and returns the integrated search result list to the client to provide the federated search service of the multi-source private domain data.
[0009] In a second aspect, the present application provides a multi-source private domain data search method for a data space, which comprises the following steps:
[0010] In a second aspect, the present application provides a multi-source private domain data search method for a data space, which comprises the following steps:
[0011] Each private domain data source stores private domain data of a data owner and establishes an index to provide a local search service and a local access service of the private domain data.
[0012] The client requests a search result list from the global agent according to the user query string, requests corresponding private domain data from a local access service of one or more private domain data sources to which the user query string is directed according to the search result list, and generates a query result corresponding to the user query string according to the obtained private domain data, wherein the search result list comprises data identifiers of private domain data related to the user query string in the one or more private domain data sources.
[0013] The global search agent requests a search result list from a local search service of one or more private domain data sources to which the user query string is directed, integrates the obtained search result list, and returns the integrated search result list to the client to provide the federated search service of the multi-source private domain data.
[0014] In a third aspect, the present application provides a computer program product comprising computer programs / instructions, which, when executed by a processor, implement the steps of the multi-source private domain data search method for a data space according to the second aspect.
[0015] In a fourth aspect, the present application provides a computer readable storage medium, which stores computer programs / instructions, which, when executed by a processor, implement the steps of the multi-source private domain data search method for a data space according to the third aspect.
[0016] In a fifth aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, wherein the processor implements the steps of the data space-oriented multi-source private domain data retrieval method according to the second aspect when executing the program.
[0017] As can be seen from the above technical solutions, the present application uses a global retrieval agent and a plurality of federated private domain data sources in the data space that are in communication connection with the global retrieval agent to provide federated retrieval services for multi-source private domain data for users, thereby using a decentralized data indexing and discovery paradigm to use the retrieval capability of the private domain data sources locally to achieve cross-domain retrieval without the dispersed private domain data leaving the local private domain data sources (i.e., without leaving the control domain of the data owners); and the present application uses the idea of decoupling retrieval and access, and designs the retrieval and access processes to be completed by two independent modules, i.e., the global retrieval agent focuses on the retrieval and discovery of data and only obtains data identifiers as retrieval results without directly accessing the private domain data; and the client as a user agent accesses the corresponding private domain data according to the data identifiers, thereby making the global retrieval agent unable to contact the private domain data in the retrieval stage, so as to protect the privacy and rights of the data owners. In this way, the present application achieves cross-domain retrieval for multi-source private domain data on the premise of meeting the demand for data rights protection. BRIEF DESCRIPTION OF DRAWINGS
[0018] In order to more clearly illustrate the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0019] Figure 1 A structural schematic diagram of a data space-oriented multi-source private domain data retrieval system provided by an embodiment of the present application;
[0020] Figure 2 A process schematic diagram of a system providing retrieval enhancement generation services online provided by an embodiment of the present application;
[0021] Figure 3 A flowchart of a data space-oriented multi-source private domain data retrieval method provided by an embodiment of the present application;
[0022] Figure 4 A schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0023] With reference to the drawings, the technical solutions in the embodiments of the present application will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0024] Organizations such as enterprises, government agencies, and research institutions accumulate a large amount of high-value and high-quality data in the process of operation. These data are usually scattered and stored in different software and hardware environments, forming a data space containing massive multi-source heterogeneous data.
[0025] Private domain data, as an important part of this space, carries key information in different fields, scenarios, and organizations, covering multiple dimensions such as business operation, social governance, and scientific research. These data have significant differences in source, format, structure, and semantics, and together form a complex and diverse data ecosystem. Private domain data contains great economic potential and social significance, and needs to be discovered and utilized through efficient technical means to fully release its value.
[0026] However, in the data space, the effective integration and discovery of multi-source heterogeneous private domain data face two important challenges:
[0027] (1) Multi-source heterogeneous private domain data is difficult to discover.
[0028] Private domain data in the data space has strong dispersion. Specifically, these data are usually scattered and stored in different software and hardware environments, belonging to different data owners, forming a highly distributed data storage and management architecture.
[0029] This dispersion makes it difficult to search cross-domain data, so it is difficult to be effectively discovered. Moreover, since private domain data may contain a large amount of private information with very high value, data owners usually do not want data to leave their control domain for data rights considerations, resulting in that the scattered private domain data is difficult to establish centralized indexing using traditional data discovery solutions, thus forming a data island phenomenon.
[0030] It should be noted that centralized indexing is a common paradigm for multi-source scattered data discovery in traditional data discovery solutions. Its technical principle is to periodically crawl public data on the Internet through a web crawler, and store these data in a centralized index library. Then, a centralized search engine establishes a mapping relationship between keywords and documents through structured processing of the crawled data by using technologies such as inverted index, so as to support users to quickly obtain relevant results through keyword query.
[0031] However, this centralized paradigm has high efficiency and convenience in processing public data, but has significant limitations in the private domain data discovery scenario in the data space. In scenarios with high privacy needs, data owners usually do not want their data to be out of the control domain, so it is difficult to establish a global data view through data crawling and centralized indexing, leading to data island phenomenon, so it is difficult to implement cross-domain retrieval for multi-source private domain data.
[0032] (2) Data rights are difficult to protect.
[0033] In the traditional data discovery scheme, the retrieval agent usually takes charge of retrieval and returning of original data at the same time, leading to high coupling between retrieval and access processes, which, if applied to retrieval of private domain data, may cause private domain data containing privacy information to be leaked to the "honest but curious" retrieval agent, thereby threatening the privacy and rights of data owners.
[0034] Based on the above analysis, in order to solve the problem that related technologies are difficult to implement cross-domain retrieval for multi-source private domain data on the premise of meeting the demand of data rights protection, the embodiments of the present application provide a multi-source private domain data retrieval scheme for data space, which realizes decentralized data indexing and discovery functions for data space on the premise of ensuring local storage and management of private domain data, to support cross-domain retrieval; and by decoupling the retrieval and access processes, the leakage of private domain data can be avoided in the retrieval stage, ensuring that the privacy and rights of data owners are effectively protected, thereby realizing cross-domain retrieval for multi-source private domain data on the premise of meeting the demand of data rights protection.
[0035] First, in order to facilitate understanding of the technical solutions provided by the present application, the main technical concepts involved in the embodiments of the present application are briefly described below.
[0036] Data space (Data Space) is a technical framework that supports safe circulation and collaborative governance of large-scale heterogeneous data. Its core is to break down data silos through the "on-demand integration" mode, and realize logical unification and sovereign controllability of cross-domain data.
[0037] Retrieval-Augmented Generation (RAG) is an emerging technology that combines information retrieval and generative models. Its core principle is to extract relevant data from an external knowledge base based on a user query string, and then input the retrieval results as context into a large language model (LLM) to enhance the accuracy and timeliness of the generation process. Specifically, when a user submits a query string, RAG first efficiently locates associated information from structured or unstructured data sources using vector search or keyword matching, then integrates this information into the prompt words of the LLM through dynamic embedding technology, and finally generates a high-quality response that incorporates external knowledge from the LLM.
[0038] Chroma: It is an open-source vector database framework that provides approximate nearest neighbor search services.
[0039] LangChain: It is a framework for developing applications driven by large language models.
[0040] Referring to Figure 1 The system includes a client, a global retrieval agent, and multiple federated private domain data sources in the data space that are communicatively connected to the global retrieval agent. The system includes the following components:
[0041] Each private domain data source is configured to store private domain data of a data owner and establish an index to provide local retrieval services and local access services for the private domain data.
[0042] The client is configured to receive a user query string, request a retrieval result list from the global agent based on the user query string, request corresponding private domain data from the local access services of one or more private domain data sources corresponding to the user query string based on the retrieval result list, and generate a query result corresponding to the user query string based on the obtained private domain data. The retrieval result list includes data identifiers of private domain data in one or more private domain data sources corresponding to the user query string that are semantically related to the user query string.
[0043] The global retrieval agent is configured to request a retrieval result list from the local retrieval services of one or more private domain data sources corresponding to the user query string, integrate the obtained retrieval result list, and return it to the client to provide federated retrieval services for multiple private domain data.
[0044] In this embodiment, the multi-source private domain data retrieval system facing the data space is mainly composed of three modules, i.e., multiple private domain data sources in the data space, a global retrieval agent (of a third party), and a client.
[0045] There are a large number of private domain data sources in the data space, which are responsible for storing private domain data of data owners and establishing indexes to provide local access services and local retrieval services to the outside. By joining multiple private domain data sources in the data space to the federation environment of the global retrieval agent, such as establishing a communication connection between the global retrieval agent and the private domain data sources and recording the private domain data sources as federated private domain data sources, the global retrieval agent can call the local retrieval services of the private domain data sources to provide federated retrieval services of multi-source private domain data. The client is responsible for providing query services to users, which obtains private domain data corresponding to a user query string (which describes the information that the user wants to query) by calling the federated retrieval services of multi-source private domain data provided by the global retrieval agent and the local access services provided by the private domain data sources, processes the obtained private domain data (such as summarizing or extracting key information, etc.) to generate a query result corresponding to the user query string. Thus, based on the federated retrieval technology, the present application realizes the discovery of multi-source private domain data in the data space without leaving the local, thereby ensuring the control ability of the private domain data of the data owner.
[0046] As can be seen from the above technical solution, the present application uses the global retrieval agent and multiple federated private domain data sources in the data space that are in communication connection with the global retrieval agent to provide federated retrieval services of multi-source private domain data to users, thereby using a decentralized data indexing and discovery paradigm to realize cross-domain retrieval using the local retrieval capability of the private domain data source without leaving the local of the private domain data (i.e., without leaving the control domain of the data owner). The present application adopts the idea of decoupling retrieval and access, and designs the retrieval and access processes to be completed by two independent modules, i.e., the global retrieval agent focuses on the retrieval and discovery of data and only obtains data identifiers as retrieval results without directly accessing private domain data; the client as a user agent accesses the corresponding private domain data according to the data identifier, thereby making the global retrieval agent unable to contact the private domain data in the retrieval stage, thereby protecting the privacy and rights of the data owner. In this way, the present application realizes cross-domain retrieval for multi-source private domain data while meeting the demand for data rights protection.
[0047] In an optional embodiment, the client is configured to generate the query result corresponding to the user query string by the following steps:
[0048] organize the obtained private domain data into a context related to the user query string, and fill a pre-defined prompt template with the user query string and the context related thereto to obtain a target prompt;
[0049] input the target prompt into a large language model to generate the query result, wherein the large language model is configured to generate the query result corresponding to the user query string through a prompt-enhanced generation process to support the client to provide the retrieval-enhanced generation service to the user.
[0050] In a specific implementation, the client serves as a user agent and is responsible for retrieving and accessing the private domain data related to the user query string and connecting an externally or locally deployed large language model to provide the retrieval-enhanced generation service. Optionally, the system supports the user to select the large language model according to the user's own needs.
[0051] Specifically, the client obtains the private domain data related to the user query string by invoking the federated retrieval service provided by the global retrieval agent for the multi-source private domain data and the local access service provided by the private domain data source; then organizes the obtained private domain data into a context related to the user query string to guide the large language model to generate the query result corresponding to the user query string, so as to utilize the powerful information mining capability of the generative large language model to integrate the information in the multi-source heterogeneous private domain data in the data space discovered through the federated retrieval and extract the key information to provide to the user.
[0052] In this embodiment, considering that the private domain data generated by different entities and environments and the management requirements and use scenarios of the data owners for the private domain data are different, the private domain data often has strong heterogeneity in terms of format, field, etc. Such heterogeneity makes it difficult for traditional data discovery methods to effectively extract information from the retrieved private domain data, further increasing the complexity of data retrieval and information integration, and limiting the usability and value of the private domain data. Moreover, traditional data discovery schemes (such as centralized search engines) usually return query results in a simple list form, and the list-form query results are difficult to meet the user's integration requirements for complex information and are difficult to effectively extract and utilize the value contained in the private domain data. To solve the foregoing problems, the retrieval-enhanced generation technology is introduced to effectively integrate and utilize the private domain data. Thus, by combining the federated retrieval technology and the retrieval-enhanced generation technology, the retrieval and utilization efficiency of the multi-source heterogeneous private domain data can be effectively improved.
[0053] In an optional embodiment, each private domain data source is further configured to perform the following steps:
[0054] In response to receiving the request for accessing the private domain data from the client, detecting access rights of a user associated with the client based on a preset access control policy;
[0055] In a case where it is detected that the user associated with the client has the access rights, returning corresponding private domain data to the client according to the request for accessing the private domain data.
[0056] In this embodiment, the application decouples the retrieval and access processes, so that the global retrieval agent cannot access the original data (i.e., the private domain data) in the retrieval stage, and the private domain data source can effectively manage the end users of the private domain data through the access control, thereby guaranteeing the data rights in the mechanism.
[0057] In an optional embodiment, the client is configured with a client retriever and a client accessor, each of the private domain data sources is configured with a digital object registry and a digital object repository, and the global retrieval agent is configured with a global retriever and a data source list configured to store information of a plurality of federated private domain data sources in communication connection with the global retrieval agent, wherein:
[0058] The client retriever is configured to, in response to receiving a user query string, encapsulate the user query string into a first retrieval request package and send the first retrieval request package to the global retriever to request a retrieval result from the global retriever;
[0059] The global retriever is configured to, in response to receiving the first retrieval request package, parse the first retrieval request package to obtain the user query string, read digital object registry addresses of one or more private domain data sources corresponding to the user query string from the data source list to generate a first address list, encapsulate the user query string into one or more second retrieval request packages, and distribute the one or more second retrieval request packages to the digital object registries of the one or more private domain data sources corresponding to the user query string according to the first address list.
[0060] The digital object registry of each of the private domain data sources is configured to, in response to receiving the second retrieval request package, parse the second retrieval request package to obtain the user query string, retrieve a digital object associated with itself according to the user query string, generate a first retrieval result list according to a digital object identifier of the retrieved digital object, encapsulate an address of a digital object repository associated with itself and the first retrieval result list into a first response package, and return the first response package to the global retriever, the first retrieval result list including a digital object identifier of a digital object related to the semantic of the user query string, the digital object being obtained by encapsulating the private domain data.
[0061] The global searcher is configured to, in response to receiving one or more first response packets, parse the one or more first response packets to obtain the addresses of one or more digital object repositories and a first search result list, integrate the addresses of the one or more digital object repositories and the first search result list to generate a second address list and a second search result list, encapsulate the second search result list and the second address list into a second response packet, and return it to the client searcher.
[0062] The client retrieval is configured to, in response to receiving the second response packet, parse the second response packet to obtain the second retrieval result list and the second address list, and send them to the client accessor;
[0063] The client accessor is configured to, in response to receiving the second search result list and the second address list, send a request packet for the data entity of the digital object associated with the second search result list to the digital object repository associated with the second address list;
[0064] Each of the private domain data source's digital object repositories is configured to return the corresponding data entity to the client accessor in response to receiving a request packet for the data entity of the digital object.
[0065] Among them, the digital object repository and digital object registry of the private domain data source each have their own unique identifier and address.
[0066] In this embodiment, the client consists of two subsystems: a client retrieval system and a client access system. These subsystems serve as the user's multi-source private domain data retrieval entry point and data access proxy, respectively, collaboratively executing user-oriented federated retrieval services. The client does not need to maintain any private domain data source information; it only needs to maintain information about the global retrieval proxy (such as the address of the global retrieval system) to send the first retrieval request packet. This allows the global retrieval system to invoke the local retrieval services of each private domain data source to provide user-oriented federated retrieval services. It is understood that the client can further invoke a large language model to provide enhanced user-oriented federated retrieval generation services.
[0067] Specifically, the global retrieval agent uses a global retrieval tool to call the local retrieval services of each private domain data source to obtain retrieval results, and integrates the obtained retrieval results to provide a federated retrieval service for multi-source private domain data. Furthermore, to accurately record data source information, this application designs the global retrieval agent to maintain a list of data sources to store information about each private domain data source in the federation, such as the identifier of each private domain data source, the address of the digital object registry, the address of the digital object repository, and the metadata stored by the private domain data source.
[0068] It can be understood that, in order to separate the data identifier and the data entity to realize the guarantee of the data right, the private domain data in the data space is encapsulated as a digital object, and then the digital object architecture is applied to the multi-source private domain data retrieval system in the data space.
[0069] The digital object model, as one of the core models of the digital object architecture, can be mainly divided into three parts: a digital object identifier, metadata, and a data entity. The digital object identifier is a digital object identity (i.e., a decentralized identity), which should be a globally unique and persistent mark of the digital object; the metadata is information containing the characteristics of the digital object (such as the creator, etc.), and thus can be used for the search and discovery of the digital object; and the data entity is the actual data part stored by the digital object, which is composed of multiple structured or unstructured data elements. In order to support the basic functions of digital object interoperability, the digital object architecture proposes three core system components, namely a digital object identifier system, a digital object registry, and a digital object repository, which respectively manage the identifier, metadata, and data entity of the digital object.
[0070] The private domain data in the data space is encapsulated as a digital object, and the separation storage and management of the data object identifier, metadata, and data entity are realized for the obtained digital object. Moreover, the index and storage capabilities of the digital object registry and the digital object repository are utilized to realize the retrieval and access separation retrieval pipeline.
[0071] Specifically, the metadata part of the digital object is stored and managed by using the digital object registry, and the digital object registry is designed to provide a digital object retrieval interface to receive a retrieval request (such as the second retrieval request package) sent by a user or other components (such as a global retriever), and can respond to the retrieval request by establishing an inverted index for the metadata (which contains information describing the characteristics of the digital object) therein, and return the digital object identifier of the relevant digital object retrieved (i.e., the first retrieval result list). At the same time, the data entity part (i.e., the private domain data) of the digital object is stored and managed by using the digital object repository, such as providing management services of addition, deletion, modification, and query, and the digital object repository is designed to provide an access interface of the digital object to receive a data request (such as the request package of the data entity in the digital object associated with the second retrieval result list) sent by a user or other components (such as a client accessor), and then return the corresponding data entity (i.e., return the corresponding private domain data). In this way, the data index and storage within the private domain data source are realized by the digital object registry and the digital object repository respectively, so as to avoid data leakage by utilizing the separation of the digital object identifier and the data entity.
[0072] Optionally, the client retriever is further configured to encapsulate the preset first retrieval parameter set and the user query string into the first retrieval request package, wherein the first retrieval parameter set at least includes a maximum total number of requested retrieval results, and other parameters can be customized and extended by a system developer, for example, the maximum number of retrieval results requested from each private domain data source can also be included.
[0073] The global retriever is further configured to parse the first retrieval request package to obtain the user query string and the first retrieval parameter set, determine, according to the first retrieval parameter set, a second retrieval parameter set corresponding to each of one or more private domain data sources to which the user query string is directed, encapsulate the second retrieval parameter set corresponding to each of the one or more private domain data sources and the user query string into the one or more second retrieval request packages respectively, and the second retrieval parameter set at least includes a maximum number of retrieval results requested from a single private domain data source.
[0074] The digital object registry of each private domain data source is further configured to parse the second retrieval request package to obtain the user query string and the second retrieval parameter set, perform retrieval on the digital object associated with itself according to the user query string and the second retrieval parameter set, and intercept each retrieval result obtained according to relevance to generate the first retrieval result list, wherein the retrieval result includes a digital object identifier of the retrieved digital object, and each retrieval result is sorted in descending order of semantic relevance between the private domain data corresponding to the retrieval result (i.e., the digital object identifier) and the user query string, and the first five retrieval results are intercepted as the first retrieval result list.
[0075] Optionally, the digital object warehouse of each private domain data source is further configured to perform encapsulation operation on each private domain data in the private domain data source associated with itself as a data entity to obtain each digital object, and embed the data entity of each digital object by a deep encoder to obtain each dense vector in a semantic space, wherein the digital object includes a digital object identifier, a data entity, and metadata.
[0076] The digital object warehouse of each private domain data source is further configured to store the digital object and the dense vector corresponding to the digital object, and construct a first mapping relationship, wherein the first mapping relationship includes a mapping relationship from a digital object identifier to a data entity (i.e., private domain data) of the digital object and a dense vector corresponding to the data entity, to support a service of accessing the private domain data according to the digital object identifier.
[0077] The digital object registry of each private domain data source is further configured to build a second mapping relationship and a dense vector index, so that after a user query string is matched in a semantic space constituted by the dense vectors to obtain a dense vector matched by the user query string in the dense vector index, the first search result list is generated according to the second mapping relationship, the second mapping relationship being a mapping relationship from a dense vector corresponding to a digital object to an identification of the digital object corresponding to the dense vector, and the dense vector index including the dense vector corresponding to each digital object associated with the digital object registry.
[0078] In this embodiment, the private domain data source acts as an agent of the data owner, and provides local search services and local access services of the private domain data by deploying a digital object registry and a digital object warehouse. In order to improve the semantic relevance of local search, the digital object warehouse encapsulates each piece of raw data (i.e. private domain data) as a digital object, and embeds the data entity part of the digital object by a deep encoder (e.g. Transformer) to obtain a dense vector in a semantic space. For a data entity of a digital object and a dense vector obtained by encoding the data entity, the data owner can uniformly assign a unique digital object identification to both.
[0079] The digital object warehouse stores the digital objects and the corresponding dense vectors, and builds a mapping from the digital object identification to the raw data (i.e. private domain data, i.e. data entity of the digital object) and the dense vector, thereby providing a service of accessing the private domain data according to the data identification (i.e. digital object identification) of the private domain data.
[0080] The digital object registry builds a mapping from the dense vector of the private domain data to the data identification of the private domain data and indexes it, such as performing k-nearest neighbor search on the input user query string in the dense vector index. It can be understood that, in order to avoid leakage of the private domain data in the search process, the digital object registry only returns the data identification (not the raw data) of the private domain data semantically related to the user query string based on the mapping, thereby making the global search agent unable to access the private domain data unrestrictedly.
[0081] Optionally, the dense vector is obtained by the following steps:
[0082] First, the data entity is divided into pieces according to a set length;
[0083] Second, each piece of natural language description of the data entity is input into a pre-trained deep encoder, so that the deep encoder encodes the data entity pieces by attention mechanism to output a dense vector of fixed length.
[0084] Optionally, the dense vector index and the digital object repository are implemented using an open-source vector database framework Chroma.
[0085] Optionally, the client accessor is further configured to provide a retrieval augmentation generation service to a user based on a large language model, wherein the client is implemented based on a large language model programming framework LangChain to implement a retrieval augmentation generation pipeline; and the client is communicatively connected with the large language model through an application programming interface (API), and the client retriever and the client accessor are implemented by decoupling retrieval functions and access functions of the retriever in LangChain, respectively.
[0086] It can be understood that, by introducing large models and vector storage capabilities through frameworks such as LangChain and Chroma, the system can effectively integrate the understanding and information integration capabilities of large models, facilitate matching of user query strings in a semantic space constituted by dense vectors, and enhance the scalability of the system.
[0087] Optionally, each of the private domain data sources and the global retriever is implemented based on a lightweight web application framework Flask.
[0088] Optionally, the information of the private domain data sources in the data source list at least includes the identities and addresses of the digital object registries and the digital object repositories in the private domain data sources.
[0089] The global retriever is further configured to, in a case where a request to join the federated environment is received from any of the private domain data sources in the data space, record the identities and addresses of the digital object registries and the digital object repositories carried by the request to join the federated environment in the data source list, to indicate that the private domain data source has joined the federated environment.
[0090] Exemplarily, the retrieval augmentation generation pipeline of the system includes two stages: a start preparation stage and an online service stage.
[0091] In the start preparation stage, the private domain data sources in the data space join the federated environment of the global retriever, and the global retriever constructs a semantic representation of the data sources, and the specific process is as follows:
[0092] 1) The private domain data sources send a request to join the federation to the global retriever;
[0093] 2) After receiving the request, the global search agent records the identification of the private domain data source and its corresponding data source address in the data source list, indicating that the private domain data source has joined the federation, thereby providing support for subsequent request distribution (such as the distribution of the second search request package, and the distribution of the request package of the private domain data associated with the second search result list, etc.).
[0094] For the online service stage, refer to the process diagram of the system online providing the search enhancement generation service shown in Figure 2 The process of the system online providing the search enhancement generation service is as follows:
[0095] (1) User query.
[0096] In this step, the user sends a query request (containing a query string) to the client, which may be for private domain data in one or more private domain data sources.
[0097] (2) Request search.
[0098] In this step, after the client receives the user's query request, the query string and the search parameter set (i.e., the preset first search parameter set) are encapsulated into a query request package (i.e., the first search request package) by the client searcher, and sent to the global search agent.
[0099] (3) Obtain data source identification and address.
[0100] In this step, after the global search agent receives the query request package, it parses the query string and search parameter set therein. The global search agent sets the maximum number of search results requested from each private domain data source, extracts other possible parameters, and reads the address of the relevant digital object registry from the data source list.
[0101] (4) Distribute search request.
[0102] In this step, the global searcher of the global search agent encapsulates the user query string and other possible search parameters into a query request package (i.e., the second search request package); then, the global searcher distributes the query request package to the digital object registry of the selected private domain data source according to the address read in the previous step.
[0103] (5) Return search result identification.
[0104] In this step, after receiving the query request package, the digital object registry parses it to extract the user query string and retrieval parameters. Then, the digital object registry embeds the user query string into the same semantic space as the private domain data and performs approximate nearest neighbor retrieval on the dense vector index to obtain an ordered retrieval result list (i.e., the first retrieval result list, sorted by relevance). The digital object registry then encapsulates it and the address of the relevant digital object repository into a response package (i.e., the first response package) and returns it to the global retriever, where the data identifier corresponds to the original data (i.e., private domain data) semantically related to the user query string.
[0105] Exemplarily, taking the case where the retrieval parameters include the maximum value of the total number of requested retrieval results, after receiving the user query string, the private domain data source uses a deep encoder (the same deep encoder as the one used to encode the data entity in the previous text) to encode the query string into a dense vector and perform nearest neighbor retrieval with the dense vectors corresponding to the data entities of the digital objects in the digital object repository. For example, the Hierarchical Navigable SmallWorld (HNSW) algorithm can be used to calculate the closest several dense vectors. If the number of retrieval results to be returned n has been specified, the first n retrieval results (i.e., the first n digital object identifiers) are returned.
[0106] Optionally, the deep encoder of the digital object registry uses the all-mpnet-base-v2 model of the SentenceTransformers series to embed natural language text (i.e., the data entity of the digital object) into a 768-dimensional dense vector. To maintain consistency with the digital object registry, the deep encoder in the global retrieval agent also uses the all-mpnet-base-v2 model of the SentenceTransformers series.
[0107] (6) Reordering to obtain a retrieval result identifier list.
[0108] In this step, after receiving the response package from each digital object registry, the global retriever extracts the retrieval result identifier list and access address list returned by each data source. Then, the global retriever reorders each retrieval result identifier list, merges them into one list, and according to the maximum number of results required by the user, truncates the first several retrieval results (i.e., digital object identifiers) in the list as the final retrieval result identifier list (i.e., the second retrieval result list), and accordingly obtains the corresponding digital object repository access address list (i.e., the second address list).
[0109] (7) Return the retrieval result identifier list.
[0110] In this step, the global retriever adds the list of returned search results and the list of access addresses to the response data packet (i.e., the second response packet) and returns them to the client retriever.
[0111] (8) Submit the search result identifier list.
[0112] In this step, the client retriever parses the response data packet to obtain the list of search results and the list of access addresses and hands them over to the client accessor.
[0113] (9) Access the search result raw data.
[0114] In this step, the client accessor sends a request packet to the corresponding digital object repository for accessing the raw data corresponding to the data identifier according to the list of search results and the corresponding list of access addresses.
[0115] (10) Return the search result raw data.
[0116] In this step, the private domain data source can perform access control on the user on its own after receiving the raw data access request from the client (the specific access control policy is not within the scope of the application) to protect data rights. When the user has access rights, the digital object repository returns the raw data corresponding to the data identifier to the client accessor.
[0117] (11) Context enhancement.
[0118] In this step, the client accessor organizes the accessed raw data into relevant contexts, fills in the pre-defined prompt word template with the user's query string, and obtains the generated prompt word for submission to the large language model. Thus, by combining the information integration capability of the large language model, the retrieved heterogeneous data is constructed as a context to guide the large language model generation, providing highly condensed natural language form answers to the user.
[0119] (12) Return the answer.
[0120] In this step, the large language model uses the prompt word enhancement generation process to return the generated natural language answer to the client. Thus, based on the context construction capability of the large language model, the retrieved heterogeneous data is converted into a condensed answer in natural language form, improving information integration and generation efficiency.
[0121] Related data examples:
[0122] Query string: Can probiotics prevent common cold?
[0123] Retrieved private domain data: Randomized, double-blind, and placebo-controlled study on the use of new probiotic lactic acid bacteria to enhance human immune defense against viral infections. Background: The aim of this study was to investigate whether the consumption of Lactobacillus plantarum HEAL 9 (DSM 15312) and Lactobacillus paracasei 8700:2 (DSM 13434) would affect the natural acquisition of common cold infections in healthy subjects. Results: The incidence of one or more common cold episodes decreased from 67% in the control group to 55% in the probiotic group (p<0.05), as well as the number of days with cold symptoms.
[0124] Large language model generated answer (i.e., the query result corresponding to the user query string): Yes, probiotics may help prevent the common cold. A study found that consuming specific probiotic strains, Lactobacillus plantarum HEAL 9 and Lactobacillus paracasei 8700:2, reduced the incidence of common cold from 67% in the control group to 55% in the probiotic group. Additionally, the probiotic group experienced a significant reduction in the number of days with common cold symptoms compared to the control group. Other studies have also shown that probiotics can reduce the incidence and duration of cold and flu-like symptoms, particularly in children.
[0125] As can be seen from the above pipeline, the idea of decoupling retrieval and access is adopted in the design of private domain data sources, global retrieval agents, and clients. For private domain data sources in the data space, the digital object repository establishes a mapping from digital object identifiers to data entities, while the digital object registry establishes a mapping from indexes to digital object identifiers, achieving the separation of local retrieval and local access; the global retrieval agent only implements federated retrieval based on the digital object registry and does not access data entities during retrieval; the client also designs the retriever and the accessor as two modules. By decoupling the retriever in the retrieval enhancement generation pipeline into two independent modules of retrieval and access, combined with the identifier and data entity separation technology of the digital object architecture, it is ensured that the retrieval agent cannot directly access the original data, thereby protecting data rights.
[0126] In the above example, the system adopts a server-client (C / S) architecture, and the retrieval and access decoupled retrieval enhancement generation pipeline is implemented for the client, global retrieval agent, and data owner repository backend, which is programmed by Python language. The system supports the following functions:
[0127] (1) Combined with retrieval enhancement technology and large models, a discovery system for private domain data in the data space is provided, which realizes efficient information integration and extraction of heterogeneous data;
[0128] (2) Combined with federal search technology, distributed information retrieval of scattered private domain data sources is realized, and the data rights problem caused by centralized indexing is effectively avoided.
[0129] (3) Combined with the digital object architecture, a retrieval enhancement generation pipeline decoupled from retrieval and access is realized, and data leakage to unauthorized third parties during retrieval enhancement generation is avoided.
[0130] Therefore, in view of the difficulty of discovering multi-source heterogeneous private domain data, the present application designs a federal search enhancement generation architecture for multi-source private domain data, builds a global search agent for the entire data space as a retriever of the search enhancement generation system, efficiently discovers multi-source heterogeneous data in a federal search manner without leaving the storage, and uses a large language model for information extraction and integration; in view of the difficulty of guaranteeing data rights, the present application designs a new search enhancement generation pipeline for the data space, separates the retrieval stage from the data access process, and combines the digital object architecture, so that the global search agent cannot access the original data in the retrieval stage, thereby ensuring that the access process to the original data is always under the control of the data owner, to meet the needs of data rights protection.
[0131] It should be noted that through the analysis of the current data discovery technology, the present application has obtained some problems that can be optimized and solved for the data space and provided corresponding solutions:
[0132] (1) Decentralized data does not leave the local environment.
[0133] Since data owners usually have strict protection requirements for data control domains and do not want data to leave their local environment, it is difficult for decentralized data to be retrieved across domains through centralized indexing, and a data discovery system for the data space should be built based on the premise of ensuring the local storage and management of data, therefore, the present application adopts a decentralized data indexing and discovery method to solve the foregoing problems.
[0134] (2) Data is multi-source and heterogeneous, and the information integration capability of the data discovery system should be improved.
[0135] The data in the data space has significant differences in source, format and semantics, such as the mixture of structured data, semi-structured data and unstructured data, making it difficult for traditional retrieval methods to effectively integrate heterogeneous data, and this heterogeneity further increases the complexity of data discovery, therefore, the present application introduces a large language model to improve the information integration capability;
[0136] (3) Data rights protection.
[0137] In the current data discovery process, the retrieval agent is usually responsible for retrieving and returning the original data at the same time, resulting in a high coupling between retrieval and access processes. This coupling makes it possible for private domain data containing privacy information to be leaked to the "honest but curious" retrieval agent. Therefore, by avoiding the retrieval agent directly obtaining the original data, the disclosure of the original data is avoided in the retrieval stage.
[0138] Therefore, the present application designs a new data discovery and information integration system oriented to data space. On the basis of ensuring data local storage and management, a decentralized data index and discovery system oriented to data space is constructed to support cross-domain retrieval. By introducing advanced technologies such as large language models, the integration capability of the data discovery system for multi-source heterogeneous data is improved, and the problem that traditional methods are difficult to handle heterogeneous data is solved. In the retrieval stage, the disclosure of the original data is avoided, and by decoupling the retrieval and access processes, the data rights are effectively protected. Therefore, the system can protect the data rights while realizing the efficient integration and utilization of multi-source heterogeneous private domain data.
[0139] The present application also provides a multi-source private domain data retrieval method oriented to data space, as shown in Figure 3 The method comprises the following steps:
[0140] Step S101: establishing a multi-source private domain data retrieval system oriented to data space, the system comprising a client, a global retrieval agent, and a plurality of federated private domain data sources in the data space in communication connection with the global retrieval agent;
[0141] Step S102: each private domain data source stores the private domain data of the data owner and establishes an index to provide local retrieval services and local access services for the private domain data;
[0142] Step S103: the client, in response to receiving a user query string, requests a retrieval result list from the global agent according to the user query string, the retrieval result list comprising data identifiers of private domain data in one or more private domain data sources to which the user query string is directed and which are semantically related to the user query string;
[0143] Step S104: the global retrieval agent requests a retrieval result list from the local retrieval services of one or more private domain data sources to which the user query string is directed, integrates the obtained retrieval result lists, and returns them to the client to provide federated retrieval services for multi-source private domain data;
[0144] Step S105: The client queries the local access service of one or more private domain data sources corresponding to the user query string according to the search result list, and generates the query result corresponding to the user query string according to the obtained private domain data.
[0145] As a possible implementation, the client generates the query result corresponding to the user query string by the following steps:
[0146] organize the obtained private domain data into a context related to the user query string, and fill in a pre-defined prompt word template using the user query string and the related context to obtain a target prompt word;
[0147] input the target prompt word into a large language model to generate the query result, wherein the large language model is configured to generate the query result corresponding to the user query string through a prompt word enhanced generation process to support the client to provide a search enhanced generation service to the user.
[0148] As a possible implementation, the client is configured with a client retriever and a client accessor, each of the private domain data sources is configured with a digital object registry and a digital object warehouse, the global retrieval agent is configured with a global retriever and a data source list, and the data source list is configured to store information of a plurality of federated private domain data sources connected to the global retrieval agent in communication, wherein:
[0149] The client retriever, in response to receiving a user query string, encapsulates the user query string into a first retrieval request package and sends it to the global retriever to request a search result from the global retriever;
[0150] The global retriever, in response to receiving the first retrieval request package, parses the first retrieval request package to obtain the user query string, reads the digital object registry addresses of one or more private domain data sources corresponding to the user query string from the data source list to generate a first address list, encapsulates the user query string into one or more second retrieval request packages, and distributes the one or more second retrieval request packages to the digital object registries of the one or more private domain data sources corresponding to the user query string according to the first address list;
[0151] The digital object registry of each private domain data source, in response to receiving the second retrieval request package, parses the second retrieval request package to obtain the user query string, retrieves the digital object associated with itself according to the user query string, generates a first retrieval result list according to the digital object identifier of the retrieved digital object, encapsulates the address of the digital object warehouse associated with itself and the first retrieval result list into a first response package, and returns the first response package to the global retriever, wherein the first retrieval result list includes the digital object identifier of the digital object related to the semantic of the user query string, and the digital object is obtained by encapsulating private domain data;
[0152] The global retriever, in response to receiving one or more first response packages, parses the one or more first response packages to obtain the address of one or more digital object warehouses and the first retrieval result list, integrates the address of the one or more digital object warehouses and the first retrieval result list respectively to generate a second address list and a second retrieval result list, encapsulates the second retrieval result list and the second address list into a second response package, and returns the second response package to the client retriever;
[0153] The client retriever, in response to receiving the second response package, parses the second response package to obtain the second retrieval result list and the second address list, and sends them to the client accessor;
[0154] The client accessor, in response to receiving the second retrieval result list and the second address list, sends a request package of the data entity of the digital object associated with the second retrieval result list to the digital object warehouse associated with the second address list;
[0155] The digital object warehouse of each private domain data source, in response to receiving the request package of the data entity of the digital object, returns the corresponding data entity to the client accessor.
[0156] As a possible implementation, the method further comprises:
[0157] The client retriever encapsulates a preset first retrieval parameter set and the user query string into the first retrieval request package, and the first retrieval parameter set at least includes the maximum total number of requested retrieval results;
[0158] The global retriever parses the first retrieval request package to obtain the user query string and the first retrieval parameter set, determines one or more second retrieval parameter sets corresponding to one or more private domain data sources respectively to which the user query string is directed according to the first retrieval parameter set, and encapsulates the one or more second retrieval parameter sets corresponding to the one or more private domain data sources respectively and the user query string as the one or more second retrieval request packages, wherein the second retrieval parameter set at least includes a maximum number of retrieval results requested from a single private domain data source.
[0159] The digital object registry of each private domain data source parses the second retrieval request package to obtain the user query string and the second retrieval parameter set, performs retrieval on the digital object associated with itself according to the user query string and the second retrieval parameter set, and intercepts each retrieval result obtained according to the relevance to generate the first retrieval result list, wherein the retrieval result includes a digital object identifier of the digital object retrieved.
[0160] As a possible implementation, the data identifier of the private domain data is that the digital object warehouse of each private domain data source encapsulates each private domain data in the private domain data source associated with itself as a data entity to obtain each digital object, and embeds the data entity of each digital object respectively by a deep encoder to obtain each dense vector in a semantic space, wherein the digital object includes a digital object identifier, a data entity, and metadata.
[0161] The digital object warehouse of each private domain data source stores the digital object and the dense vector corresponding thereto, and constructs a first mapping relationship, wherein the first mapping relationship includes a mapping relationship from a digital object identifier to a data entity of the digital object and a dense vector corresponding thereto, to support a service of accessing private domain data according to the digital object identifier;
[0162] The digital object registry of each private domain data source constructs a second mapping relationship and a dense vector index, so that after obtaining the dense vector matched by the user query string in the dense vector constituted semantic space in the dense vector index by matching the user query string on the dense vector constituted semantic space, the first retrieval result list is generated according to the second mapping relationship, wherein the second mapping relationship is a mapping relationship from a dense vector corresponding to a digital object to a digital object identifier corresponding to the digital object, and the dense vector index includes each dense vector corresponding to each digital object associated with the digital object registry.
[0163] As a possible implementation, the dense vector index and the digital object warehouse are implemented by using an open source vector database framework Chroma.
[0164] As a possible implementation, the information of the private domain data source in the data source list at least includes: the identification and address of the digital object registry and the digital object warehouse in the private domain data source; the method further includes:
[0165] In the case that the global retriever receives a request for joining the federated environment from any private domain data source in the data space, the global retriever records the identification and address of the digital object registry and the digital object warehouse carried by the request for joining the federated environment in the data source list, to indicate that the private domain data source has joined the federated environment.
[0166] As a possible implementation, the method further includes:
[0167] The client accessor provides a retrieval enhancement generation service to the user based on a large language model, wherein the client is implemented based on a large language model programming framework LangChain, and is connected with the large language model through an application programming interface API, and the client retriever and the client accessor are implemented by decoupling the retrieval function and the access function of the retriever in LangChain.
[0168] As a possible implementation, the method further includes:
[0169] Each private domain data source, in response to receiving a request for accessing private domain data from the client, detects the access authority of the user associated with the client based on a preset access control policy;
[0170] Each private domain data source, in the case that the user associated with the client is detected to have the access authority, returns corresponding private domain data to the client according to the request for accessing private domain data.
[0171] As a possible implementation, each private domain data source and the global retrieval agent are implemented based on a lightweight web application framework Flask.
[0172] It should be noted that, for the method embodiments, in order to simply describe, they are all described as a series of action combinations, but those skilled in the art should know that the application embodiments are not limited to the action order described, because according to the application embodiments, certain steps can be performed in other order or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions involved are not necessarily necessary for the application embodiments.
[0173] The application embodiments also provide an electronic device, which refers to Figure 4 , Figure 4is a schematic diagram of an electronic device proposed by an embodiment of the present application. As shown in Figure 4 The electronic device 100 includes a memory 110 and a processor 120, the memory 110 and the processor 120 are communicatively connected through a bus, and the memory 110 stores a computer program, the computer program can run on the processor 120, and then the steps in the data space-oriented multi-source private domain data retrieval method disclosed by the embodiment of the present application are implemented.
[0174] The embodiment of the present application also provides a computer readable storage medium, which stores a computer program / instruction, and the computer program / instruction is executed by a processor to implement the data space-oriented multi-source private domain data retrieval method disclosed by the embodiment of the present application.
[0175] The embodiment of the present application also provides a computer program product, which includes a computer program / instruction, and the computer program / instruction is executed by a processor to implement the data space-oriented multi-source private domain data retrieval method disclosed by the embodiment of the present application.
[0176] Each embodiment in the specification is described in a progressive manner, and each embodiment focuses on the difference from other embodiments. The same and similar parts between each embodiment can be referred to each other.
[0177] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, device or computer program product. Therefore, the embodiments of the present application can adopt a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application can adopt the form of a computer program product implemented on one or more computer usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer usable program code.
[0178] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams according to the method, system, device, storage medium and program product of the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing terminal equipment to produce a machine, so that the instructions executed by the computer or other programmable data processing terminal equipment produce a machine that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The device that implements the functions specified in one flow or multiple flows and / or blocks Figure 1 The device that implements the functions specified in one flow or multiple flows and / or blocks
[0179] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the flow Figure 1 one or more flows and / or blocks Figure 1 one or more blocks or multiple blocks.
[0180] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the flow Figure 1 one or more flows and / or blocks Figure 1 one or more blocks or multiple blocks.
[0181] Finally, it should be noted that the terms "first" and "second" and the like are used merely to distinguish one element from another, and do not necessarily indicate a physical or chronological priority of one element over another. Furthermore, the terms "comprises", "comprising", or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. The terms "includes", "including", or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus.
[0182] The above provides a kind of multi-source private domain data retrieval system for data space provided by the present application, the principle and implementation mode of the present application are described in this paper by specific examples, the above example is only for helping to understand the method of the present application and its core idea;For the general technical personnel in the art, according to the idea of the present application, there will be changes in specific implementation mode and application range, and the above description should not be understood as the limitation of the present application.
Claims
1. A data space oriented multi-source private domain data retrieval system, characterized in that, The system comprises a client, a global retrieval agent, and a plurality of federated private domain data sources in a data space in communication connection with the global retrieval agent, wherein: Each of the private domain data sources is configured to store private domain data of a data owner and build an index to provide local retrieval service and local access service of the private domain data; The client is configured to receive a user query string, request a retrieval result list from the global retrieval agent according to the user query string, request corresponding private domain data from the local access service of one or more private domain data sources to which the user query string is directed according to the retrieval result list, and generate a query result corresponding to the user query string according to the obtained private domain data, wherein the retrieval result list comprises data identifiers of private domain data in the one or more private domain data sources to which the user query string is directed and which is semantically related to the user query string; The global retrieval agent is configured to request a retrieval result list from the local retrieval service of one or more private domain data sources to which the user query string is directed, and return the integrated retrieval result list to the client to provide federated retrieval service of multi-source private domain data; The client is configured to generate the query result corresponding to the user query string by the following steps: Organize the obtained private domain data into a context related to the user query string, and fill in a pre-defined prompt word template using the user query string and the related context to obtain a target prompt word; Input the target prompt word into a large language model to generate the query result, wherein the large language model is configured to generate the query result corresponding to the user query string through a prompt word enhanced generation process to support the client to provide retrieval enhanced generation service to the user; the client is configured with a client retriever and a client accessor, each of the private domain data sources is configured with a digital object registry and a digital object warehouse, and the global retrieval agent is configured with a global retriever and a data source list, wherein the data source list is configured to store information of the plurality of federated private domain data sources in communication connection with the global retrieval agent.
2. The system of claim 1, wherein, The client retriever is configured to, in response to receiving a user query string, encapsulate the user query string into a first retrieval request package and send it to the global retriever to request a retrieval result from the global retriever; The global retriever is configured to, in response to receiving the first retrieval request package, parse the first retrieval request package to obtain the user query string, read digital object registry addresses of one or more private domain data sources to which the user query string is directed from the data source list to generate a first address list, encapsulate the user query string into one or more second retrieval request packages, and distribute the one or more second retrieval request packages to the digital object registries of the one or more private domain data sources to which the user query string is directed according to the first address list. A digital object registry of each private domain data source is configured to, in response to receiving the second retrieval request package, parse the second retrieval request package to obtain the user query string, retrieve a digital object associated with the private domain data source according to the user query string, generate a first retrieval result list according to a digital object identifier of the retrieved digital object, encapsulate an address of a digital object repository associated with the private domain data source and the first retrieval result list into a first response package, and return the first response package to the global retriever, where the first retrieval result list includes a digital object identifier of a digital object related to the semantic of the user query string, and the digital object is obtained by encapsulating private domain data; The global retriever is configured to, in response to receiving one or more first response packages, parse the one or more first response packages to obtain one or more digital object repository addresses and first retrieval result lists, integrate the one or more digital object repository addresses and first retrieval result lists respectively to generate a second address list and a second retrieval result list, encapsulate the second retrieval result list and the second address list into a second response package, and return the second response package to the client retriever; The client retriever is configured to, in response to receiving the second response package, parse the second response package to obtain the second retrieval result list and the second address list, and send the second retrieval result list and the second address list to the client accessor; The client accessor is configured to, in response to receiving the second retrieval result list and the second address list, send a request package of a data entity of a digital object associated with the second retrieval result list to a digital object repository associated with the second address list; The digital object repository of each private domain data source is configured to, in response to receiving the request package of the data entity of the digital object, return a corresponding data entity to the client accessor.
3. The system of claim 2, wherein, The client retriever is further configured to encapsulate a preset first retrieval parameter set and the user query string into the first retrieval request package, where the first retrieval parameter set at least includes a maximum total number of requested retrieval results; The global retriever is further configured to parse the first retrieval request package to obtain the user query string and the first retrieval parameter set, determine one or more second retrieval parameter sets corresponding to one or more private domain data sources to which the user query string is directed according to the first retrieval parameter set, encapsulate the one or more second retrieval parameter sets corresponding to the one or more private domain data sources respectively and the user query string into the one or more second retrieval request packages, and the second retrieval parameter set at least includes a maximum number of requested retrieval results for a single private domain data source; The digital object registry of each private domain data source is further configured to parse the second search request package to obtain the user query string and the second search parameter set, search for digital objects associated with the private domain data source according to the user query string and the second search parameter set, and intercept each search result according to relevance to generate the first search result list. The search result includes a digital object identifier of the searched digital object.
4. The system of claim 2, wherein, The digital object warehouse of each private domain data source is further configured to encapsulate each private domain data in the private domain data source as a data entity to obtain a digital object, and embed the data entity of the digital object in a semantic space through a deep encoder to obtain a dense vector in the semantic space. The digital object includes a digital object identifier, a data entity, and metadata. The digital object warehouse of each private domain data source is further configured to store the digital object and the corresponding dense vector, and construct a first mapping relationship. The first mapping relationship includes a mapping relationship from a digital object identifier to a data entity of the digital object and a corresponding dense vector, to support a service of accessing private domain data according to a digital object identifier. The digital object registry of each private domain data source is further configured to construct a second mapping relationship and a dense vector index. After matching a user query string in a semantic space constituted by dense vectors to obtain a dense vector matched by the user query string in the dense vector index, the first search result list is generated according to the second mapping relationship. The second mapping relationship is a mapping relationship from a dense vector corresponding to a digital object to a digital object identifier corresponding to the digital object. The dense vector index includes a dense vector corresponding to each digital object associated with the digital object registry.
5. The system of claim 4, wherein, The dense vector index and the digital object warehouse are implemented by using an open source vector database framework Chroma.
6. The system of claim 2, wherein, The information of the private domain data source in the data source list at least includes identifiers and addresses of a digital object registry and a digital object warehouse in the private domain data source. The global retriever is further configured to, in a case where a request for joining a federated environment is received from any private domain data source in the data space, record identifiers and addresses of a digital object registry and a digital object warehouse carried by the request for joining the federated environment in the data source list, to indicate that the private domain data source has joined the federated environment.
7. The system of claim 2, wherein, The client accessor is further configured to provide a search enhancement generation service to a user based on a large language model. The client is implemented based on a large language model programming framework LangChain, and is communicatively connected with the large language model through an application programming interface (API). The client retriever and the client accessor are implemented by decoupling a search function and an access function of a retriever in LangChain.
8. The system of claim 1, wherein, Each private domain data source is further configured to perform the following steps: In response to receiving the request for accessing the private domain data from the client, detecting, based on a preset access control policy, an access right of a user associated with the client; In a case where it is detected that the user associated with the client has the access right, returning corresponding private domain data to the client according to the request for accessing the private domain data.
9. The system of any of claims 1-8, wherein, Each of the private domain data sources and the global retrieval agent is implemented based on a lightweight network application framework Flask.
Citation Information
Patent Citations
Privacy protection range aggregation query method of spatial data federation
CN115905317A
Data space-oriented digital object meta-registry system and search method
CN117931812A