Efficient management method and system for data directory in data space
By combining distributed data storage systems and large language models, the semantic description, information update and security control issues of data directory systems in cross-organizational collaboration scenarios are solved, efficient data management and unified retrieval are achieved, and data interoperability and sharing capabilities are improved.
Patent Information
- Application Number
- CN202511292565.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-11
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-09-11
AI Technical Summary
Existing data directory systems have problems in cross-organizational and cross-platform data collaboration scenarios, such as insufficient semantic description capabilities, untimely information updates, weak security and sovereignty control capabilities, and difficulty in supporting multi-organization collaboration, resulting in limited data interoperability and sharing.
A distributed data storage system is used to incrementally update metadata, build full-text indexes and vector indexes, call a large language model to extract semantic entity lists and vectors, determine retrieval response results based on a list of semantically equivalent fields, and obtain data directory related information through multiple channels for centralized storage and management.
It achieves the real-time and completeness of data, improves retrieval efficiency and accuracy, deepens the understanding of data, provides more comprehensive and valuable information, and enhances user experience.
Smart Images

Figure CN120821725A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data management, and in particular to an efficient management method and system for data directories in a data space. Background Art
[0002] With the rapid development of cutting-edge technologies such as big data, cloud computing, artificial intelligence, and the Industrial Internet, the value of data is becoming increasingly prominent, making it a crucial strategic resource. To fully tap the potential of data and enable data sharing, collaboration, and value unlocked across organizations, the innovative concept of "DataSpace" has emerged in recent years and has gradually gained widespread attention. DataSpaces emphasize distributed data ownership, collaborative governance, semantic interoperability, and data sovereignty protection, demonstrating significant application value in numerous key scenarios, such as government collaboration, industrial chain data interoperability, and medical research data sharing. Within the overall DataSpace architecture, the Data Catalog system plays a fundamental and crucial role. Acting as a precise "data resource map," it performs crucial tasks such as metadata extraction, semantic modeling, quality assessment, and categorization of accessed data. It also provides data consumers with convenient access to search, browsing, application, and analysis, enabling efficient exploration and utilization of data resources. Through the Data Catalog system, organizations gain a clearer understanding of their own and others' data assets, providing strong support for data-driven decision-making and promoting cross-organizational data collaboration and innovation. As data space application scenarios continue to expand and deepen, the need for efficient management of data catalog systems is becoming increasingly urgent. A scientific and efficient data catalog management method and system can not only improve data availability and usability, but also further promote the healthy development of the data space ecosystem and inject strong impetus into the digital transformation of various industries.
[0003] However, existing data catalog systems suffer from numerous limitations. Regarding semantic description, inconsistent data field naming across multiple organizations and systems, coupled with significant metadata omissions, makes it difficult for data catalogs to accurately understand and match data from diverse sources. This lack of unified semantic description capabilities significantly hinders data interoperability and integration. Regarding information updates, as data sources continuously and dynamically change, key data such as field structure, data volume, and interface information can easily become invalid. However, existing systems lack effective incremental update mechanisms, resulting in untimely or inaccurate catalog updates, which in turn prevents them from reflecting the latest data status and hinders their effective utilization. Regarding security and sovereignty control, data catalogs in data space scenarios must not only describe data but also be tightly integrated with data access permissions, data sovereignty declarations, and authorization agreements. However, existing systems generally lack these capabilities, making it difficult to meet the stringent data security and sovereignty protection requirements of data spaces. Furthermore, most existing data catalog platforms utilize a centralized, single-tenant architecture, which struggles to support collaborative data governance across multiple organizations and lacks the ability to connect heterogeneous systems, limiting widespread data sharing and collaboration within data spaces.
[0004] To sum up, there is an urgent need for a new data catalog management method and system to meet the needs of efficient construction and sustainable operation of data space, realize the automation, intelligent and standardized description, classification and governance of data resources, and have semantic connectivity, sovereign control, cross-domain indexing and a programmable publishing mechanism.
[0005] Therefore, the present invention proposes an efficient management method and system for data directories in a data space. Summary of the Invention
[0006] The present invention provides an efficient management method and system for data directories in a data space, which is used to unify, describe, catalog, index and retrieve multi-source heterogeneous data resources in cross-organizational and cross-platform data collaboration scenarios, thereby solving at least one of the above-mentioned defects.
[0007] The present invention provides an efficient management method for a data directory in a data space, comprising: Perform incremental metadata updates on metadata information related to data directories, data product registration information, and cross-node data directory information obtained through multiple channels, and centrally store them in a distributed data storage system; Construct a full-text index and vector index for metadata in a distributed data storage system. Based on the search mode corresponding to the request type of the user's search request, match the vector generated based on the request content with the full-text index or vector index to obtain a search response result. Call the large language model to extract the semantic entity list and semantic vector of each dataset, mark semantically equivalent datasets based on the semantic vectors of all datasets, build a dataset node graph based on all semantic entity lists, and output a list of semantically equivalent fields based on the semantically equivalent datasets and dataset node graph; Determine the associated recommendation results of the search response results based on the list of semantically equivalent fields, and return the search response results and the associated recommendation results synchronously.
[0008] Preferably, the distributed data storage system includes a distributed storage subsystem compatible with the object storage protocol, a distributed relational data subsystem and a vector index subsystem. The distributed storage subsystem compatible with the object storage protocol is used to store unstructured data and serves as the underlying storage support for the distributed relational data subsystem and the vector index subsystem.
[0009] Preferably, metadata information related to the data directory obtained through multiple channels, data product registration information, and cross-node data directory information are incrementally updated and centrally stored in the distributed data storage system, including: Receive manually registered metadata information through the data connector's directory upload interface. At the same time, as a back-end reverse proxy for the data connector, directly access the data provider's application system to extract metadata and generate data product registration information based on the data product registration template. When the metadata update service is deployed near a regional functional node or a global functional node, cross-node data directory information is retrieved through the functional node's query service; Encode metadata information, data product registration information, and cross-node data directory information into metadata operation logs and write them into the distributed data storage system to complete incremental metadata updates.
[0010] Preferably, a full-text index and a vector index of metadata in a distributed data storage system are constructed, and based on a search mode corresponding to a request type of a user search request, a vector generated based on the request content is matched with the full-text index or the vector index to obtain a search response result, including: Build full-text indexes and vector indexes based on metadata in distributed data storage systems; If the user's search request is a keyword or semi-structured keyword request, the full-text index matching process prioritizes precise matching based on the keywords in the user's search request to generate a first search response result. Simultaneously, the user's search request is converted into a basic semantic vector using a large language model, and semantically relevant datasets of the basic semantic vector are matched in the vector index to generate a second search response result. If the user's search request is a semantic demand request, the high-precision semantic vector conversion based on the user's search request is prioritized during the vector index matching process to generate a second search response result. At the same time, the core keywords in the user's search request are extracted and matched against a dataset containing the core keywords in the full-text index to generate a first search response result. The retrieval response result includes a first retrieval response result and a second retrieval response result that are generated synchronously.
[0011] Preferably, a large language model is called to extract a semantic entity list and semantic vectors of each data set, semantically equivalent data sets are marked based on the semantic vectors of all data sets, a data set node graph is constructed based on all semantic entity lists, and a semantically equivalent field list is output based on the semantically equivalent data sets and the data set node graph, including: Call the large language model to extract the semantic entity list and semantic vector of each data set; Calculate the cosine similarity of the semantic vectors of any two data sets, and mark any two data sets with similarity greater than a preset similarity threshold as semantically equivalent data sets; Analyze the field names and field annotations of structured data sets and extract semantic entities to add to the semantic entity list; Construct a dataset node graph with the dataset as the vertex and the semantic entity list as the condition for the intersection; Verify the semantic consistency of fields for semantically equivalent datasets and datasets whose correlation in the dataset node graph is greater than a preset correlation threshold and output a list of semantically equivalent fields.
[0012] Preferably, determining the associated recommendation results of the search response results based on the semantically equivalent field list, and synchronously returning the search response results and the associated recommendation results, includes: Provides list category browsing and graph node navigation based on the dataset node graph; The first search response result and the second search response result in the search response result are fused based on the result fusion algorithm to obtain the final search result; Generate related recommendation results based on the correlation between the node corresponding to the final search result and the remaining nodes in the dataset node graph and the list of semantically equivalent fields; The final search results and related recommendation results are returned synchronously.
[0013] Preferably, the specific execution method of the result fusion algorithm includes: Calculate the inverse ranking of each data set in the first search response result and the inverse ranking of each data set in the second search response result respectively; Determining weights of the first search response result and the second search response result; Perform a weighted operation on the two inverse rankings based on the weights of the first search response result and the second search response result to obtain a fusion score of the corresponding data set; The top-k datasets are selected in descending order of fusion scores as the final retrieval results.
[0014] Preferably, determining the weights of the first search response result and the second search response result includes: Determine the current application scenario of the data space and define multiple types of search demand indicators for the current application scenario, including accuracy indicators, comprehensiveness indicators, and timeliness indicators; Based on the historical search logs in the latest historical period of the current application scenario, determine the actual values of the accuracy index, comprehensiveness index, timeliness index and user satisfaction of each historical search result; Taking the actual values of the accuracy, comprehensiveness, and timeliness indicators of each historical search result as independent variables and user satisfaction as the dependent variable, a multivariate linear regression model is used to determine the linear relationship between multiple types of search demand indicators and user satisfaction in the current application scenario. Based on the linear relationship between multiple retrieval demand indicators and user satisfaction in the current application scenario, the optimal weight combination of multiple retrieval demand indicators in the current application scenario is determined; Based on the optimal weight combination of multiple retrieval demand indicators in the current application scenario, an indicator weight matrix is constructed. Based on the correlation coefficients between multiple retrieval demand indicators and full-text indexes and vector indexes, an indicator-index correlation matrix is constructed. Based on the constraint factors of full-text indexes and vector indexes in the current application scenario, a scenario constraint matrix is constructed. The result obtained by multiplying the indicator weight matrix, the indicator-index association matrix and the scenario constraint matrix in sequence is used as the optimized index weight matrix. Based on the dual index weights and the normalization operation of all elements in the optimized index weight matrix, the final index weight matrix is obtained, and the weights of the first retrieval response result and the second retrieval response result are determined based on the elements in the final index weight matrix.
[0015] Preferably, based on the historical search logs in the latest historical period under the current application scenario, the actual values of the accuracy index, comprehensiveness index, timeliness index and user satisfaction of each historical search result are determined, including: Based on the historical search logs in the latest historical period under the current application scenario, determine the ratio of the number of results containing core keywords or semantic entities in each historical search result to the total number of results as the search result matching degree, and determine the ratio of the number of results with a click matching degree not less than a preset matching degree threshold to the total number of clicks in each historical search result as the precise click-through rate, and determine the number of business dimensions covered by the results returned by a single search request in each historical search result as the number of associated dimensions, and determine the ratio of the number of requests using associated results to the total number of requests in each historical search result as the associated usage rate, and determine the proportion of data in the search results in each historical search result whose update time does not exceed the preset cycle time of the current application scenario as the freshness, and determine the ratio of the number of clicks on results with a freshness not less than the preset freshness threshold to the total number of clicks in each historical search result as the real-time click-through rate; The actual value of the precision index is determined based on the search result matching degree and the accurate click rate of each historical search result, the actual value of the comprehensiveness index is determined based on the number of associated dimensions and the associated usage rate of each historical search result, and the actual value of the timeliness index is determined based on the freshness and real-time click rate of each historical search result; The user satisfaction of each historical search result is calculated based on the click rate, collection rate and usage rate of each historical search result in the historical search log within the latest historical period in the current application scenario.
[0016] The present invention provides an efficient management system for data directories in a data space, comprising: The metadata collection and update module is used to incrementally update metadata information related to the data directory obtained through multiple channels, data product registration information, and cross-node data directory information, and centrally store them in the distributed data storage system; The index construction and retrieval module is used to build full-text indexes and vector indexes of metadata in the distributed data storage system. Based on the retrieval mode corresponding to the request type of the user's retrieval request, the vector generated based on the request content is matched with the full-text index or vector index to obtain the retrieval response result; The consistency analysis and fusion module is used to call the large language model to extract the semantic entity list and semantic vector of each dataset, mark semantically equivalent datasets based on the semantic vectors of all datasets, build a dataset node graph based on all semantic entity lists, and output a list of semantically equivalent fields based on the semantically equivalent datasets and dataset node graph; The associated recommendation and output module is used to determine the associated recommendation results of the search response results based on the list of semantically equivalent fields, and synchronously return the search response results and the associated recommendation results.
[0017] The beneficial effects of the present invention compared to the prior art are: incremental metadata updates are performed on metadata information related to the data directory, data product registration information, and cross-node data directory information obtained through multiple channels, and stored centrally in a distributed data storage system, ensuring the real-time and integrity of the data, and facilitating unified management and call of data. A full-text index and vector index of metadata in a distributed data storage system are constructed, and according to the retrieval mode corresponding to the user's retrieval request type, the vector generated based on the request content is matched with the corresponding index to obtain the retrieval response result, thereby improving the retrieval efficiency and accuracy, and being able to quickly meet the user's diverse retrieval needs. A large language model is called to extract the semantic entity list and semantic vector of each data set, and semantically equivalent data sets are marked accordingly, a data set node graph is constructed, and a list of semantically equivalent fields is output, thereby mining the semantic associations between data and deepening the understanding of the data. Based on the list of semantically equivalent fields, the associated recommendation results of the retrieval response results are determined and returned synchronously, providing users with more comprehensive and valuable information, further improving the user experience, and helping users to more deeply explore and utilize the data directory in the data space.
[0018] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present invention. The purpose and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in this application document.
[0019] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings: Figure 1 Flowchart of a method for efficiently managing a data directory in a data space according to an embodiment of the present invention; Figure 2 This is a flowchart of metadata consistency analysis in an embodiment of the present invention; Figure 3 is a semantic network diagram of a data set in an embodiment of the present invention; Figure 4 This is a flowchart of intelligent semantic retrieval and recommendation in an embodiment of the present invention; Figure 5 This is a system description diagram in an embodiment of the present invention. DETAILED DESCRIPTION
[0021] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0022] like Figure 1 As shown, the present invention provides an efficient management method for a data directory in a data space, comprising: Perform incremental metadata updates on metadata information related to data directories, data product registration information, and cross-node data directory information obtained through multiple channels, and centrally store them in a distributed data storage system; Construct a full-text index and vector index for metadata in a distributed data storage system. Based on the search mode corresponding to the request type of the user's search request, match the vector generated based on the request content with the full-text index or vector index to obtain a search response result. Call the large language model to extract the semantic entity list and semantic vector of each dataset, mark semantically equivalent datasets based on the semantic vectors of all datasets, build a dataset node graph based on all semantic entity lists, and output a list of semantically equivalent fields based on the semantically equivalent datasets and dataset node graph; Determine the associated recommendation results of the search response results based on the list of semantically equivalent fields, and return the search response results and the associated recommendation results synchronously.
[0023] In this embodiment, metadata related to the data catalog includes information about the data's industry, source, product interaction method, data type, format, and structured data field definitions. Data product registration information is generated based on the data product registration template and metadata extracted from the data provider's application system. It includes information related to data product transactions. Cross-node data catalog information refers to data catalog information between different regions or global functional nodes within the data space architecture.
[0024] In this embodiment, the user retrieval request is a data query demand issued by the user to the data directory management system. For example, if the user wants to obtain sales data of a specific time and place, he or she will submit such a request to the system.
[0025] In this embodiment, the request content is the specific demand description in the user's search request, for example, "query the clothing sales in Beijing in the first quarter of 2024" is the request content.
[0026] In this embodiment, the vector generated based on the request content is converted into a vector using a large language model. For example, the query "Query clothing sales in Beijing in the first quarter of 2024" is converted into a vector containing semantic features using a large language model. The values in the vector represent the feature values of different semantic dimensions.
[0027] In this embodiment, the search response result is the result returned to the user by the system after matching the full-text index with the vector index according to the user's search request.
[0028] In this embodiment, the large language model is a model with powerful language understanding and generation capabilities. In this embodiment, it is used to convert user search requests into semantic vectors, extract semantic entity lists and semantic vectors of data sets, etc. Models such as ChatGPT can understand natural language and perform semantic analysis and conversion.
[0029] In this embodiment, a data set is a collection of data. In the data directory management system, each data set has a specific metadata description. For example, an employee information data set includes employee name, age, position and other data.
[0030] In this embodiment, the semantic entity list is a list of meaningful entities extracted from the dataset name, description, and structured data fields using a large language model. For example, from the "employee salary table (employee number, salary, bonus)" table, semantic entities such as "employee," "salary," and "bonus" are extracted to form a list. A semantic vector is a vector generated by converting the dataset name and description using a large language model. It is used to represent the semantic features of the dataset, and the vector value reflects the semantic characteristics of the dataset in different dimensions.
[0031] In this embodiment, calling the large language model to extract the semantic entity list and semantic vector of each data set is to use the large language model to analyze the name and description information of each data set, output the semantic entity list, and convert the data set name and description into a semantic vector.
[0032] In this embodiment, the data directory plus the corresponding description information and tags constitute the overall description of the corresponding data, which is metadata. Figure 5 As shown in the figure, the data catalog management system consists of the following five blocks: data storage service; metadata update service; metadata indexing and query service; data catalog browsing service; metadata consistency analysis and fusion service. Among them: Data storage services: This system uses a distributed system to store metadata, index data, and metadata update logs. Through abstract storage interfaces, it can flexibly adapt to distributed key-value storage systems and relational database systems, providing flexible scalability and high-performance read and write support. Currently, we use the S3-compatible distributed key-value storage system, the distributed relational data system Apache Doris, and Milvus for vector indexing. The S3-compatible distributed key-value storage system can not only directly store unstructured data such as metadata update logs and indexes, but can also serve as the underlying file system for Apache Doris and Milvus.
[0033] Metadata update service: According to the latest national data infrastructure settings / data space standards, data is first connected to the data space through the data connector from the data provider's business system / or data sharing service system. Figure 5In connection 2, the metadata service accepts metadata information about data resources / products manually registered by users through the data connector's data catalog upload interface. This includes information about the industry, data source, product interaction method, data type, format (structured data: field definition), and other transaction-related information. The metadata service encodes this information into metadata operation logs and writes them to the underlying data storage system through the storage service.
[0034] In addition to the above push-based streaming update solution, the metadata update service can also serve as a backend reverse proxy for data connectors, such as Figure 5 Connection 3 in the diagram directly accesses the data provider's application system, such as accessing the backend relational database system through JDBC, or accessing the data provider's Hive metadata service through the HiveMetaStore service. In this way, metadata from the service provider's business system can be directly extracted, and data product registration information can be generated based on the data product registration template. Figure 5 Line 4 in the data connector is written to the data connector, which greatly reduces the workload when registering data products.
[0035] In the national data infrastructure / data space standards, the data connector's data directory upload interface requires that it upload its own data directory information to the regional function node, which in turn uploads it to the global function node. Figure 5 When the metadata update service is deployed near the above two nodes, it can retrieve the data directory information through the query service of the above functional node and then submit it to the data storage service for storage.
[0036] Metadata indexing and query services: There are two types of metadata indexing: one is full-text indexing, which is supported by the built-in full-text indexing function of Apache Doris; the second is vector indexing based on large language models (currently supporting DeepSeek and ChatGPT), which is supported by Milvus.
[0037] The search service supports: a) simple keyword searches, such as "Weather cloud map for the Beijing area." b) semi-structured keyword searches, such as "Region: Beijing; Data type: Weather cloud map." c) intelligent semantic search and recommendations. When a user submits an analysis or modeling request, the search service returns relevant datasets based on the semantics of the request.
[0038] Data catalog browsing service: Although the descriptive information of various data sources in the data space is stored in a table format, in reality, data is actually interconnected in a networked manner. For example, personal information from the public security department can be associated with medical treatment information, consumer information from the commercial department, and traffic information from the transportation department. Traffic information can be associated with vehicle information, as well as various accessories, manufacturers, and other related information. Based on this understanding, the data catalog description information obtained from data connectors, regional functional nodes, and global functional nodes will be semantically analyzed using a large model and assigned different semantic entity labels, such as category, region, vehicle, medical, and so on. Different data sets with the same semantic labels can be associated with each other.
[0039] For structured data, we use the subsequent metadata consistency analysis service to identify fields with the same semantics but different names, or fields with the same name but different semantics. We then establish associations based on the fields with the same semantics in the dataset.
[0040] Through the above two levels of semantic analysis, the data sets in the data space form a graph with multiple connection relationships. Based on the above processing results, the data directory browsing service provides list-based classification browsing and graph-based node navigation starting from a certain node, such as Figure 3 As shown in the dataset semantic network diagram, nodes with the same color represent semantically equivalent datasets, and other nodes represent semantically related data. Users can navigate from one node to other related nodes on this diagram by following the node associations.
[0041] like Figure 2 As shown in the figure, the steps of metadata consistency analysis and fusion service are as follows: S101: Use a large language model to extract the semantic information of the dataset name and dataset description, and output: semantic entity list (entity_list); Dataset name + semantic vector corresponding to the dataset description (dataset_semantic_vector); S102: Calculate the cosine similarity of all dataset_semantic_vectors. Datasets with a similarity greater than a specified threshold (e.g., 0.95) are considered equivalent datasets, describing the same information from different sources. Examples include weather cloud map data from different provinces, or traffic road snapshot data, or road mapping maps published by different organizations.
[0042] S103: For structured datasets, analyze their field annotations and names to extract their semantic entities. For example, in an order table (order number, product number, shipping method, delivery date), the output semantic entities are: order and product. These semantic entities are added to the corresponding entity list (entity_list) extracted from the dataset name and description.
[0043] S104: Calculate the union of the semantic entity lists on each data set, and finally construct a graph with one data set as a vertex. If the vertices corresponding to two vertices have an intersection, each identical element is converted into a table connecting the two vertices.
[0044] S105: For each node pair whose correlation is greater than a specified threshold, add the node pair to the semantically related vertex pair list related_pair_list.
[0045] S106: For semantically equivalent structured data sets, if the field names are the same, they can be considered to represent the same information (if there are field descriptions, the semantics of the field descriptions can be further compared to determine whether they are completely consistent; if not, a sampling method can be used to compare the field values); for fields with similar field names, such as "username, user_name", "user_id, userId", their field annotations or field values are compared to determine whether they have the same semantics.
[0046] For related datasets A and B, the field semantics of their connected nodes are compared based on their association on the node graph, and the comparison method is similar to the method in the above steps.
[0047] In summary, the efficient management method for data directories in this data space achieves the following technical effects: Consistency: Through metadata consistency analysis and integration, a unified semantic model view is provided to users. At the dataset level, datasets with different names are given the same semantic name. For example, vehicle capture information from traffic control departments may be named differently in different provinces, with some named "vehicle capture" and others "road capture." Through semantic analysis of data descriptions, this is unified as "traffic control vehicle capture." At the field level, different fields may have different names. For example, different ID card information may be named "user_id" in some datasets and "userID" in others. After consistency analysis, this is unified as "user_id_code." After unified field naming and semantic equivalence analysis, different datasets can be linked. For example, hotel accommodation registration information and public security personnel information can be linked using ID card identifiers. This provides a more complete dataset for subsequent data analysis, which in turn generates more analytical needs and unlocks more data value.
[0048] Immediacy: When data sources change, the metadata update service instantly detects the changes and automatically or semi-automatically updates the data catalog information on the data connector and the relevant metadata items in the metadata storage system. This prevents users from querying outdated information, resulting in data and catalog information mismatches and, in turn, the failure of data transaction contracts.
[0049] Flexibility: The service-based architecture allows the system to be flexibly deployed alongside different components in different data spaces, serving as ancillary public support services. Based on standardized data infrastructure interfaces, the data catalog management system can adapt to various heterogeneous data sources. A flexible user interface supports everything from simple keyword searches to complex semantic natural language queries, as well as intuitive data graph-based browsing and navigation. This improves the efficiency and convenience of user access to the data catalog.
[0050] To enable a distributed data storage system to effectively store various types of data and provide underlying support for other components, a distributed data storage system is proposed, consisting of a distributed storage subsystem compatible with the object storage protocol, a distributed relational data subsystem, and a vector indexing subsystem. The object storage subsystem plays a crucial role, capable of storing unstructured data such as metadata update logs and indexes. It also provides underlying storage support for the distributed relational data subsystem and the vector indexing subsystem. These subsystems rely on this subsystem for stable operation. In practical applications, distributed key-value storage systems like S3-compatible distributed key-value storage systems are examples of this object storage subsystem. These subsystems provide the underlying storage support for the distributed relational data system Apache Doris and the vector indexing system Milvus.
[0051] In order to comprehensively and timely update and centrally store metadata, product registration, and cross-node information related to data catalogs obtained through multiple channels, it is proposed to incrementally update metadata information related to data catalogs, data product registration information, and cross-node data catalog information obtained through multiple channels, and centrally store them in a distributed data storage system, including: Receiving manually registered metadata information through the directory upload interface of the data connector means that the data provider or user can manually input data-related metadata information, such as industry, data source, product interaction form, data type, format (for structured data, also including field definition) and other transaction-related information, into the system through the directory upload interface specially set up by the data connector, providing basic information for subsequent data management and processing. At the same time, as the back-end reverse proxy of the data connector, it directly accesses the data provider's application system to extract metadata and generates data product registration information according to the data product registration template. This means that after obtaining the metadata extracted from the data provider's application system, these metadata are sorted and filled in according to the pre-established data product registration template to generate data product registration information containing content related to data product transactions, such as data product specifications, usage restrictions, prices and other related registration information.
[0052] When the metadata update service is deployed near a regional or global functional node, it retrieves cross-node data directory information through the functional node's query service. This means that in the data space architecture, when the metadata update service is located in a specific location, it can use the query service provided by the regional or global functional node to obtain cross-node data directory information shared between different nodes.
[0053] Metadata information, data product registration information, and cross-node data directory information are encoded as metadata operation logs and written to the distributed data storage system to complete incremental metadata updates. This involves converting the collected information into metadata operation logs according to specific encoding rules and then storing them in the distributed data storage system.
[0054] In order to efficiently and accurately obtain search response results using full-text indexes and vector indexes based on different types of user search requests, this paper proposes to build full-text indexes and vector indexes for metadata in a distributed data storage system. Based on the search mode corresponding to the request type of the user search request, the vector generated based on the request content is matched with the full-text index or vector index to obtain the search response results, including: Build full-text indexes and vector indexes based on metadata in distributed data storage systems; If the user's search request is a keyword or semi-structured keyword request, the full-text index matching process prioritizes precise matching based on the keywords in the user's search request to generate a first search response result. Simultaneously, the user's search request is converted into a basic semantic vector using a large language model, and semantically relevant datasets of the basic semantic vector are matched in the vector index to generate a second search response result. If the user's search request is a semantic demand request, the high-precision semantic vector conversion based on the user's search request is prioritized during the vector index matching process to generate a second search response result. At the same time, the core keywords in the user's search request are extracted and matched against a dataset containing the core keywords in the full-text index to generate a first search response result. The retrieval response result includes a first retrieval response result and a second retrieval response result that are generated synchronously.
[0055] In this embodiment, full-text indexing utilizes the built-in functionality of Apache Doris to comprehensively index the text content in metadata, facilitating rapid location of data containing specific keywords. Vector indexing utilizes Milvus to convert metadata into vector form, enabling semantic-based retrieval.
[0056] In this embodiment, a keyword or semi-structured keyword request is a type of user search request. A keyword request refers to a user directly entering a simple keyword to search, such as "sales data." A semi-structured keyword request refers to a user entering keywords in a specific structured format, such as "Region: Beijing; Data Type: Sales Data."
[0057] In this embodiment, the full-text index matching process prioritizes precise matching of keywords in the user's search request to generate the first search response result. For example, if a user enters the keyword "electronic product sales," the full-text index will search the metadata of the distributed data storage system for data records that precisely contain these keywords and present the matching data to the user as the first search response result.
[0058] In this embodiment, the large language model converts the user's search request into a basic semantic vector. The vector index then matches the semantically relevant datasets for the basic semantic vectors to generate a second search response. For example, if a user requests "analyze mobile phone sales trends," the large language model converts this request into a basic semantic vector. The vector index then uses semantic similarity to find datasets related to the semantics, such as historical mobile phone sales data or market share data, as the second search response.
[0059] In this example, semantic demand requests are another type of user search request. This type of request focuses more on expressing complex semantic needs, proposing analysis and modeling requirements in natural language, such as "Predict changes in the new energy vehicle market share over the next year."
[0060] In this embodiment, the vector index matching process prioritizes precise matching based on the high-precision semantic vectors converted from the user's search request to generate a second search response result. When a user makes a semantic request, the large language model converts it into a high-precision semantic vector. Based on this high-precision vector, the vector index accurately matches highly semantically relevant datasets in the dataset. For example, for a request to "predict changes in the market share of new energy vehicles in the next year," the vector index can find relevant datasets such as new energy vehicle production, sales, and market competition trends as the second search response result.
[0061] When generating high-precision semantic vectors, a three-step process is performed: data space domain vocabulary construction - model fine-tuning - vector optimization: Construct a dynamically updated "data space domain vocabulary", which is divided into a three-level structure: core layer, extension layer, and association layer: the core layer contains basic data space entities, such as cross-node directories, data sovereignty statements, authorization agreement identifiers, and metadata operation logs. Each entity is annotated with an industry benchmark weight, such as a data sovereignty statement weight of 0.9 in government scenarios and a patient privacy identifier weight of 0.95 in medical scenarios; the extension layer contains derived concepts of entities, such as data sovereignty statement derivatives: cross-border data sovereignty and hierarchical sovereignty management, with weights of 0.6-0.8 times the weight of the core layer entity; the association layer contains business terms associated with entities, such as "authorization agreement identifier" associated with "data usage period" and "access permission granularity", with weights of 0.3-0.5 times the weight of the core layer entity; the vocabulary is automatically updated every 7 days based on the frequency of new metadata entities in the data space, and extension layer / association layer terms with an appearance frequency below the threshold (e.g., <5 times per week) are removed; Based on the domain vocabulary, the large language model is incrementally fine-tuned, and the "domain sample generation-contrastive learning training" strategy is adopted: first, 100,000+ semantic demand requests are extracted from the historical retrieval log of the data space, and each request is annotated with 3-5 core entities in the domain vocabulary as labels; then a "domain entity masking task" is constructed to randomly mask the domain entities in the input text, allowing the model to predict the masked entities. At the same time, contrastive learning is introduced, and texts containing domain entities and synonymous texts without domain entities are used as positive and negative samples to optimize the model's ability to capture the semantics of domain entities; a dynamic learning rate attenuation strategy is adopted during the fine-tuning process, wherein the initial learning rate is 5e -5 , decaying by 10% every 2000 steps, ensuring that the model retains general semantic understanding capabilities while strengthening semantic sensitivity in the data space domain; The initial semantic vector output by the fine-tuned model is optimized using attention weighting. The weight of each domain entity in the vocabulary in the user's search request is calculated, and the weight is mapped to an adjustment coefficient of the vector dimension through the attention mechanism. The feature dimensions of the corresponding domain entities in the initial semantic vector are weighted and enhanced. For example, the value of the dimension corresponding to cross-border data sovereignty is multiplied by 1.2-1.5 times, and finally a high-precision semantic vector is generated. After testing, the retrieval recall rate of this vector in the data space scenario is improved compared with the general semantic vector, and the mismatch rate is reduced.
[0062] In this embodiment, the core keywords in the user's search request are extracted and matched against datasets containing the core keywords in the full-text index to generate a first search response result. For example, for a request to "predict changes in the market share of new energy vehicles over the next year," the system extracts core keywords such as "new energy vehicles" and "market share," and searches the full-text index for datasets containing these core keywords, such as new energy vehicle market research reports and relevant policy documents, as the first search response result.
[0063] like Figure 2 As shown in the figure, in order to deeply explore the semantic relationship between datasets, output a list of semantically equivalent fields, and provide a basis for data association analysis, it is proposed to call a large language model to extract the semantic entity list and semantic vector of each dataset, mark semantically equivalent datasets based on the semantic vectors of all datasets, build a dataset node graph based on all semantic entity lists, and output a list of semantically equivalent fields based on the semantically equivalent datasets and dataset node graph, including: Call the large language model to extract the semantic entity list and semantic vector of each data set; Calculate the cosine similarity of the semantic vectors of any two data sets, and mark any two data sets with similarity greater than a preset similarity threshold as semantically equivalent data sets; Analyzing the field names and field annotations of a structured dataset and extracting semantic entities to add to the semantic entity list is an operation that deeply explores the semantic information of the dataset. For a structured dataset, such as an employee information table, field names may include "employee number," "name," and "department," and field annotations may further explain these names. Meaningful semantic entities such as "employee" and "department" are extracted from these field names and annotations and added to the semantic entity list previously extracted from the dataset name and description using a large language model. This enriches the semantic entity list and more comprehensively reflects the semantics of the dataset.
[0064] Constructing a dataset node graph with datasets as vertices and semantic entity lists having intersections as conditions is to construct a graphical structure that can intuitively display the semantic associations between datasets. Consider each dataset as a vertex. If the semantic entity lists of two datasets have the same elements, that is, there is an intersection, these same elements are used as connections to establish a line between the two vertices. For example, the semantic entity list of dataset A includes "product" and "sales", and the semantic entity list of dataset B includes "sales" and "market". Since both have the intersection of "sales", a line is established between the corresponding vertices of datasets A and B. Many datasets are constructed into graphs in this way; Verifying field semantic consistency and outputting a list of semantically equivalent fields for semantically equivalent datasets and datasets with correlations greater than a preset correlation threshold within the dataset node graph is done to identify semantically identical fields across datasets. For semantically equivalent datasets and datasets with high correlations (greater than a preset correlation threshold), detailed field analysis is performed. For example, for two semantically equivalent employee information datasets, one with the "Employee ID" field and the other with the "Employee Number" field, semantic consistency is verified by comparing field annotations and values. Semantically consistent fields are organized into a list and output as the semantically equivalent field list.
[0065] In this embodiment, the cosine similarity of the semantic vectors of any two data sets is calculated in order to measure the similarity between the two data sets at the semantic level.
[0066] In this embodiment, the preset similarity threshold is a pre-set standard value used to determine whether two data sets are semantically equivalent. For example, the preset similarity threshold is set to 0.95.
[0067] In this embodiment, the preset correlation threshold is a preset numerical standard, for example, set to 0.8.
[0068] like Figure 4 As shown in the figure, in order to provide related recommendations for retrieval response results based on semantic analysis results and improve the comprehensiveness and practicality of retrieval services, it is proposed to determine the related recommendation results of retrieval response results based on a list of semantically equivalent fields, and synchronously return the retrieval response results and related recommendation results, including: Providing list category browsing and graph node navigation based on the dataset node graph is to use the constructed dataset node graph to provide users with a convenient data browsing method. List category browsing is to classify the datasets according to the characteristics or associations of different datasets in the dataset node graph, and present them to the user in the form of a list. For example, according to the industry to which the data belongs, all datasets are divided into categories such as medical, financial, and education. Users can quickly browse the datasets under this category by clicking on different categories. Graph node navigation allows users to start from a certain node in the dataset (i.e., a certain dataset) and navigate to other related dataset nodes along the lines between the nodes in the node graph (representing semantic associations). For example, starting from the "Medical Patient Information Dataset" node, users can navigate to related nodes such as the "Medical Diagnosis Result Dataset" through the associations between nodes; The first search response result and the second search response result in the search response result are fused based on the result fusion algorithm to obtain the final search result; Generating associated recommendation results based on the correlation between the node corresponding to the final search result and the remaining nodes in the dataset node graph and the list of semantically equivalent fields is the process of providing relevant data recommendations to users. When the system gives the final search results, the datasets corresponding to these results are represented as corresponding nodes in the dataset node graph. The correlation between these nodes and the remaining nodes in the graph is calculated. The higher the correlation, the closer the semantic correlation between the dataset corresponding to the remaining node and the dataset of the final search result. At the same time, combined with the list of semantically equivalent fields, the semantically equivalent fields in these related datasets are analyzed to determine which datasets to recommend to users. For example, if the final search result is a dataset related to "electronic product sales data", by analyzing the node graph correlation and the list of semantically equivalent fields, it is found that the "electronic product production data" dataset has a high correlation with it and there are semantically equivalent fields. The "electronic product production data" dataset is then provided to the user as an associated recommendation result. The final search results and related recommendation results are returned synchronously.
[0069] like Figure 4 As shown in the figure, in order to scientifically and rationally integrate the results generated by different search paths and obtain higher-quality final search results, a specific implementation method of the result fusion algorithm is proposed, including: Calculate the last ranking of each dataset in the first search response result and the last ranking in the second search response result respectively; assuming that there are 10 datasets in the first search response result, sort them from high to low according to rules such as relevance. If a dataset ranks 3rd, its last ranking is 1 / 3; similarly, in the second search response result, if the dataset ranks 5th, its last ranking is 1 / 5.
[0070] Determine the weights of the first search response result and the second search response result; for example, after a series of analyses, determine that the weight of the first search response result is 0.6, and the weight of the second search response result is 0.4.
[0071] Perform a weighted operation on the two inverse rankings based on the weights of the first search response result and the second search response result to obtain a fusion score of the corresponding data set; The top k datasets are selected as the final search results, sorted by fusion score from high to low. This approach involves filtering out the datasets that best meet user needs. For example, if k is set to 5, all datasets are sorted by fusion score from high to low, and the top five datasets are selected as the final search results presented to the user.
[0072] In this embodiment, we take user intelligent semantic retrieval as an example, such as query Q1: "During the epidemic, analyze the spread speed and scope of the epidemic, and track possible close contacts." The specific processing flow is as follows: Figure 4 Intelligent semantic retrieval is shown.
[0073] S401: Process the query Q1 using the full-text search system and return the top-k results R1; S402: Using the large language model, convert the query Q1 into a semantic vector, query from the vector index database, and return the top-k query results R2; S403: Use the query result fusion algorithm “reciprocal rank fusion” to merge the query results R1 and R2 into the final top-k results R3.
[0074] S404: Using R3, based on the dataset semantic association graph, the associations between nodes in the R3 set and other nodes, and the list of semantically equivalent fields, another result, R4, is returned. Finally, result R3 is returned as the query result, and R4 is returned as the recommended association.
[0075] After the above processing, the result set R3 of Q1 is: hospital treatment information, mobile phone trajectory movement information, shopping mall and supermarket shopping information; the recommended data set R4 is: highway vehicle capture information, public security personnel information, and vehicle registration information.
[0076] In order to accurately determine the weights of the first search response result and the second search response result according to the current application scenario of the data space and optimize the fusion effect of the search results, it is proposed to determine the weights of the first search response result and the second search response result, including: Determine the current application scenario of the data space, which means to determine which specific scenario the current data retrieval operation is in within the data space, such as the government data collaboration scenario, the medical data sharing scenario, or the industrial chain data interoperability scenario. This can be done by identifying the type of organization to which the user belongs, or automatically judging from the scenario entity in the search keyword. And define multiple categories of retrieval demand indicators for the current application scenario, including: accuracy indicators, comprehensiveness indicators, and timeliness indicators; The historical search logs in the latest historical period based on the current application scenario refer to the search logs generated in the most recent specific time period in the determined current application scenario, such as the government data collaboration scenario. According to the previous settings, this historical period is generally 30 days. These logs record the specific content of the search request, the feedback of the search results, including whether it is clicked, collected, used, and the attributes of the data, such as update time, related dimensions, matching degree and other key information. Determine the actual values of the accuracy index, comprehensiveness index, timeliness index and user satisfaction of each historical search result; Using the actual values of the precision, comprehensiveness, and timeliness indicators for each historical search result as independent variables and user satisfaction as the dependent variable, a multiple linear regression model was used to determine the linear relationship between multiple search demand indicators and user satisfaction in the current application scenario. Specifically, the actual values of the precision, comprehensiveness, and timeliness indicators were extracted and calculated from the historical search logs and used as input variables. User satisfaction was calculated by weighting the "search result usage rate plus the collection rate" in the logs. In other words, satisfaction equals 0.6 times the usage rate plus 0.4 times the collection rate, and used as the output variable. Using a multiple linear regression model, satisfaction equals β1 times the precision score, β2 times the comprehensiveness score, and β3 times the timeliness score, and finally adds the error term ε (where β1, β2, and β3 are standardized regression coefficients, and the precision, comprehensiveness, and timeliness scores are the normalized values of each indicator). This method identifies the linear relationship between these search demand indicators and user satisfaction.
[0077] Based on the linear relationship between the multi-category retrieval demand indicators and user satisfaction in the current application scenario, the optimal weight combination of the multi-category retrieval demand indicators in the current application scenario is determined, that is, the weight combination when user satisfaction is the highest. After obtaining the above linear relationship, the standardized regression coefficients β1, β2, and β3 output by the model are the weights of the accuracy, comprehensiveness, and timeliness indicators, respectively, and the accuracy weight plus the comprehensiveness weight plus the timeliness weight equals 1. Under the premise of meeting this condition, the weight combination that maximizes user satisfaction is the optimal weight combination of the multi-category retrieval demand indicators in the current application scenario.
[0078] Based on the optimal weight combination of multiple retrieval demand indicators in the current application scenario, an indicator weight matrix is constructed. For example, in the government data collaboration scenario, if the optimal weight combination is an accuracy weight of 0.4, a comprehensiveness weight of 0.3, and a timeliness weight of 0.3, then the constructed indicator weight matrix W is a matrix with 1 row and 3 columns, written as W equals [0.4, 0.3, 0.3], where the rows represent the scenarios and the columns correspond to the accuracy, comprehensiveness, and timeliness indicators respectively.
[0079] Based on the correlation coefficients between multiple categories of retrieval demand indicators and the full-text index and vector index, an indicator-index correlation matrix is constructed. Element A[i][j] in the matrix represents the strength of the correlation between the i-th indicator and the j-th index. This strength ranges from 0 to 1 and is determined based on the characteristics of the full-text index and vector index. For example, the full-text index performs well in terms of precision, meaning it has strong keyword matching capabilities, but is relatively weak in terms of comprehensiveness and timeliness. The vector index, on the other hand, performs well in terms of comprehensiveness, meaning it has strong semantic association capabilities, but is relatively weak in terms of precision. The resulting indicator-index correlation matrix A might look like this: the first row indicates that the correlation strength between precision and the full-text index is 0.8, and the correlation strength with the vector index is 0.2; the second row indicates that the correlation strength between comprehensiveness and the full-text index is 0.3, and the correlation strength with the vector index is 0.7; the third row indicates that the correlation strength between timeliness and both the full-text index and the vector index is 0.5.
[0080] And based on the constraint factors of full-text index and vector index in the current application scenario, a scenario constraint matrix is constructed; the diagonal elements of this matrix are the basic constraint coefficients of the scenario on the index, the purpose of which is to ensure that the index weight is non-negative and reasonable. The non-diagonal elements are 0, indicating that there is no direct constraint relationship between the indexes. Its setting is based on the compliance requirements of the scenario. For example, in the government data collaboration scenario, in order to strengthen the constraints of full-text index and weaken the constraints of vector index, the constructed scenario constraint matrix Cgovernment is equal to .
[0081] The optimized index weight matrix is obtained by sequentially multiplying the indicator weight matrix, the indicator-index association matrix, and the scenario constraint matrix. Using matrix multiplication, the indicator weight matrix W is multiplied by the indicator-index association matrix A to obtain the initial index weight matrix S. This matrix has one row and two columns, with rows representing scenarios and columns representing indexes. The principle is to distribute the scenario weights for the indicators to the full-text index and vector index based on the strength of the correlation between the indicators and the indexes.
[0082] Based on the sum of the dual index weights, all elements in the optimized index weight matrix are normalized to obtain the final index weight matrix. To ensure that the final sum of the dual index weights is a value that conforms to the actual business situation, this sum is set to a specific value based on the search log verification. Based on this specific value, all elements in the optimized index weight matrix are normalized to obtain the final index weight matrix. For example, assuming that the specific value of the sum of the dual index weights is 2.2, the elements in the optimized index weight matrix are S'Government Affairs[0][0] equal to 0.672, S'Government Affairs[0][1] equal to 0.396, and their sum is 1.068. Then the elements in the final index weight matrix are calculated as follows: the first element in the final index weight matrix is equal to 0.672 divided by 1.068 and then multiplied by 2.2, which is approximately equal to 1.38, and is approximately 1.4; the second element is equal to 0.396 divided by 1.068 and then multiplied by 2.2, which is approximately equal to 0.82, and is approximately 0.8. In this way, the final index weight matrix is obtained.
[0083] The weights of the first and second search response results are determined based on the elements in the final index weight matrix. The elements in the final index weight matrix directly correspond to the weights of the first search response result (full-text index) and the second search response result (vector index). For example, in the final index weight matrix calculated above, the first element, 1.4, is the weight of the first search response result, and the second element, 0.8, is the weight of the second search response result.
[0084] In this embodiment, the accuracy index, comprehensiveness index, and timeliness index are important metrics used to evaluate retrieval requirements and retrieval results. The accuracy index refers to the degree of accuracy of the match between the retrieval results and the user's needs, which can be reflected by the matching degree of the retrieval results and the user's accurate click-through rate. The comprehensiveness index mainly refers to the coverage of the associated data. In the medical scenario, it is reflected in the number of business dimensions covered by the results returned by a single search request, such as medical treatment, medication, trajectory and other dimensions. It can also be measured by the usage rate of associated results. The timeliness index focuses on the freshness of the data update and the click-through rate of real-time results.
[0085] In this embodiment, the sum of the dual index weights refers to the sum of the weight of the first search response result corresponding to the full-text index and the weight of the second search response result corresponding to the vector index.
[0086] In order to accurately extract the actual values of the accuracy, comprehensiveness, and timeliness indicators of each historical search result and user satisfaction from the historical search logs and provide an effective basis for weight determination, it is proposed to determine the actual values of the accuracy, comprehensiveness, and timeliness indicators of each historical search result and user satisfaction based on the historical search logs in the latest historical period under the current application scenario, including: Based on the historical search logs within the most recent historical period in the current application scenario, the ratio of the number of results containing core keywords or semantic entities in each historical search result to the total number of results is determined as the search result match degree. This means that in the current specific data space application scenario (such as government affairs, medical care, or industrial chain scenarios), based on the search logs of the most recent historical period (usually 30 days), for each historical search, the number of results containing core keywords or semantic entities is counted, and this number is then divided by the total number of search results. The resulting ratio is the search result match degree, which is used to measure the degree of match between the search results and the user's core needs. For example, in a government data search, there are 100 total results, of which 80 contain the core keyword "tax data". In this case, the search result match degree is 80 ÷ 100 = 0.8.
[0087] The precise click-through rate (CTR) is determined by calculating the ratio of the number of results with a click-through rate of at least a preset match threshold to the total number of clicks in each historical search. This means that for each historical search, a match threshold (e.g., 90%) is set. The number of clicks on results that meet or exceed this threshold is counted. This number of clicks is then divided by the total number of clicks. The resulting ratio is the precise CTR, reflecting the user's interest in precise matching results. For example, if a search results in 50 total clicks, and 30 of them have a match of at least 90%, the precise CTR is 30 ÷ 50 = 0.6.
[0088] The number of business dimensions covered by a single search request in each historical search result is determined as the number of associated dimensions. In each historical search result, the number of business dimensions covered by the results returned by a single search request is considered the number of associated dimensions, which is used to assess the comprehensiveness of the search results. For example, in a medical scenario, if a search request returns results covering three business dimensions: patient visit information, medication information, and disease trajectory information, the number of associated dimensions is 3.
[0089] The association utilization rate is determined by calculating the ratio of the number of requests that used the associated results to the total number of requests in each historical search. For each historical search, the number of requests that used the associated results is counted and divided by the total number of search requests. The resulting ratio is the association utilization rate, which reflects the extent to which users utilize the associated results. Assuming there are 100 search requests over a period of time, and 60 of them use the associated results, the association utilization rate is 60 ÷ 100 = 0.6.
[0090] The freshness is determined by determining the percentage of data in each historical search result that has been updated within the preset period of the current application scenario. For each historical search result, a time period is set based on the current application scenario (12 hours for industrial chain scenarios, 7 days for government scenarios, and 3 days for medical scenarios). The amount of data updated within this period is counted and divided by the total amount of data. The resulting percentage is the freshness, which is used to measure the timeliness of the search result data. For example, in a search for a medical scenario, there are a total of 200 data items, of which 150 have been updated within 3 days. The freshness is 150 ÷ 200 = 0.75.
[0091] The real-time click-through rate (CTR) is determined by calculating the ratio of clicks on results with a freshness threshold of at least 80% to the total number of clicks in each historical search result. For each historical search result, a freshness threshold (e.g., 80%) is set. The number of clicks on results with a freshness that meets or exceeds this threshold is counted and then divided by the total number of clicks. The resulting ratio is the real-time CTR, which is used to further measure user attention to fresh data. For example, if the total number of clicks is 80, and 50 of them are on results with a freshness of at least 80%, the real-time CTR is 50 ÷ 80 = 0.625.
[0092] The actual value of the precision index is determined based on the search result matching degree and the precise click-through rate of each historical search result. The two are combined through weighted calculation or other statistical methods to obtain the actual value of the precision index. For example, the weight of the search result matching degree can be set to 0.6, and the weight of the precise click-through rate can be set to 0.4. The actual value of the comprehensiveness index is determined based on the number of associated dimensions and the associated usage rate of each historical search result. It is also possible to combine the two through weighted calculation and other methods to obtain the actual value of the comprehensiveness index. For example, the weight of the number of associated dimensions is set to 0.5, and the weight of the associated usage rate is set to 0.5. The actual value of the timeliness index is determined based on the freshness and real-time click-through rate of each historical search result. Through similar weighted calculation and other means, these two data are integrated to obtain the actual value of the timeliness index. For example, assume that the weight of freshness is 0.6 and the weight of real-time click-through rate is 0.4.
[0093] The user satisfaction of each historical search result is calculated based on the click rate, collection rate and usage rate of each historical search result in the historical search log within the latest historical period in the current application scenario.
[0094] like Figure 5 As shown, the present invention proposes an implementation method of an efficient management system for a data directory in a data space, including: The metadata collection and update module is used to incrementally update metadata information related to the data directory obtained through multiple channels, data product registration information, and cross-node data directory information, and centrally store them in the distributed data storage system; The index construction and retrieval module is used to build full-text indexes and vector indexes of metadata in the distributed data storage system. Based on the retrieval mode corresponding to the request type of the user's retrieval request, the vector generated based on the request content is matched with the full-text index or vector index to obtain the retrieval response result; The consistency analysis and fusion module is used to call the large language model to extract the semantic entity list and semantic vector of each dataset, mark semantically equivalent datasets based on the semantic vectors of all datasets, build a dataset node graph based on all semantic entity lists, and output a list of semantically equivalent fields based on the semantically equivalent datasets and dataset node graph; The associated recommendation and output module is used to determine the associated recommendation results of the search response results based on the list of semantically equivalent fields, and synchronously return the search response results and the associated recommendation results.
[0095] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the present invention and its equivalents, the present invention is intended to include these modifications and variations.
Claims
1. An efficient management method for a data directory in a data space, characterized in that: include: Perform incremental metadata updates on metadata information related to data directories, data product registration information, and cross-node data directory information obtained through multiple channels, and centrally store them in a distributed data storage system; Construct a full-text index and vector index for metadata in a distributed data storage system. Based on the search mode corresponding to the request type of the user's search request, match the vector generated based on the request content with the full-text index or vector index to obtain a search response result. Call the large language model to extract the semantic entity list and semantic vector of each dataset, mark semantically equivalent datasets based on the semantic vectors of all datasets, build a dataset node graph based on all semantic entity lists, and output a list of semantically equivalent fields based on the semantically equivalent datasets and dataset node graph; Determine the associated recommendation results of the search response results based on the list of semantically equivalent fields, and return the search response results and the associated recommendation results synchronously.
2. The efficient management method for data directories in a data space according to claim 1, characterized in that: The distributed data storage system includes a distributed storage subsystem compatible with the object storage protocol, a distributed relational data subsystem and a vector index subsystem. The distributed storage subsystem compatible with the object storage protocol is used to store unstructured data and serves as the underlying storage support for the distributed relational data subsystem and the vector index subsystem.
3. The efficient management method for data directories in a data space according to claim 1, characterized in that: Perform incremental metadata updates on metadata information related to data catalogs, data product registration information, and cross-node data catalog information obtained through multiple channels, and centrally store them in a distributed data storage system, including: Receive manually registered metadata information through the data connector's directory upload interface. At the same time, as a back-end reverse proxy for the data connector, directly access the data provider's application system to extract metadata and generate data product registration information based on the data product registration template. When the metadata update service is deployed near a regional functional node or a global functional node, cross-node data directory information is retrieved through the functional node's query service; Encode metadata information, data product registration information, and cross-node data directory information into metadata operation logs and write them into the distributed data storage system to complete incremental metadata updates.
4. The efficient management method for data directories in a data space according to claim 1, characterized in that: Build a full-text index and vector index for metadata in the distributed data storage system. Based on the search mode corresponding to the request type of the user's search request, match the vector generated based on the request content with the full-text index or vector index to obtain the search response results, including: Build full-text indexes and vector indexes based on metadata in distributed data storage systems; If the user's search request is a keyword or semi-structured keyword request, the full-text index matching process prioritizes precise matching based on the keywords in the user's search request to generate a first search response result. Simultaneously, the user's search request is converted into a basic semantic vector using a large language model, and semantically relevant datasets of the basic semantic vector are matched in the vector index to generate a second search response result. If the user's search request is a semantic demand request, the high-precision semantic vector conversion based on the user's search request is prioritized during the vector index matching process to generate a second search response result. At the same time, the core keywords in the user's search request are extracted and matched against a dataset containing the core keywords in the full-text index to generate a first search response result. The retrieval response result includes a first retrieval response result and a second retrieval response result that are generated synchronously.
5. The efficient management method for data directories in a data space according to claim 1, characterized in that: Call the large language model to extract the semantic entity list and semantic vector of each dataset, mark semantically equivalent datasets based on the semantic vectors of all datasets, build a dataset node graph based on all semantic entity lists, and output a list of semantically equivalent fields based on the semantically equivalent datasets and dataset node graph, including: Call the large language model to extract the semantic entity list and semantic vector of each data set; Calculate the cosine similarity of the semantic vectors of any two data sets, and mark any two data sets with similarity greater than a preset similarity threshold as semantically equivalent data sets; Analyze the field names and field annotations of structured data sets and extract semantic entities to add to the semantic entity list; Construct a dataset node graph with the dataset as the vertex and the semantic entity list as the condition for the intersection; Verify the semantic consistency of fields for semantically equivalent datasets and datasets whose correlation in the dataset node graph is greater than a preset correlation threshold and output a list of semantically equivalent fields.
6. The efficient management method for data directories in a data space according to claim 1, characterized in that: Determine the related recommendation results of the search response results based on the list of semantically equivalent fields, and return the search response results and the related recommendation results simultaneously, including: Provides list category browsing and graph node navigation based on the dataset node graph; The first search response result and the second search response result in the search response result are fused based on the result fusion algorithm to obtain the final search result; Generate related recommendation results based on the correlation between the node corresponding to the final search result and the remaining nodes in the dataset node graph and the list of semantically equivalent fields; The final search results and related recommendation results are returned synchronously.
7. The efficient management method for data directories in a data space according to claim 6, characterized in that: The specific implementation method of the result fusion algorithm includes: Calculate the inverse ranking of each data set in the first search response result and the inverse ranking of each data set in the second search response result respectively; Determining weights of the first search response result and the second search response result; Perform a weighted operation on the two inverse rankings based on the weights of the first search response result and the second search response result to obtain a fusion score of the corresponding data set; The top-k datasets are selected in descending order of fusion scores as the final retrieval results.
8. The efficient management method for data directories in a data space according to claim 7, characterized in that: Determining the weights of the first search response result and the second search response result includes: Determine the current application scenario of the data space and define multiple types of search demand indicators for the current application scenario, including accuracy indicators, comprehensiveness indicators, and timeliness indicators; Based on the historical search logs in the latest historical period of the current application scenario, determine the actual values of the accuracy index, comprehensiveness index, timeliness index and user satisfaction of each historical search result; Taking the actual values of the accuracy, comprehensiveness, and timeliness indicators of each historical search result as independent variables and user satisfaction as the dependent variable, a multivariate linear regression model is used to determine the linear relationship between multiple types of search demand indicators and user satisfaction in the current application scenario. Based on the linear relationship between multiple retrieval demand indicators and user satisfaction in the current application scenario, the optimal weight combination of multiple retrieval demand indicators in the current application scenario is determined; Based on the optimal weight combination of multiple retrieval demand indicators in the current application scenario, an indicator weight matrix is constructed. Based on the correlation coefficients between multiple retrieval demand indicators and full-text indexes and vector indexes, an indicator-index correlation matrix is constructed. Based on the constraint factors of full-text indexes and vector indexes in the current application scenario, a scenario constraint matrix is constructed. The result obtained by multiplying the indicator weight matrix, the indicator-index association matrix and the scenario constraint matrix in sequence is used as the optimized index weight matrix. Based on the dual index weights and the normalization operation of all elements in the optimized index weight matrix, the final index weight matrix is obtained, and the weights of the first retrieval response result and the second retrieval response result are determined based on the elements in the final index weight matrix.
9. The efficient management method for data directories in a data space according to claim 8, characterized in that: Based on the historical search logs in the latest historical period of the current application scenario, determine the actual values of the accuracy index, comprehensiveness index, timeliness index and user satisfaction of each historical search result, including: Based on the historical search logs in the latest historical period under the current application scenario, determine the ratio of the number of results containing core keywords or semantic entities in each historical search result to the total number of results as the search result matching degree, and determine the ratio of the number of results with a click matching degree not less than a preset matching degree threshold to the total number of clicks in each historical search result as the precise click-through rate, and determine the number of business dimensions covered by the results returned by a single search request in each historical search result as the number of associated dimensions, and determine the ratio of the number of requests using associated results to the total number of requests in each historical search result as the associated usage rate, and determine the proportion of data in the search results in each historical search result whose update time does not exceed the preset cycle time of the current application scenario as the freshness, and determine the ratio of the number of clicks on results with a freshness not less than the preset freshness threshold to the total number of clicks in each historical search result as the real-time click-through rate; The actual value of the precision index is determined based on the search result matching degree and the accurate click rate of each historical search result, the actual value of the comprehensiveness index is determined based on the number of associated dimensions and the associated usage rate of each historical search result, and the actual value of the timeliness index is determined based on the freshness and real-time click rate of each historical search result; The user satisfaction of each historical search result is calculated based on the click rate, collection rate and usage rate of each historical search result in the historical search log within the latest historical period in the current application scenario.
10. An efficient management system for data directories in a data space, characterized in that include: The metadata collection and update module is used to incrementally update metadata information related to the data directory obtained through multiple channels, data product registration information, and cross-node data directory information, and centrally store them in the distributed data storage system; The index construction and retrieval module is used to build full-text indexes and vector indexes of metadata in the distributed data storage system. Based on the retrieval mode corresponding to the request type of the user's retrieval request, the vector generated based on the request content is matched with the full-text index or vector index to obtain the retrieval response result; The consistency analysis and fusion module is used to call the large language model to extract the semantic entity list and semantic vector of each dataset, mark semantically equivalent datasets based on the semantic vectors of all datasets, build a dataset node graph based on all semantic entity lists, and output a list of semantically equivalent fields based on the semantically equivalent datasets and dataset node graph; The associated recommendation and output module is used to determine the associated recommendation results of the search response results based on the list of semantically equivalent fields, and synchronously return the search response results and the associated recommendation results.
Citation Information
Patent Citations
Sentence recommendation method and device, electronic device, and storage medium
CN109542247A
Multi-source mobile application-oriented knowledge graph construction method
CN115292520A
Medical data search method and device, electronic equipment and storage medium
CN115547443A
Task execution method and device, storage medium and electronic equipment
CN118193757A
Composite symbolic and non-symbolic artificial intelligence system for advanced reasoning and semantic search
US20240386015A1
Cited By
Automobile after-sales service information retrieval method and system based on cloud database
CN121722960A