Data retrieval method and device, equipment and storage medium

By storing multi-source heterogeneous data in regions and vectorized and cross-level correlation, the problems of data retrieval efficiency and comprehensiveness of large models in complex and professional tasks are solved, and an efficient and flexible data retrieval method is realized.

CN120256545APending Publication Date: 2025-07-04SERVYOU SOFTWARE GRP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510422444.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the development of existing technology, the efficient fusion of multi-source heterogeneous vertical data and the flexible and controllable index enhancement effects have not been effectively solved, resulting in insufficient data retrieval efficiency and comprehensiveness of complex and professional tasks.

Method used

Cross-level data retrieval is achieved by storing multi-source heterogeneous data into the data knowledge layer, industry knowledge layer and policy knowledge layer, and using entity relationship extraction instructions and preset walk strategies for vectorization and association.

Benefits of technology

It improves the efficiency and comprehensiveness of multi-source heterogeneous data retrieval, enhances the level of data understanding and retrieval flexibility, and adapts to the needs of complex business scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256545A_ABST
    Figure CN120256545A_ABST
Patent Text Reader

Abstract

The invention discloses a data retrieval method and device, equipment and a storage medium, and relates to the technical field of data retrieval, and the method comprises the steps: obtaining tax data to be retrieved, and storing the tax data to corresponding positions in a preset data storage layer according to corresponding data types; extracting entities in the to-be-retrieved tax data and relationships between the entities, and vectorizing the to-be-retrieved tax data in a preset data storage layer according to the entities and the relationships between the entities to obtain corresponding target vector data; respectively selecting shared entities from the data knowledge layer, the industry knowledge layer and the policy knowledge layer according to the matching degree of the entities among the different layers so as to associate the entities among the different layers; and obtaining a data retrieval instruction, and retrieving the target vector data based on a corresponding retrieval range and a preset migration strategy to obtain target tax data. The multi-source heterogeneous data are associated and fused, so that the efficiency and comprehensiveness of multi-source heterogeneous data retrieval are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data retrieval technology, and in particular to a data retrieval method, device, equipment and storage medium. Background Art

[0002] Currently in the field of large model application development, the fastest and lowest-cost way to develop professional vertical large model applications is to use the RAG (Retrieval-augmented Generation) framework to provide general large models with professional knowledge references related to specific tasks to complete tasks in professional fields.

[0003] At present, this RAG architecture solves the problem of insufficient professional capabilities of large models at a lower cost, but compared with the standard large model training paradigm, that is, using massive data for pre-training, instruction fine-tuning, and reinforcement learning, it still has the problem of insufficient knowledge. Because in different professional fields, what needs to be completed is not just a simple conceptual description task of one question and one answer, but may also involve more flexible and complex professional tasks such as specific "problem analysis", "business guidance", "relevant case references", and "applicable policy recommendations". These complex tasks often require more complex, continuous, and multi-source data to supplement the basic capabilities of large models. Therefore, how to efficiently integrate complex multi-source heterogeneous vertical domain data and achieve more flexible and controllable index enhancement effects is a technical problem that needs to be solved urgently. Summary of the invention

[0004] In view of this, the purpose of the present invention is to provide a data retrieval method, device, equipment and storage medium, which can improve the efficiency and comprehensiveness of multi-source heterogeneous data retrieval by associating and fusing multi-source heterogeneous data and searching data with different search scopes and search strategies. The specific scheme is as follows:

[0005] In a first aspect, the present application provides a data retrieval method, comprising:

[0006] Obtain the tax data to be retrieved, and store each of the tax data to be retrieved in a corresponding position in a preset data storage layer according to the data type of the tax data to be retrieved; wherein the preset data storage layer includes a data knowledge layer, an industry knowledge layer, and a policy knowledge layer, the data knowledge layer is used to store metadata corresponding to the tax data to be retrieved, the industry knowledge layer is used to store industry knowledge data in the tax data to be retrieved, and the policy knowledge layer is used to store policy knowledge data in the tax data to be retrieved;

[0007] Receive the entity relationship extraction instruction input by the user, and use the extraction mode corresponding to the entity relationship extraction instruction to extract the entities and the relationships between entities in the to-be-retrieved tax data, so as to store the to-be-retrieved data in the form of entity relationships, and vectorize the to-be-retrieved tax data in the preset data storage layer according to the entities and the relationships between entities to obtain the corresponding target vector data;

[0008] Select the shared entities between layers from the data knowledge layer, the industry knowledge layer, and the policy knowledge layer according to the matching degree of entities between different layers, so as to associate the entities between different layers;

[0009] Obtain the data retrieval instruction input by the user, and retrieve the target vector data based on the retrieval range and the preset random walk strategy corresponding to the data retrieval instruction to obtain the target tax data corresponding to the data retrieval instruction.

[0010] Optionally, the extracting the entities and the relationships between entities in the to-be-retrieved tax data by using the extraction mode corresponding to the entity relationship extraction instruction includes:

[0011] If the extraction mode corresponding to the entity relationship extraction instruction is the passive autonomous mode, use the preset large language model and the preset prompt words to extract all the entities and the corresponding relationships between entities in the to-be-retrieved tax data;

[0012] If the extraction mode corresponding to the entity relationship extraction instruction is the active non-expansion mode, generate the first prompt word according to the entity relationship extraction instruction, and use the preset large language model and the first prompt word to extract the target entities corresponding to the entity relationship extraction instruction and the target relationships between the target entities in the to-be-retrieved tax data;

[0013] If the extraction mode corresponding to the entity relationship extraction instruction is the active expansion mode, generate the first prompt word according to the entity relationship extraction instruction, use the first prompt word to supplement the preset prompt word to obtain the second prompt word, and use the preset large language model and the second prompt word to extract the target entities and the target relationships between the target entities.

[0014] Optionally, the vectorizing the to-be-retrieved tax data in the preset data storage layer according to the entities and the relationships between entities includes:

[0015] Obtain the association strength between entities based on the relationships between entities, so as to vectorize the to-be-retrieved tax data in each preset data storage layer according to the association strength between entities.

[0016] Optionally, the data retrieval method further includes:

[0017] Determine target multi-source heterogeneous data from the data to be retrieved; wherein, the target multi-source heterogeneous data is non-text data in the data to be retrieved;

[0018] Determine whether the target multi-source heterogeneous data is the target tax data corresponding to the data retrieval instruction. If the target multi-source heterogeneous data is not the target tax data corresponding to the data retrieval instruction, then associate the target multi-source heterogeneous data with the corresponding entity, so that the user can retrieve the corresponding target multi-source heterogeneous data after retrieving the target tax data.

[0019] Optionally, retrieving the target vector data based on the retrieval range corresponding to the data retrieval instruction and a preset traversal strategy to obtain the target tax data corresponding to the data retrieval instruction includes:

[0020] Perform syntax verification on the data retrieval instruction. If the data retrieval instruction passes the syntax verification, then perform vectorization processing on the data retrieval instruction to obtain a corresponding vectorized instruction;

[0021] Match the target vector data with the vectorized instruction based on the retrieval range corresponding to the data retrieval instruction and the preset traversal strategy to obtain the target tax data corresponding to the data retrieval instruction.

[0022] Optionally, retrieving the target vector data based on the retrieval range corresponding to the data retrieval instruction and a preset traversal strategy includes:

[0023] Determine whether the retrieval range corresponding to the data retrieval instruction only includes a single level in the preset data storage layer. If the retrieval range corresponding to the data retrieval instruction only includes a single level in the preset data storage layer, then use the preset traversal strategy to perform retrieval within a single level on the target vector data in the preset data storage layer corresponding to the data retrieval instruction;

[0024] If the retrieval range corresponding to the data retrieval instruction includes different levels in the preset data storage layer, then use the preset traversal strategy and the shared entity to perform cross-level retrieval on the target vector data in each preset data storage layer corresponding to the data retrieval instruction.

[0025] Optionally, retrieving the target vector data based on the retrieval range corresponding to the data retrieval instruction and a preset traversal strategy includes:

[0026] Determine the retrieval range corresponding to the data retrieval instruction;

[0027] Retrieve the target vector data within the retrieval range corresponding to the data retrieval instruction by using the preset random walk strategy; wherein, the preset random walk strategy includes a meta-path random walk strategy and a random walk strategy.

[0028] If the preset random walk strategy corresponding to the data retrieval instruction is a random walk strategy, determine the first target random walk path corresponding to each entity based on the weight corresponding to each entity, and retrieve the target vector data according to the first target random walk path.

[0029] If the preset random walk strategy corresponding to the data retrieval instruction is a meta-path random walk strategy, set the random walk condition according to the data type corresponding to each entity, and determine the second target random walk path corresponding to each entity according to the random walk condition and the weight corresponding to each entity, and retrieve the target vector data according to the second target random walk path.

[0030] In a second aspect, the present application provides a data retrieval device, including:

[0031] A data storage module, configured to obtain the tax data to be retrieved, and store each piece of the tax data to be retrieved at a corresponding position in a preset data storage layer according to the data type of the tax data to be retrieved; wherein, the preset data storage layer includes a data knowledge layer, an industry knowledge layer, and a policy knowledge layer, the data knowledge layer is used to store the metadata corresponding to the tax data to be retrieved, the industry knowledge layer is used to store the industry knowledge data in the tax data to be retrieved, and the policy knowledge layer is used to store the policy knowledge data in the tax data to be retrieved.

[0032] A data vectorization module, configured to receive an entity relationship extraction instruction input by a user, extract entities and relationships between entities in the tax data to be retrieved by using an extraction mode corresponding to the entity relationship extraction instruction, so as to store the tax data to be retrieved in the form of entity relationships, and vectorize the tax data to be retrieved in the preset data storage layer according to the entities and the relationships between entities to obtain corresponding target vector data.

[0033] An entity selection module, configured to respectively select shared entities between layers from the data knowledge layer, the industry knowledge layer, and the policy knowledge layer according to the matching degree of entities between different layers, so as to associate entities between different layers.

[0034] A data retrieval module, configured to obtain a data retrieval instruction input by a user, and retrieve the target vector data based on the retrieval range and the preset random walk strategy corresponding to the data retrieval instruction, so as to obtain the target tax data corresponding to the data retrieval instruction.

[0035] In a third aspect, the present application provides an electronic device, including:

[0036] a memory for storing a computer program;

[0037] a processor for executing the computer program to implement the foregoing data retrieval method.

[0038] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program, where the computer program, when executed by a processor, implements the foregoing data retrieval method.

[0039] In the present application, first, tax data to be retrieved needs to be obtained, and each of the tax data to be retrieved is respectively stored in a corresponding position in a preset data storage layer according to the data type of the tax data to be retrieved; wherein, the preset data storage layer includes a data knowledge layer, an industry knowledge layer, and a policy knowledge layer, the data knowledge layer is used to store metadata corresponding to the tax data to be retrieved, the industry knowledge layer is used to store industry knowledge data in the tax data to be retrieved, and the policy knowledge layer is used to store policy knowledge data in the tax data to be retrieved; then an entity relationship extraction instruction input by a user is received, and an extraction mode corresponding to the entity relationship extraction instruction is used to extract entities and relationships between entities in the tax data to be retrieved, so as to store the tax data to be retrieved in the form of entity relationships, and vectorize the tax data to be retrieved in the preset data storage layer according to the entities and the relationships between entities to obtain corresponding target vector data; then, shared entities between layers are respectively selected from the data knowledge layer, the industry knowledge layer, and the policy knowledge layer according to the matching degree of entities between different layers, so as to associate entities between different layers; finally, a data retrieval instruction input by the user is obtained, and the target vector data is retrieved based on the retrieval range corresponding to the data retrieval instruction and a preset walking strategy to obtain target tax data corresponding to the data retrieval instruction. Thus, it can be seen that in the present application, by storing multi-source heterogeneous data in different regions, extracting entity relationships of the multi-source heterogeneous data in each region, and vectorizing the multi-source heterogeneous data in each region, the understanding degree of the model for the data is deepened; by constructing connections between the multi-source heterogeneous data within each region and constructing connections between different regions, the efficient fusion of multi-source heterogeneous data is achieved; by retrieving the multi-source heterogeneous data according to the data retrieval range required by the user and the walking path during data retrieval, the flexibility of data retrieval is improved. Description of the Drawings

[0040] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.

[0041] Figure 1 Flowchart of a data retrieval method disclosed in this application;

[0042] Figure 2 Schematic diagram of the structure of a data retrieval system disclosed in this application;

[0043] Figure 3 Schematic diagram of the process of a specific data retrieval method disclosed in this application;

[0044] Figure 4 Flowchart of a method for extracting entity relationships in different modes disclosed in this application;

[0045] Figure 5 Schematic diagram of the process of a method for processing multi-source heterogeneous data disclosed in this application;

[0046] Figure 6 Schematic diagram of a data cross-layer association structure disclosed in this application;

[0047] Figure 7 Flowchart of a specific data retrieval method disclosed in this application;

[0048] Figure 8 Schematic diagram of a single-layer data retrieval method disclosed in this application;

[0049] Figure 9 Schematic diagram of a cross-layer data retrieval method disclosed in this application;

[0050] Figure 10 Schematic diagram of a random walk method disclosed in this application;

[0051] Figure 11 Schematic diagram of a meta-path walk method disclosed in this application;

[0052] Figure 12 Schematic diagram of the structure of a data retrieval device disclosed in this application;

[0053] Figure 13 Schematic diagram of the structure of an electronic device disclosed in this application. Detailed implementation manners

[0054] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0055] Currently, for data retrieval, a retrieval enhancement framework is usually adopted to complete it. This method has the problem of being unable to process multi-source heterogeneous data, and the retrieval effect is uncontrollable. Therefore, this application provides a data retrieval method. By associating and fusing multi-source heterogeneous data, and retrieving the data with different retrieval scopes and retrieval strategies, the controllability of the retrieval process is ensured, and the efficiency and comprehensiveness of retrieving multi-source heterogeneous data are improved.

[0056] See Figure 1 As shown, an embodiment of the present invention discloses a data retrieval method, including:

[0057] Step S11, obtain the tax data to be retrieved, and store each of the tax data to be retrieved in the corresponding position in the preset data storage layer according to the data type of the tax data to be retrieved; wherein, the preset data storage layer includes a data knowledge layer, an industry knowledge layer, and a policy knowledge layer. The data knowledge layer is used to store the metadata corresponding to the tax data to be retrieved, the industry knowledge layer is used to store the industry knowledge data in the tax data to be retrieved, and the policy knowledge layer is used to store the policy knowledge data in the tax data to be retrieved.

[0058] In the data preprocessing stage, this embodiment will hierarchically define the knowledge dimension based on different multi-source heterogeneous data from a macroscopic perspective. For example, Figure 2 as shown in the data retrieval system in Figure 3 in the tax field, knowledge can be defined as three layers: "data knowledge layer", "industry knowledge layer", and "policy knowledge layer". Each layer contains multi-source heterogeneous data from the tax and related fields. Through the multi-modal data processing module, feature encoding is performed for different data types for association and search. It should be noted that each layer is at the same level and non-isolated, and they are associated through shared nodes and shared edges. And in the data retrieval stage, as

[0059] It should be noted that the data types stored in the three layers of the data knowledge layer, the industry knowledge layer, and the policy knowledge layer are not the same. Among them, the data knowledge layer is used to store the metadata corresponding to the tax data to be retrieved. That is, the tax data to be retrieved is not directly stored in the data knowledge layer, but the information used to describe the data attributes is stored to support functions such as indicating the storage location, historical data, resource search, and file records. For example, if the data to be retrieved is a certain table in the database, the table is not directly stored in the data knowledge layer, but the table name or table index link corresponding to the table and other data that can index the table are stored; the industry knowledge layer is used to store the industry knowledge data in the tax data to be retrieved, and the policy knowledge layer is used to store the policy knowledge data in the tax data to be retrieved; in this embodiment, the data stored in the industry knowledge layer and the policy knowledge layer can also be set as the metadata corresponding to the industry knowledge and policy knowledge according to the requirements. By storing the metadata corresponding to the required data in each layer, the storage space occupancy is greatly reduced, thereby reducing the data storage cost.

[0060] Step S12, receive the entity relationship extraction instruction input by the user, and use the extraction mode corresponding to the entity relationship extraction instruction to extract the entities and the relationships between entities in the tax data to be retrieved, so as to store the data to be retrieved in the form of entity relationships, and vectorize the tax data to be retrieved in the preset data storage layer according to the entities and the relationships between entities to obtain the corresponding target vector data.

[0061] In this embodiment, the process of using the extraction mode corresponding to the entity relationship extraction instruction to extract the entities and the relationships between entities in the tax data to be retrieved may specifically include: if the extraction mode corresponding to the entity relationship extraction instruction is the passive autonomous mode, use the preset large language model and the preset prompt words to extract all the entities and the corresponding relationships between entities in the tax data to be retrieved; if the extraction mode corresponding to the entity relationship extraction instruction is the active non-expansion mode, generate the first prompt word according to the entity relationship extraction instruction, and use the preset large language model and the first prompt word to extract the target entities and the target relationships between the target entities in the tax data to be retrieved corresponding to the entity relationship extraction instruction; if the extraction mode corresponding to the entity relationship extraction instruction is the active expansion mode, generate the first prompt word according to the entity relationship extraction instruction, use the first prompt word to supplement the preset prompt word to obtain the second prompt word, and use the preset large language model and the second prompt word to extract the target entities and the target relationships between the target entities, that is, Figure 4As shown in the figure, the entity relationship extraction in this embodiment is divided into three modes, namely, passive autonomous mode, active non-expansion mode, and active expansion mode. In the stage of constructing the graph data structure, the system supports the user to form instruction extraction prompts by defining entities and relationship types, so as to improve the effectiveness of the constructed knowledge graph. When the active definition mode is enabled by the user, it is also possible to set whether it is an expansion mode. If it is enabled, the system will automatically enrich and supplement the elements other than the entities and relationship content defined by the user. When the user does not actively define, the default is the automatic mode, and the large model will extract the entities and relationships in the multi-source heterogeneous data converted into text data based on the instructions designed by the prompt engineering. Through the passive autonomous mode, the system will automatically perform entity and relationship extraction, and extract all the entities and relationships corresponding to the data stored in the three-layer data storage layer. Through the active non-expansion mode and the active expansion mode, the user can set the prompts by themselves and only extract the required entities and the relationships between entities, reducing the data calculation volume.

[0062] In this embodiment, after obtaining the relationship between entities, it is also necessary to obtain the association strength between the data in the preset data storage layer according to the relationship between entities, so as to vectorize the data according to the association strength between the data. Correspondingly, the process of vectorizing the to-be-retrieved tax data in the preset data storage layer according to the entity and the relationship between entities may specifically include: obtaining the association strength between entities based on the relationship between entities, so as to vectorize the to-be-retrieved tax data in each preset data storage layer according to the association strength between entities. Specifically, for relational data, this embodiment classifies it in the "data knowledge layer", and this embodiment associates the association relationships of each data set in this layer in a more accurate manner. Among them, the "field name" and "field meaning description" calculate the association strength of each data set through the "weighted association strength algorithm" for vectorization and use it as the edge weight for subsequent graph walking actions. The above weighted association strength algorithm is: in the data knowledge layer, match the field name and field meaning description between each data respectively to obtain the association strength between each data.

[0063] For text data, including policy knowledge and industry knowledge, this embodiment classifies it into the "industry knowledge layer" and the "policy knowledge layer". This embodiment first converts these data into text data with a unified format through a preset multi-modal data preprocessing system; then uses a controllable entity relationship extraction module to extract the entities and relationships, and calculates the softmax (an activation function) probability as the association strength according to the occurrence frequency of the basic triples of (node, edge, node) for use as the edge weight for subsequent graph walking actions.

[0064] The data retrieval method in this embodiment also includes: determining target multi-source heterogeneous data from the data to be retrieved; wherein the target multi-source heterogeneous data is non-text data in the data to be retrieved; determining whether the target multi-source heterogeneous data is the target tax data corresponding to the data retrieval instruction, and if the target multi-source heterogeneous data is not the target tax data corresponding to the data retrieval instruction, then associating the target multi-source heterogeneous data with the corresponding entity, so that the user can obtain the corresponding target multi-source heterogeneous data after retrieving the target tax data. That is, for other multi-source heterogeneous data, such as pictures, audio, video, etc., it is necessary to decide whether to associate them with entities in the graph as metadata or to perform feature encoding as the object of vector search according to their specific uses. If it is the latter, Figure 5 As shown, the multimodal data feature encoding module in this embodiment can use the corresponding large model to process different types of multi-source heterogeneous data, and will convert these data into vector data, and internally associate and store them in the specified industry knowledge layer; specifically, if the above-mentioned other multi-source heterogeneous data is data that the user may need to retrieve, it will be stored in the corresponding position in the preset data storage layer according to its data type, and vectorized so that the user can retrieve it; if the above-mentioned other multi-source heterogeneous data is not data that the user directly needs to retrieve, it will be associated with the corresponding vectorized data, so that the user can obtain these multi-source heterogeneous data at the same time when retrieving the required data, thereby improving the comprehensiveness of the data retrieval results.

[0065] Step S13: selecting shared entities between layers from the data knowledge layer, the industry knowledge layer and the policy knowledge layer respectively according to the matching degree of entities between different layers, so as to associate entities between different layers.

[0066] It should be noted that if Figure 6 As shown, the multi-layer knowledge architecture is parallel and interconnected. Knowledge retrieval can be started from any layer and related data can be searched. When the wandering path needs to be a global mode (i.e. wandering from one layer to another), specific shared nodes are required to connect different knowledge layers. In the multi-source heterogeneous tax knowledge base system in this embodiment, the "data knowledge layer" and the other two layers use the "field name" node as the associated node. When the entity names of the other two layers have nodes that fully match them or have a semantic relevance that exceeds the threshold, these nodes are shared nodes, and the edges are called shared edges; similarly, the other two layers are associated according to the mixed association strength algorithm of the entity name. By setting shared nodes between layers, users can perform cross-layer detection of data, thereby obtaining different types of data, which improves the comprehensiveness of data retrieval.

[0067] Step S14: Obtain the data retrieval instruction input by the user, and retrieve the target vector data based on the retrieval range corresponding to the data retrieval instruction and the preset traversal strategy to obtain the target tax data corresponding to the data retrieval instruction.

[0068] In this embodiment, the process of retrieving the target vector data based on the retrieval range corresponding to the data retrieval instruction and the preset traversal strategy to obtain the target tax data corresponding to the data retrieval instruction includes: performing a syntax check on the data retrieval instruction. If the data retrieval instruction passes the syntax check, perform a vectorization process on the data retrieval instruction to obtain the corresponding vectorized instruction; match the target vector data with the vectorized instruction based on the retrieval range corresponding to the data retrieval instruction and the preset traversal strategy to obtain the target tax data corresponding to the data retrieval instruction. That is, the essence of the data retrieval process in this embodiment is a process of matching the vectorized instruction corresponding to the user's needs with the vectorized data in the preset data storage layer according to the preset traversal path. In this embodiment, the retrieval range of the data can be selected. For example, in this embodiment, the retrieval can be performed only in any one of the data knowledge layer, the industry knowledge layer, or the policy knowledge layer, or any two or all three of them can be selected for data retrieval; in addition, in this embodiment, different traversal strategies can also be selected for data retrieval, that is, different data retrieval paths can be selected for data retrieval. By customizing the data retrieval range, the obtained data can be highly relevant to the user's needs, avoiding the problem of decreased user experience caused by introducing other irrelevant data, and by customizing the data retrieval range, the computational amount of the system can be reduced to a certain extent, thereby improving the data retrieval efficiency; by selecting different traversal strategies for data retrieval, complex business scenarios can be adapted, and the usability of the data retrieval method in this embodiment is improved.

[0069] It can be seen that in this application, by storing multi-source heterogeneous data in different regions, extracting the entity relationships of the multi-source heterogeneous data in each region, and vectorizing the multi-source heterogeneous data in each region, the model's understanding of the data is deepened; by constructing the connections between the multi-source heterogeneous data within each region and constructing the connections between different regions, the efficient fusion of multi-source heterogeneous data is realized; by retrieving the multi-source heterogeneous data according to the data retrieval range required by the user and the traversal path during data retrieval, the flexibility of data retrieval is improved.

[0070] Based on the foregoing embodiments, this application describes the overall process of data retrieval. To make the technical solutions in this application more complete, next, this application will elaborate in detail on the process of retrieving based on the retrieval range and traversal strategy. Refer to Figure 7 As shown, this application discloses a specific data retrieval process, including:

[0071] Step S21: Determine the retrieval range and the random walk strategy corresponding to the data retrieval instruction.

[0072] In this embodiment, the retrieval range of the data may include one or more layers in the preset data storage layer. Therefore, before performing data retrieval, it is necessary to determine the retrieval range required by the user. Specifically, it is necessary to determine whether the retrieval range corresponding to the data retrieval instruction only includes a single level in the preset data storage layer. If the retrieval range corresponding to the data retrieval instruction only includes a single level in the preset data storage layer, then use the preset random walk strategy to retrieve the target vector data within a single level in the preset data storage layer corresponding to the data retrieval instruction; if the retrieval range corresponding to the data retrieval instruction includes different levels in the preset data storage layer, then use the preset random walk strategy and the shared entity to perform cross-level retrieval of the target vector data in each preset data storage layer corresponding to the data retrieval instruction. For example, in a specific embodiment, as Figure 8 shown, this embodiment can perform data retrieval only in the policy knowledge layer; in another specific embodiment, as Figure 9 shown, this embodiment can perform cross-level retrieval between the data knowledge layer and the policy knowledge layer.

[0073] In addition, it is also necessary to determine the random walk strategy corresponding to the data retrieval instruction. The data random walk strategy in this embodiment includes a random walk strategy and a meta-path random walk strategy. It should be noted that the above random walk strategy only walks according to the weights of each node, that is, each entity, while the meta-path random walk strategy needs to set a pre-walking condition corresponding to the data type in advance and walk according to the above pre-walking condition and the weights of the nodes. By determining the data retrieval range and the random walk strategy, the retrieved data can be made more in line with the user's needs and the user experience can be improved.

[0074] Step S22: If the random walk strategy corresponding to the data retrieval instruction is a random walk strategy, then determine the first target random walk path corresponding to each entity based on the weights of each entity, and retrieve the target vector data according to the first target random walk path and the retrieval range.

[0075] In this embodiment, the random walk strategy is to generate a knowledge sequence according to the node weights as the random walk probability. As Figure 10 shown, a random walk path is composed of three entities in the industry knowledge layer. The random walk strategy in this embodiment additionally introduces a retraction probability p and a depth search probability q. When calculating the transition probability from node v to its neighbor node x, if x is the previous visited node t, then the transition probability is the edge weight divided by p; if x is adjacent to t, then the transition probability is the edge weight; if x is not adjacent to t, then the transition probability is the edge weight divided by q. Then, these unnormalized probabilities are normalized (softmax) to obtain the final transition probability.

[0076] Step S23: If the random walk strategy corresponding to the data retrieval instruction is a meta-path random walk strategy, set random walk conditions according to the data types corresponding to each entity, determine second target random walk paths corresponding to each entity according to the random walk conditions and the weights corresponding to each entity, and retrieve the target vector data according to the second target random walk paths and the retrieval range.

[0077] In this embodiment, the meta-path random walk strategy is a random walk with the order of the random walk sequence of node meta-data types as the pre-random walk condition, indexing to the knowledge cluster of the specified sequence. For example, as Figure 11 shown, if the pre-random walk condition is set to be from data to policy and then to regulations, when retrieving data, it will first enter a certain node in the data knowledge layer to obtain data, and then cross-domain to the policy knowledge layer to obtain policy data and regulation data in sequence. It should be noted that when performing the meta-path random walk, it is still necessary to perform the random walk according to the weight of the foregoing node and the node transition probability calculation method.

[0078] It can be seen that in this application, by storing multi-source heterogeneous data in regions, extracting the entity relationships of the multi-source heterogeneous data in each region, and vectorizing the multi-source heterogeneous data in each region, the understanding degree of the model for the data is deepened; by constructing the connections between the multi-source heterogeneous data within each region and constructing the connections between different regions, the efficient fusion of multi-source heterogeneous data is realized; by retrieving the multi-source heterogeneous data according to the data retrieval range required by the user and the random walk path during data retrieval, the flexibility of data retrieval is improved.

[0079] Refer to Figure 12 shown, an embodiment of the present invention discloses a data retrieval device, including:

[0080] A data storage module 11, configured to obtain tax data to be retrieved, and store each of the tax data to be retrieved at a corresponding position in a preset data storage layer according to the data type of the tax data to be retrieved; wherein, the preset data storage layer includes a data knowledge layer, an industry knowledge layer, and a policy knowledge layer, the data knowledge layer is used to store the metadata corresponding to the tax data to be retrieved, the industry knowledge layer is used to store the industry knowledge data in the tax data to be retrieved, and the policy knowledge layer is used to store the policy knowledge data in the tax data to be retrieved;

[0081] The data vectorization module 12 is configured to receive an entity relationship extraction instruction input by a user, extract entities and relationships between entities in the to-be-retrieved tax data by using an extraction pattern corresponding to the entity relationship extraction instruction, so as to store the to-be-retrieved data in the form of entity relationships, and vectorize the to-be-retrieved tax data in the preset data storage layer according to the entities and the relationships between the entities, so as to obtain corresponding target vector data;

[0082] The entity selection module 13 is configured to separately select shared entities between layers from the data knowledge layer, the industry knowledge layer, and the policy knowledge layer according to the matching degree of entities between different layers, so as to associate the entities between different layers;

[0083] The data retrieval module 14 is configured to obtain a data retrieval instruction input by a user, and retrieve the target vector data based on a retrieval range corresponding to the data retrieval instruction and a preset walking strategy, so as to obtain target tax data corresponding to the data retrieval instruction.

[0084] It can be seen that in this application, by storing multi-source heterogeneous data in different regions, extracting entity relationships of the multi-source heterogeneous data in each region, and vectorizing the multi-source heterogeneous data in each region, the understanding degree of the model for the data is deepened; by constructing the connections between the multi-source heterogeneous data within each region and constructing the connections between different regions, the efficient fusion of multi-source heterogeneous data is realized; by retrieving the multi-source heterogeneous data according to the data retrieval range required by the user and the walking path during data retrieval, the flexibility of data retrieval is improved.

[0085] In some specific embodiments, the data vectorization module 12 may specifically include:

[0086] The first entity relationship extraction unit is configured to, if the extraction pattern corresponding to the entity relationship extraction instruction is the passive autonomous mode, extract all entities and corresponding relationships between entities in the to-be-retrieved tax data by using a preset large language model and a preset prompt word;

[0087] The second entity relationship extraction unit is configured to, if the extraction pattern corresponding to the entity relationship extraction instruction is the active non-expansion mode, generate a first prompt word according to the entity relationship extraction instruction, and extract target entities corresponding to the entity relationship extraction instruction and target relationships between the target entities in the to-be-retrieved tax data by using the preset large language model and the first prompt word;

[0088] The third entity relationship extraction unit is configured to, if the extraction mode corresponding to the entity relationship extraction instruction is the active expansion mode, generate the first prompt word according to the entity relationship extraction instruction, supplement the preset prompt word with the first prompt word to obtain the second prompt word, and use the preset large language model and the second prompt word to extract the target entity and the target relationship between the target entities.

[0089] In some specific embodiments, the data vectorization module 12 may specifically include:

[0090] The data vectorization unit is configured to obtain the association strength between entities based on the relationship between entities, and vectorize the to-be-retrieved tax data in each preset data storage layer according to the association strength between entities.

[0091] In some specific embodiments, the data retrieval device further includes:

[0092] The data determination module is configured to determine the target multi-source heterogeneous data from the to-be-retrieved data; wherein, the target multi-source heterogeneous data is non-text data in the to-be-retrieved data;

[0093] The data association module is configured to determine whether the target multi-source heterogeneous data is the target tax data corresponding to the data retrieval instruction. If the target multi-source heterogeneous data is not the target tax data corresponding to the data retrieval instruction, the target multi-source heterogeneous data is associated with the corresponding entity, so that the user can obtain the corresponding target multi-source heterogeneous data after retrieving the target tax data.

[0094] In some specific embodiments, the data retrieval module 14 may specifically include:

[0095] The syntax verification unit is configured to perform syntax verification on the data retrieval instruction. If the data retrieval instruction passes the syntax verification, the data retrieval instruction is vectorized to obtain the corresponding vectorized instruction;

[0096] The data matching unit is configured to match the target vector data with the vectorized instruction based on the retrieval range corresponding to the data retrieval instruction and the preset random walk strategy to obtain the target tax data corresponding to the data retrieval instruction.

[0097] In some specific embodiments, the data retrieval module 14 may specifically include:

[0098] A first data retrieval unit, configured to determine whether the retrieval range corresponding to the data retrieval instruction only includes a single level in the preset data storage layer. If the retrieval range corresponding to the data retrieval instruction only includes a single level in the preset data storage layer, then use the preset random walk strategy to perform in-layer retrieval of the target vector data in the preset data storage layer corresponding to the data retrieval instruction;

[0099] A second data retrieval unit, configured to, if the retrieval range corresponding to the data retrieval instruction includes different levels in the preset data storage layer, use the preset random walk strategy and the shared entity to perform cross-layer retrieval of the target vector data in each preset data storage layer corresponding to the data retrieval instruction.

[0100] In some specific embodiments, the data retrieval module 14 may specifically include:

[0101] A retrieval range determination unit, configured to determine the retrieval range corresponding to the data retrieval instruction;

[0102] A third data retrieval unit, configured to use the preset random walk strategy to perform retrieval of the target vector data within the retrieval range corresponding to the data retrieval instruction; wherein, the preset random walk strategy includes a meta-path random walk strategy and a random walk strategy;

[0103] A fourth data retrieval unit, configured to, if the preset random walk strategy corresponding to the data retrieval instruction is a random walk strategy, determine a first target random walk path corresponding to each entity based on the weight corresponding to each entity, and perform retrieval of the target vector data according to the first target random walk path;

[0104] A fifth data retrieval unit, configured to, if the preset random walk strategy corresponding to the data retrieval instruction is a meta-path random walk strategy, set a random walk condition according to the data type corresponding to each entity, and determine a second target random walk path corresponding to each entity according to the random walk condition and the weight corresponding to each entity, and perform retrieval of the target vector data according to the second target random walk path.

[0105] Furthermore, an embodiment of the present application also discloses an electronic device, Figure 13 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment. The content in the figure should not be regarded as any limitation on the scope of use of the present application.

[0106] Figure 13Schematic diagram of the structure of an electronic device 20 provided by an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the data retrieval method disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0107] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of the present application, and specific limitations are not imposed here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitations are imposed here.

[0108] In addition, as a carrier for resource storage, the memory 22 can be a read-only memory, a random access memory, a magnetic disk, or an optical disc, etc., and the resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be short-term storage or permanent storage.

[0109] Among them, the operating system 221 is used to manage and control each hardware device and the computer program 222 on the electronic device 20, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the data retrieval method executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 can further include computer programs that can be used to complete other specific tasks.

[0110] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the data retrieval method disclosed above is implemented. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated here.

[0111] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments, and the same or similar parts among the various embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0112] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered as exceeding the scope of this application.

[0113] The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of both. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0114] Finally, it should also be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0115] The technical solutions provided in this application have been introduced in detail above. Specific examples have been used in this article to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. A data retrieval method, characterized in that, Including: Obtain the tax data to be retrieved, and store each of the tax data to be retrieved in the corresponding position in the preset data storage layer according to the data type of the tax data to be retrieved; wherein, the preset data storage layer includes a data knowledge layer, an industry knowledge layer, and a policy knowledge layer, the data knowledge layer is used to store the metadata corresponding to the tax data to be retrieved, the industry knowledge layer is used to store the industry knowledge data in the tax data to be retrieved, and the policy knowledge layer is used to store the policy knowledge data in the tax data to be retrieved; Receive the entity relationship extraction instruction input by the user, and use the extraction mode corresponding to the entity relationship extraction instruction to extract the entities and the relationships between entities in the tax data to be retrieved, so as to store the tax data to be retrieved in the form of entity relationships, and vectorize the tax data to be retrieved in the preset data storage layer according to the entities and the relationships between entities, so as to obtain the corresponding target vector data; Select the shared entities between layers from the data knowledge layer, the industry knowledge layer, and the policy knowledge layer according to the matching degree of entities between different layers, so as to associate the entities between different layers; Obtain the data retrieval instruction input by the user, and retrieve the target vector data based on the retrieval range and the preset walking strategy corresponding to the data retrieval instruction, so as to obtain the target tax data corresponding to the data retrieval instruction.

2. The data retrieval method according to claim 1, wherein The extracting the entities and the relationships between entities in the tax data to be retrieved by using the extraction mode corresponding to the entity relationship extraction instruction includes: If the extraction mode corresponding to the entity relationship extraction instruction is the passive autonomous mode, use the preset large language model and the preset prompt words to extract all the entities and the corresponding relationships between entities in the tax data to be retrieved; If the extraction mode corresponding to the entity relationship extraction instruction is the active non-expansion mode, generate the first prompt word according to the entity relationship extraction instruction, and use the preset large language model and the first prompt word to extract the target entity corresponding to the entity relationship extraction instruction and the target relationship between the target entities in the tax data to be retrieved; If the extraction mode corresponding to the entity relationship extraction instruction is the active expansion mode, generate the first prompt word according to the entity relationship extraction instruction, use the first prompt word to supplement the preset prompt word to obtain the second prompt word, and use the preset large language model and the second prompt word to extract the target entity and the target relationship between the target entities.

3. The data retrieval method according to claim 1, wherein The vectorizing the tax data to be retrieved in the preset data storage layer according to the entities and the relationships between entities includes: Obtain the association strength between entities based on the relationships between entities, so as to vectorize the tax data to be retrieved in each preset data storage layer according to the association strength between entities.

4. The data retrieval method according to claim 1, wherein Also including: Determine the target multi-source heterogeneous data from the data to be retrieved; wherein, the target multi-source heterogeneous data is the non-text data in the data to be retrieved; Determine whether the target multi-source heterogeneous data is the target tax data corresponding to the data retrieval instruction. If the target multi-source heterogeneous data is not the target tax data corresponding to the data retrieval instruction, then associate the target multi-source heterogeneous data with the corresponding entity so that the user can obtain the corresponding target multi-source heterogeneous data after retrieving the target tax data.

5. The data retrieval method according to claim 1, characterized in that, Retrieving the target vector data based on the retrieval range corresponding to the data retrieval instruction and a preset random walk strategy to obtain the target tax data corresponding to the data retrieval instruction, including: Perform a syntax check on the data retrieval instruction. If the data retrieval instruction passes the syntax check, then perform a vectorization process on the data retrieval instruction to obtain a corresponding vectorized instruction; Based on the retrieval range corresponding to the data retrieval instruction and a preset random walk strategy, match the target vector data with the vectorized instruction to obtain the target tax data corresponding to the data retrieval instruction.

6. The data retrieval method according to claim 1, wherein Retrieving the target vector data based on the retrieval range corresponding to the data retrieval instruction and a preset random walk strategy, including: Determine whether the retrieval range corresponding to the data retrieval instruction only includes a single level in the preset data storage layer. If the retrieval range corresponding to the data retrieval instruction only includes a single level in the preset data storage layer, then use the preset random walk strategy to perform in-layer retrieval of the target vector data in the preset data storage layer corresponding to the data retrieval instruction; If the retrieval range corresponding to the data retrieval instruction includes different levels in the preset data storage layer, then use the preset random walk strategy and the shared entity to perform cross-level retrieval of the target vector data in each preset data storage layer corresponding to the data retrieval instruction.

7. The data retrieval method according to any one of claims 1 to 6, characterized in that Retrieving the target vector data based on the retrieval range corresponding to the data retrieval instruction and a preset random walk strategy, including: Determine the retrieval range corresponding to the data retrieval instruction; Use the preset random walk strategy to retrieve the target vector data within the retrieval range corresponding to the data retrieval instruction; where the preset random walk strategy includes a meta-path random walk strategy and a random walk strategy; If the preset random walk strategy corresponding to the data retrieval instruction is a random walk strategy, then determine a first target random walk path corresponding to each entity based on the weight corresponding to each entity, and retrieve the target vector data according to the first target random walk path; If the preset random walk strategy corresponding to the data retrieval instruction is a meta-path random walk strategy, then set a random walk condition according to the data type corresponding to each entity, and determine a second target random walk path corresponding to each entity according to the random walk condition and the weight corresponding to each entity, and retrieve the target vector data according to the second target random walk path.

8. A data retrieval device, characterized in that, including: A data storage module, configured to obtain tax data to be retrieved and store each of the tax data to be retrieved at corresponding positions in a preset data storage layer according to the data types of the tax data to be retrieved; wherein, the preset data storage layer includes a data knowledge layer, an industry knowledge layer, and a policy knowledge layer, the data knowledge layer is used to store metadata corresponding to the tax data to be retrieved, the industry knowledge layer is used to store industry knowledge data in the tax data to be retrieved, and the policy knowledge layer is used to store policy knowledge data in the tax data to be retrieved; A data vectorization module, configured to receive an entity relationship extraction instruction input by a user, extract entities and relationships between entities in the tax data to be retrieved by using an extraction pattern corresponding to the entity relationship extraction instruction, so as to store the tax data to be retrieved in the form of entity relationships, and vectorize the tax data to be retrieved in the preset data storage layer according to the entities and the relationships between the entities, so as to obtain corresponding target vector data; An entity selection module, configured to respectively select shared entities between layers from the data knowledge layer, the industry knowledge layer, and the policy knowledge layer according to the matching degree of entities between different layers, so as to associate entities between different layers; A data retrieval module, configured to obtain a data retrieval instruction input by a user, and retrieve the target vector data based on a retrieval range corresponding to the data retrieval instruction and a preset walking strategy, so as to obtain target tax data corresponding to the data retrieval instruction.

9. An electronic device, characterized in that, Comprising: A memory, configured to store a computer program; A processor, configured to execute the computer program to implement the data retrieval method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, For storing a computer program, which when executed by a processor implements the data retrieval method according to any one of claims 1 to 7.