A method, device and equipment for batch query of association relationships
Through offline pre-calculation, the coverage of related parties is determined and combined with the association prediction model, the problems of low efficiency and low recall of batch association relationship query are solved, and efficient query performance and recall improvement are achieved.
Patent Information
- Application Number
- CN202210298538.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-25
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-03-25
AI Technical Summary
It is difficult for the existing technology to efficiently conduct batch association query, resulting in low query efficiency and low recall rate.
The coverage of related parties is determined through offline pre-calculation, and combined with the pre-trained association prediction model, the prediction node pairs are predicted to improve recall and query performance.
It realizes the improvement of recall rate on the basis of ensuring query time consumption, and solves the problem of limited recall rate and complex links in the existing methods.
Smart Images

Figure CN114625783B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of graph processing technology, and in particular to a method, device and equipment for batch query of association relationships. Background Art
[0002] An association relationship refers to a relationship between nodes that are interconnected and whose strength of association meets the requirements of the business. For example, the relationship between the controlling shareholder, actual controller, director, supervisor and other senior executives of an enterprise and the enterprises directly or indirectly controlled by them, as well as other relationships that may lead to the transfer of corporate interests constitute an association relationship between enterprises. With the development of social economy, the association relationships in various fields are becoming increasingly complex. For example, banks conduct credit qualification reviews on enterprises, financial offices dig out clues to corporate risks, and intermediary institutions and sponsors conduct due diligence and material reviews on proposed listed companies. It is necessary to conduct association queries on a large number of nodes. However, due to insufficient identification during node query, various risks and even actual losses can be easily caused. Therefore, the effective identification of associated nodes has become a key factor in many fields such as risk control and telecommunications fraud prevention.
[0003] Currently, online graph query can easily query the relationship between a single node pair or several nodes, but it cannot meet the needs of batch query. When batch query is performed based on the combination of real-time query and near-line calculation, or offline pre-calculation and real-time query, the query efficiency cannot be met, and the recall rate is low and the link is complex.
[0004] Therefore, a batch relationship query solution that can improve query performance is now needed. Summary of the invention
[0005] One or more embodiments of the present specification provide a method, device and apparatus for batch query of association relationships to solve the following technical problem: a batch query solution for association relationships that can improve query performance is required.
[0006] To solve the above technical problems, one or more embodiments of this specification are implemented as follows:
[0007] One or more embodiments of the present specification provide a method for batch querying association relationships, the method comprising: obtaining a query data set comprising a plurality of node pairs submitted by a user, the query data set being used to request whether there is an association relationship between nodes in each of the node pairs;
[0008] Determine the graph data to which the plurality of node pairs belong, and determine the coverage of associated parties in the graph data;
[0009] According to the associated party coverage, query the query data set to obtain a first query result;
[0010] According to the first query result, determining, among the multiple node pairs, a node pair that is outside the coverage of the associated party as a candidate node pair;
[0011] Determine whether the nodes in the candidate node pair belong to the same connected component in the graph data, and if so, determine the candidate node pair as the node pair to be predicted;
[0012] Using a pre-trained association relationship prediction model to predict the node pair to be predicted, to obtain a second query result;
[0013] The user is responded to according to the first query result and the second query result.
[0014] One or more embodiments of the present specification provide a batch query device for association relationships, the device comprising:
[0015] An acquisition unit, used to acquire a query data set comprising a plurality of node pairs submitted by a user, wherein the query data set is used to request whether there is an association relationship between nodes in each of the node pairs;
[0016] A first determining unit, configured to determine the graph data to which the plurality of node pairs belong, and determine a coverage range of associated parties in the graph data;
[0017] A query unit, configured to query the query data set according to the associated party coverage to obtain a first query result;
[0018] A second determining unit, configured to determine, according to the first query result, a node pair among the multiple node pairs that is outside the coverage of the associated party as a candidate node pair;
[0019] A judging unit, configured to judge whether the nodes in the candidate node pair belong to the same connected component in the graph data, and if so, determine the candidate node pair as a node pair to be predicted;
[0020] A prediction unit, configured to use a pre-trained association relationship prediction model to predict the pair of nodes to be predicted, and obtain a second query result;
[0021] A returning unit is used to respond to the user according to the first query result and the second query result.
[0022] One or more embodiments of the present specification provide a batch query device for association relationships, including:
[0023] at least one processor; and,
[0024] a memory communicatively connected to the at least one processor; wherein,
[0025] The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to:
[0026] Acquire a query data set comprising a plurality of node pairs submitted by a user, wherein the query data set is used to request whether there is an association relationship between nodes in each of the node pairs;
[0027] Determine the graph data to which the plurality of node pairs belong, and determine the coverage of associated parties in the graph data;
[0028] According to the associated party coverage, query the query data set to obtain a first query result;
[0029] According to the first query result, determining, among the multiple node pairs, a node pair that is outside the coverage of the associated party as a candidate node pair;
[0030] Determine whether the nodes in the candidate node pair belong to the same connected component in the graph data, and if so, determine the candidate node pair as the node pair to be predicted;
[0031] Using a pre-trained association relationship prediction model to predict the node pair to be predicted, to obtain a second query result;
[0032] The user is responded to according to the first query result and the second query result.
[0033] At least one of the above technical solutions adopted in one or more embodiments of this specification can achieve the following beneficial effects:
[0034] By judging whether the node pair belongs to the coverage of the associated party, the first query result is queried, and offline pre-computation is used to determine whether there are associated parties in multiple node pairs, avoiding the problem of long query time. For node pairs that are not covered by the associated party, the connected component is used to judge whether there is a path between the nodes in the node pair, and the node pairs with the same connected component are selected as the node pairs to be predicted, which realizes the screening of data and reduces the computing cost. The node pairs to be predicted are predicted through the relationship prediction model, which improves the recall rate while ensuring the query performance. Based on the combination of offline pre-computation and relationship prediction model, both the query time and the query recall rate are guaranteed, solving the problem of limited recall rate and complex links of the existing method. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0036] Figure 1 A schematic diagram of a method flow of a batch query method for association relationships provided for one or more embodiments of this specification;
[0037] Figure 2 A schematic diagram of a scenario in which a user submits a node pair in an application scenario provided for one or more embodiments of this specification;
[0038] Figure 3 A schematic diagram of the architecture of a batch query system for association relationships in an application scenario provided for one or more embodiments of this specification;
[0039] Figure 4 A schematic diagram of displaying query results in an application scenario provided for one or more embodiments of this specification;
[0040] Figure 5 A schematic diagram of association details of a second query result in an application scenario provided for one or more embodiments of this specification;
[0041] Figure 6 A schematic diagram of the internal structure of a device for batch query of association relationships provided for one or more embodiments of this specification;
[0042] Figure 7 A schematic diagram of the internal structure of a device for batch query of association relationships provided for one or more embodiments of the present specification. DETAILED DESCRIPTION
[0043] The embodiments of this specification provide a method, device and equipment for batch query of association relationships.
[0044] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this application, not all of them. Based on the embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this application.
[0045] Figure 1A flow chart of a batch query method for association relationships provided for one or more embodiments of this specification. The method can be applied to different business fields, such as: in the financial field, by querying the association relationships of enterprises, the ability to prevent risks such as anti-fraud and credit risk control can be improved; in the Internet e-commerce field, intelligent product recommendations can be made to users through the query of association relationships; in the telecommunications field, the prevention of telecommunications fraud of associated users can be achieved through the query of the association relationships of each account; in the industrial field, the management of complex and rapidly changing supply relationships can be achieved through batch query of the association relationships of each node. Therefore, the method can be executed by computing devices in the corresponding fields.
[0046] In one or more embodiments of the present specification, based on offline pre-computation, a relationship prediction model with high accuracy is constructed to query the association relationships between nodes in node pairs uploaded in batches by users, thereby achieving the effect of improving the recall rate while ensuring the query time consumption.
[0047] In one possible implementation method, a single node is used to query the internal node relationship through graph data such as neo4j and JanusGraph. In actual applications, for example, when conducting enterprise risk prevention and control, it is necessary to query the relationship details between any M enterprises and another N enterprises or the relationship details between any K enterprises, that is, it is necessary to query the relationship between enterprises in real time. At this time, if the query is based on the online graph query method, the query consumption of the online graph query depends on the scale of the graph and the correlation between the nodes to be queried, so the time consumption will be significantly increased when performing batch queries. In addition, for batch query requirements, multi-threading needs to be used on the server side for query. For high concurrency scenarios, N is generally limited to <= 10, otherwise the machine thread pool is easily filled. Therefore, online graph queries based on graph databases have certain limitations in query performance and functions in batch query scenarios. In addition, based on the near-line graph computing method, some distributed graph computing systems, such as Spark GraphX based on the pregel framework, can process graph query requests under large-scale data. When a request comes, a graph computing task can be started, and the result will be returned when the task ends. In this asynchronous request mode, although complex queries can be processed, the query waiting time is significantly increased (generally at the minute level), which greatly affects the user experience. Although the query efficiency problem is solved to a certain extent through the "offline pre-computation + real-time query" method, the recall rate is limited and the link is complex. This solution is also committed to solving the problems existing in the above methods and providing a batch relationship query solution that guarantees query efficiency and query recall rate.
[0048] Figure 1 The process in may include the following steps:
[0049] S101: Acquire a query data set including a plurality of node pairs submitted by a user, wherein the query data set is used to request whether there is an association relationship between nodes in each of the node pairs.
[0050] The user submits multiple node pairs that need to be batch queried for association as a query data set. It should be noted that the query data set is used by the user to request whether there is an association between the nodes in each node pair. For example, if you need to query the one-to-one association between M companies and another N companies, then the query data set includes the association between the M company nodes that need to be queried and the multiple node pairs obtained by taking one node from each of the other N company nodes, that is, the query data set includes multiple node pairs consisting of the nodes whose associations need to be queried. Figure 2 The schematic diagram of a scenario in which a user submits a node pair in an application scenario is shown. The user can submit multiple node pairs manually, for example, "A Technology Group Co., Ltd.-A Technology Group Investment Consulting Co., Ltd." is an input node pair. In addition, multiple node pairs can be uploaded in batches by downloading templates based on Excel and other methods. It should also be noted that the association relationship refers to the existence of at least one association path between the nodes in the node pair, and when the strength of the association between them also meets the definition of the business, the nodes in the node pair are associated. For example: Node pair AB is composed of node A and node B, and there is an association path between node A and node B such as "A->(100% shareholding)->C->(50% shareholding)->D->(executive management)->B". If the definition of the business is met between node A and node B, it means that there is an association relationship between node A and node B in node pair AB.
[0051] S102: Determine the graph data to which the multiple node pairs belong, and determine the coverage of associated parties in the graph data.
[0052] As can be seen from the above, although the current online graph query method can easily query the relationship between a single node or several nodes, its query time depends on the scale of the graph and the degree of correlation between the nodes to be queried. When the graph scale is large and the degree of correlation between the nodes to be queried is deep, the query time will increase significantly. Spark GraphX, based on the pregel-like framework, can handle graph query requests under large-scale data. When a request comes, a graph calculation task can be started, and the result will be returned when the task ends. In this asynchronous request mode, although complex queries can be processed, the query waiting time increases significantly.
[0053] Therefore, in order to reduce the time consumption of batch query of association relationships, in one embodiment of the present specification, before obtaining the query data set containing multiple node pairs submitted by the user, offline pre-calculation is required to ensure that the query time consumption meets the requirements, which specifically includes the following steps: First, based on the association relationship to be queried, the determination strategy for the important association relationship in the query process is determined. For example, when a bank evaluates the credit risk of an enterprise, it determines that the determination strategy for the important association relationship is to penetrate the enterprise with a shareholding greater than 5%. At this time, its important association relationship is the enterprise with a shareholding greater than 5%. Because in different application scenarios, the determination strategy for important association relationships may be based on different factors. Therefore, the determination strategy for important association relationships is not specifically limited here. After obtaining the determination strategy for important association relationships in the association relationship, multiple designated nodes belonging to the graph data are determined as reference nodes, so as to determine the nodes in the graph data that have important association relationships with the reference nodes according to the determination strategy as associated nodes. Take each designated node and the associated node of the designated node as an associated party, then all the determined associated parties are the coverage of the associated parties in the graph data. The coverage of related parties is determined based on offline pre-calculation to facilitate subsequent queries within the coverage of related parties, thus avoiding the time-consuming problem of online graph query.
[0054] S103: querying the query data set according to the associated party coverage to obtain a first query result.
[0055] After determining the associated party coverage in the graph data based on the above S102, the query data set is first queried within the associated party coverage to obtain the first query result belonging to the associated party coverage. For example: the associated party coverage includes: node A, node B, node C, node D, node E. If the query data set includes node pair AB, the association relationship between the nodes AB can be determined based on the associated party coverage corresponding to the graph data. However, if a query is performed on node pair AF or node pair FG in the query data set, since node F and node G do not belong to the associated party coverage of the graph data, the association relationship between nodes in node pairs such as AF or node pair FG cannot be determined based on the graph data.
[0056] S104: According to the first query result, determine, among the multiple node pairs, node pairs that are outside the coverage of the associated party as candidate node pairs.
[0057] Currently, the relationship between nodes is calculated based on offline graph computing, and then imported into online storage. During online query, the query method directly queries whether there is an association between node pairs. Although this can ensure query efficiency, the recall rate is low.
[0058] It can be known from the above step S103 that the first query result obtained by querying the coverage of the associated parties of the graph data belonging to multiple node pairs cannot query the node pairs that are not within the coverage of the associated parties. In practical applications, for example, the supply chain nodes in the industrial field are affected by various factors and change rapidly, and the coverage of the associated parties determined offline cannot cover all nodes. Therefore, according to the difference between the query data set and the first query result, that is, the nodes in the node pairs that are outside the coverage of the associated parties in the multiple node pairs, there may also be an association relationship. Therefore, these node pairs that cannot be queried based on the coverage of the associated parties are used as candidate node pairs, so as to query these candidate node pairs in an appropriate manner. It avoids the problem that the query is only based on offline pre-calculation and the recall rate of the query cannot be effectively improved.
[0059] S105: Determine whether the nodes in the candidate node pair belong to the same connected component in the graph data; if so, determine the candidate node pair as the node pair to be predicted.
[0060] Because there may be node pairs with associated relationships in the candidate nodes, in order to improve the recall rate of the query results and reduce the extra time spent on querying unnecessary node pairs, for example: it can be clearly determined that the nodes in the node pair are unreachable, then querying and calculating the associated relationship of the node pair will cause unnecessary computational costs. Therefore, in order to improve the query efficiency on the basis of ensuring the recall rate, in one embodiment of the present specification, by judging whether the nodes in the candidate node pair belong to the same connected component in the graph data, it is determined whether the candidate node needs to be subsequently predicted and calculated. If the nodes in the candidate node pair belong to the same connected component, it means that there is at least one associated path for the shoulder point in the node pair, and it is necessary to perform subsequent prediction queries on the candidate node, so the candidate node pairs belonging to the same connected component are used as predicted node pairs. It should be noted that the connected component is a subgraph data with an associated path in the graph data, for example: there is a path between any two nodes in a subgraph of the graph data G, that is, if any two nodes are reachable, then this subgraph is a connected component of the graph data G.
[0061] S106: using a pre-trained association relationship prediction model to predict the pair of nodes to be predicted, and obtaining a second query result.
[0062] In order to improve the recall rate of batch query results, in one embodiment of the present specification, for the node pairs to be predicted determined based on the above step S105, the node pairs to be predicted are predicted by pre-training to obtain the association relationship prediction model, and the prediction results are obtained as the second query results. The coverage of the associated party is determined by integrating offline calculations, and the first query result is obtained by performing a first query on multiple node pairs within the coverage of the associated party based on the online graph query, thereby ensuring the efficiency of the query. For the node pairs that cannot be queried from the coverage of the associated party, the batch node pairs are queried through the association relationship prediction model, which can ensure the query efficiency while ensuring the accuracy and recall rate.
[0063] S107: Respond to the user according to the first query result and the second query result.
[0064] The first query result and the second query result determined by the above query steps respond to the query submitted by the user. Because the first query result and the second query result may have inconsistent data types, when the data types of the two are inconsistent, the first query result and the second query result can be converted into the data type required by the user before returning to the user. Because different response situations may exist in different application scenarios, the process of returning the specific query result response to the user is not specifically limited here.
[0065] based on Figure 1 This specification also provides some specific implementation plans and extension plans of the method, which will be described below.
[0066] In one or more embodiments of the present specification, in order to improve the efficiency of batch relationship query, before the user performs a batch query, offline calculation is first performed to determine that important related parties are stored in the expected database, so that when the user performs a batch query, a quick response can be made to improve the efficiency of the batch query of the relationship.
[0067] Specifically, for example, in risk analysis and early warning scenarios, it is necessary to mine the related parties of the enterprise and then analyze the needs of related risks. For example, when a bank conducts a credit qualification review on an enterprise, it is necessary to mine and analyze the related enterprises of the enterprise. If the related enterprise exists in the credit blacklist, the bank's standards for credit qualification review of the enterprise will be raised, thereby ensuring the bank's existing rights and interests and reducing unnecessary economic losses. For such inquiries, in order to ensure comprehensive coverage of the query and improve risk control capabilities, a large-scale batch query will be conducted on each enterprise. If only one query is performed based on an online single node at this time, the query efficiency will be too low for a large number of data to be queried, and thus normal business needs cannot be met.
[0068] Therefore, before obtaining the query data set submitted by the user containing multiple node pairs, the method also includes: first obtaining the determination strategy for the important association relationships in the association relationships in various scenarios, wherein it can be understood that the determination strategy for the important association relationships is obvious and easy to judge, such as taking the enterprise shareholding as the determination strategy for the important association relationship. If the shareholding relationship between the enterprises shows that node A holds 100% of the shares of node B and node B holds 90% of the shares of node C, it means that there is an important association relationship between node A, node B, and node C. After determining the determination strategy for the important association relationship, multiple designated nodes belonging to the graph data are determined in the graph data as reference nodes. Thus, according to the determination strategy for the important association relationship, the nodes that belong to the graph data and have an important association relationship with the reference node are determined as association nodes. For example, taking holding shares as the determination strategy of important association relationships, multiple designated nodes A, B, and C in the graph data are determined as reference nodes. If it is determined in the graph data that node A holds 100% of the shares of node D, node A holds 50% of the shares of node B, node B holds 100% of the shares of node E, and node C holds 100% of the shares of node F, then nodes D, E, and F are the associated nodes of reference nodes A, B, and C. The designated nodes and associated nodes are both associated parties, that is, the associated party coverage in the graph data can be determined based on the designated nodes and associated nodes. Continuing with the above example, the designated nodes are A, B, and C, and the associated nodes are D, E, and F. Then the associated parties in the graph data are: nodes A, B, C, D, E, and F, and the associated party coverage is the nodes A, B, C, D, E, and F covered by the associated parties.
[0069] Furthermore, the associated nodes are determined offline in advance. It can be seen from the above that the strategy for determining important associated relationships is obvious and simple, and important associated nodes can be quickly and easily determined in various ways offline. Continuing with the above example, if the strategy for determining important associated relationships is the shareholding situation between companies, then the associated nodes can be determined based on the public information of each company. Offline determination provides strong support for online queries and solves the problem of slow query efficiency during online queries. In one embodiment of the present specification, after determining a node that belongs to the graph data and has the important associated relationship with the reference node as an associated node, the method also includes:
[0070] According to the identifier of the designated node and the identifier of the associated node of the designated node, the query identifier corresponding to each node is determined. At the same time, the association path between the designated node and the associated node is obtained. Then, according to the query identifier and the associated path, storage is performed based on the preset storage structure to obtain a pre-calculated database that can represent the coverage of the associated party. For example: if Hbase is used as the pre-calculated database, the rowkey is the query identifier determined by the identifier of each node. As shown in Table 1 below, this is a storage structure instance table corresponding to an application scenario provided in an embodiment of this specification when Hbase is used as a pre-calculated database.
[0071] Table 1. An example of the storage structure corresponding to an application scenario when Hbase is used as the pre-computation database
[0072]
[0073] As can be seen from Table 1, in this application scenario, Hbase, a highly reliable, high-performance, column-oriented, and scalable distributed storage system, is used as a pre-calculated database. The nodeId corresponding to each node is concatenated as the rowkey row key, i.e., the query identifier, so as to find the corresponding node pair in the Hbase database based on the query identifier. The path_json in the storage structure is used to store the association path information between nodes in order to obtain the coverage range of the covered party. In addition, the minimum degree of association can be a parameter used to screen the strength of the association, and the associated party label can be used to define the characteristics of each node so as to quickly obtain the corresponding association details based on the database.
[0074] In one or more embodiments of the present specification, in order to solve the problem that the query result does not match the expected result due to the user uploading the wrong node pair information, for example: Figure 2 As shown, a node pair to be determined is the association relationship of "A Technology Group Co., Ltd.-A Technology Group Investment Information Co., Ltd.". If the name of the node in the node pair is wrong due to objective factors during the upload process, the returned query result is not the result required for query. Therefore, after the node obtains the query data set containing multiple node pairs submitted by the user, the method also includes the following steps:
[0075] First, obtain a set of standard names that are determined offline in advance, wherein the standard name set stores the standard node names of each node, which can be the standard node names of each node to be queried, obtained based on various public information or various memories. Then match the node names of the nodes in each node pair with the standard node names in the standard name set. If the node names in the node pair do not match the standard node names in the standard name set, it means that there may be errors in the node names in the node pair, and the node names of the nodes in the node pair need to be corrected to ensure the correctness of the query results. The node name correction process here can directly replace the erroneous node names based on the standard name pair, or the node names can be corrected according to the set correction rules as needed, so the correction method is not limited here.
[0076] In one or more embodiments of the present specification, querying a query data set in a predicted database to obtain a first query result specifically includes:
[0077] After obtaining the query data submitted by the user and containing multiple node pairs, in order to facilitate querying the query data in the predicted database, firstly, according to the node identifiers of the nodes in each node pair in the query data set, such as node IDs, a query identifier list corresponding to the query data set is constructed. Then, according to the query identifier list, query extraction is performed in the pre-calculated database, and node pairs corresponding to each query identifier in the pre-calculated database and the query identifier list are obtained as the first query result.
[0078] Further, in one embodiment of the present specification, constructing query data and a query identifier list according to the identifiers of nodes in each node pair includes the following steps: concatenating the node identifiers of two nodes in each node pair to form a query identifier, and then sorting the query identifiers according to a preset sorting method to obtain query data and a query identifier list.
[0079] Furthermore, in order to solve the problem that the recall rate of the first query result obtained by querying the associated party coverage range based only on the graph data cannot meet the requirements, it is also necessary to determine, based on the first query result, nodes outside the associated party coverage range among multiple node pairs in the query data set as candidate nodes, so as to improve the recall rate of the batch query process by continuing to perform query calculations on the candidate nodes.
[0080] Specifically, since the first query structure is obtained based on the nodes determined offline in the estimated database, it belongs to the coverage of the associated party. For example, the coverage of the associated party is nodes A, B, C, D, E, and the query data set contains node pairs AB, AC, AE, BC, BF, EF. Then, based on the query identifier of each node pair in the query data set, only the query results of the node pairs AB, AC, and BC within the coverage range can be obtained as the first query result. However, the node pairs AE, BF, and EF cannot be determined based on the first query result, and at the same time, because the first query result is obtained based on the pre-calculated database determined offline, and the important associated parties in the calculation database are obtained through an obvious and relatively easy-to-determine determination strategy. Therefore, it cannot be ruled out that there are still node pairs that have not been determined in the offline calculation or have a more complex association relationship when the user performs a query, and after obtaining the first query result, subsequent predictions need to be performed to calculate and mine the association relationships of other node pairs. Therefore, in order to improve the recall rate of the query results, in one embodiment of the present specification, after determining the first query result, it is also necessary to determine the node pairs that are outside the coverage of the associated party among multiple node pairs according to the first query result as candidate node pairs. The specific process of determining the candidate node pair is: calculating the difference between the query data set and the first query result, and taking the node pairs included in the difference as candidate node pairs. If the query data set includes multiple node pairs represented by Q1 and the first query result includes node pairs represented by R1, then the candidate node pair is Q1-R1.
[0081] Further, such as Figure 3 As shown in the schematic diagram of the architecture of a batch query system for association relationships in an application scenario, an embodiment of the present description provides that the association relationship includes a graph service and a storage layer in the association query of the graph service. As a functional module in the graph service, the association query consists of four parts: the associated party, the connected component, the associated prediction, and the associated details. The storage layer includes: the Hbase database is used to store the associated party data, the model is not used to train the associated relationship prediction model that meets the requirements to predict the association relationship of the nodes within the calculation node pair, and the GraphDB is a graph database for predicting the associated details of the associated node pair. The batch query process shown in the architecture of the batch query system for association relationships in the application scenario is: first, the node pairs within the coverage range of the associated party are queried based on the Hbase database. After determining the candidate node pairs, the node pairs belonging to the same connected component with a connected path between the nodes are screened out as the node pairs to be predicted by checking whether the candidate node pairs exist in the same connected component. Thereby reducing the number of node pairs that need to be associated predicted, so that the calculation cost is reduced. At the same time, by performing predictive calculations on the nodes to be predicted, the recall rate of batch queries is improved.
[0082] In one or more embodiments of the present specification, in order to ensure both query efficiency and accuracy, recall rate and explainability, a pre-trained association relationship prediction model is used to predict the determined node pairs to be predicted to obtain a second query result.
[0083] Specifically, when querying batch association relationship nodes based on the "offline pre-calculation + real-time query" method, the association relationship between nodes is calculated offline through graph calculation, and then imported into online storage. When querying online, it is directly queried whether the node pairs are associated, so that the query efficiency is guaranteed. However, for large-scale graphs with hundreds of millions of nodes and edges, there is usually a super-large connected subgraph with 10 million nodes. At this time, there are 50 trillion pairs of association relationships in the subgraph alone, and the storage space and data synchronization and update efficiency will limit the final pre-calculated output relationship pairs. Generally, important association relationships of 10 billion to 100 billion scales can be stored through sub-libraries and sub-tables, which can solve the query efficiency problem to a certain extent, but the recall rate of this method is limited and the link is complex. In order to solve the problem of low recall rate and complex links in the method based on offline pre-calculation plus real-time query in the prior art, this manual first obtains the associated party coverage range offline based on an easy important association relationship determination strategy, thereby determining the first query result in the associated party coverage range, and then determining the node pairs belonging to the same connected component in the difference between the query data set and the first query result as the node pairs to be predicted. Based on the pre-trained association relationship prediction model, the association degree of the node pair to be predicted is predicted to see whether it is greater than a preset threshold. If the association degree of the node pair to be predicted is greater than the preset threshold, a second query result is generated based on the node pair to be predicted. By using the pre-trained association relationship prediction model to perform prediction calculations on the node pairs to be predicted, the node pairs with higher association degrees among the node pairs to be predicted are recalled, thereby improving the recall rate of the query.
[0084] In one or more embodiments of the present specification, in order to obtain a satisfactory association relationship prediction model, a pre-trained association relationship prediction model is used, and before predicting the node pair to be predicted, the method further includes the following steps:
[0085] Collect node pairs with association relationships as positive samples. For example, taking enterprise risk control as an example, when it is necessary to find enterprises with association relationships. Positive samples are node pairs AB, AC, BE with share penetration between enterprises, or node pair EF where one enterprise is a subsidiary of another enterprise. Such node pairs with interactive relationships are used as positive samples. After determining the positive samples, select samples that do not belong to the positive sample range and have no association relationship as negative samples. For example, the positive samples determined in the above process are AB, AC, BE, EF, and the coverage range of the positive samples is A, B, C, E, F. Then the negative samples do not belong to the node pairs composed of these nodes and there is no association relationship between the node pairs of the negative samples. Use the positive samples and negative samples as training samples for the association prediction model, determine the sample features corresponding to the training samples, and then train the graph neural network model according to the sample features to obtain a graph neural network model that meets the requirements as the association prediction model. In this process, by selecting positive samples and negative samples, the training process can clearly obtain the recall rate and accuracy of the model output results, which is convenient for selecting a model that matches the business needs as the association prediction model based on the requirements for recall rate and accuracy in different application scenarios.
[0086] Furthermore, a graph neural network model is trained based on sample features to obtain a graph neural network model that meets the requirements as an association relationship prediction model, which specifically includes the following steps:
[0087] First, the sample features of the training samples are constructed. Among them, the sample features include: vector cosine similarity, node in-degree, node out-degree, etc. Then, according to the determined sample features and the preset training strategy, the graph neural network model is trained to learn the node identification to obtain the trained graph neural network model. If it is detected that the trained graph neural network model meets the threshold set based on the performance curve, then the graph neural network model is used as the association relationship prediction model. Among them, it should be noted that the performance curve is used to set the accuracy and recall rate of the graph neural network model.
[0088] In another embodiment of the present specification, the process of obtaining the association relationship prediction model is as follows: first, based on the graph neural network model GNN such as GraphSAGE, GCN, GAT, and Geniepath, learn the embedded identification of the nodes in the graph data. Construct sample features according to the determined positive samples and negative samples, such as vector cosine similarity, node in-degree and out-degree and other features. According to the sample features and the node identification learned by GNN, input into the decision tree model for training to obtain an association relationship model for predicting whether there is an association between nodes. At the same time, a threshold that meets business needs can be set based on the recall rate-accuracy curve of the training model to limit the association relationship prediction model. Through the association relationship prediction model based on GNN, a large number of node pairs with association degrees are filtered out on the basis of the node pairs to be predicted based on the screening of the connected components. In order to facilitate the understanding of the association relationship prediction model, it can be understood as a model of whether the nodes in a node pair to be predicted are associated within N readings, where the size of the N value depends on the random walk or subgraph sampling strategy set for model training.
[0089] like Figure 4 FIG. 1 is a schematic diagram showing a query result display in an application scenario provided by one or more embodiments of this specification. Figure 2 It can be seen that when querying the association relationship of 192 groups of node pairs, the query results show that there are 182 groups with associations, and the association details between nodes in the node pair can be obtained by clicking on different node pairs.
[0090] Therefore, in order to ensure the interpretability of the query results, in one or more embodiments of the present specification, after returning the first query result and the second query result to the user, the method further includes: obtaining association details of the node pairs in the first query result and the second query result.
[0091] Specifically, in one or more embodiments of the present specification, obtaining association details of node pairs in the first query result and the second query result specifically includes the following steps:
[0092] If the node pair belongs to the first query result, then the storage structure instance table of Table 1 above can obtain the associated path corresponding to the node pair based on the pre-calculated database, for example Figure 4An association path of a node pair shown in is: "A Group Technology Group Co., Ltd.-S Venture Capital Co., Ltd.-R Network Technology Co., Ltd.-E Group Equity Investment Partnership". Based on the association path, the association details of the node pair can be determined as "A Group Technology Group Co., Ltd. is 100% of the shares of S Venture Capital Co., Ltd., S Venture Capital Co., Ltd. is a shareholder of R Network Technology Co., Ltd. and holds 31.43% of its shares, and R Network Technology Co., Ltd. is a shareholder of E Group Equity Investment Partnership and holds 8.5% of its shares". Therefore, the share information between the companies constitutes the association details between the nodes in the node pair. If the node pair belongs to the second query result, since the second query result does not belong to the coverage of the associated party, it is necessary to return the association details of the node pair according to the preset graph database.
[0093] Furthermore, if the node pair belongs to the second query result, the association details of the node pair are returned based on the preset graph database, specifically including: obtaining the second query result output by the association relationship prediction model, and determining whether the second query result is associated. Because the node pair predicted based on the association relationship prediction model is the node to be predicted selected only based on the connected component, when the node to be predicted is selected, it belongs to the same connected component, that is, there is a reachable path between the nodes and it is used as the node to be predicted for prediction calculation. As a result, there may be a type of node pair where there is an association path between the two nodes in the node pair, and the depth between the two nodes in the node pair is greater than the preset threshold. Even if the association relationship prediction model predicts that there is an association between the nodes in the node pair, the model prediction result may not be accurate because the node pair is too deep, and thus the association details of the node pair cannot be accurately determined. In order to overcome this problem. When it is determined that the second query result shows that the nodes in the node pair are associated, and the association depth of the second query result is greater than the preset depth threshold, the preset graph database will set the association details of the second query result as suspected associations. If Figure 5 The diagram shows a scenario in which nodes A and B are suspected to be associated with each other in an embodiment of the present specification. By querying the path between any node A and node B based on the preset graph database, the model prediction results are explained, and the query of the association details between nodes with a low association depth is realized. At the same time, for nodes with an association depth greater than a preset threshold, the suspected association is also explained, solving the problem of the explainability of the query results.
[0094] Based on the same idea, one or more embodiments of this specification also provide devices and apparatuses corresponding to the above methods, such as Figure 6 , Figure 7 shown.
[0095] Figure 6 A schematic diagram of a structure of a batch query device for association relationships provided in one or more embodiments of this specification, the device comprising:
[0096] An acquisition unit 601 is used to acquire a query data set comprising a plurality of node pairs submitted by a user, wherein the query data set is used to request whether there is an association relationship between nodes in each of the node pairs;
[0097] A first determining unit 602 is used to determine the graph data to which the plurality of node pairs belong, and determine the coverage of associated parties in the graph data;
[0098] A query unit 603, configured to query the query data set according to the associated party coverage to obtain a first query result;
[0099] A second determining unit 604 is configured to determine, according to the first query result, a node pair among the multiple node pairs that is outside the coverage of the associated party as a candidate node pair;
[0100] A judging unit 605 is used to judge whether the nodes in the candidate node pair belong to the same connected component in the graph data, and if so, determine the candidate node pair as the node pair to be predicted;
[0101] A prediction unit 606 is used to use a pre-trained association relationship prediction model to predict the node pair to be predicted to obtain a second query result;
[0102] The returning unit 607 is configured to return the first query result and the second query result to the user.
[0103] Optionally, the device further includes: a policy acquisition unit, a third determination unit, and a fourth determination unit;
[0104] The strategy acquisition unit is used to acquire a determination strategy for an important association relationship among the association relationships;
[0105] The third determining unit is used to determine a plurality of designated nodes belonging to the graph data as reference nodes;
[0106] The fourth determining unit is used to determine, according to the determination strategy, a node belonging to the graph data and having the important association relationship with the reference node as an associated node;
[0107] The determining of the coverage of the associated parties in the graph data specifically includes:
[0108] Each of the designated nodes and its associated nodes are taken as associated parties, and the associated party coverage in the graph data is determined according to the associated parties.
[0109] Optionally, the associated node is determined offline in advance; the device further comprises: a fifth determination unit, a path acquisition unit, and a storage unit;
[0110] The fifth determining unit is used to determine a corresponding query identifier according to the identifiers of the designated node and its associated node;
[0111] The path acquisition unit is used to acquire the associated path between the designated node and its associated node;
[0112] The storage unit is used to store according to the query identifier and the association path based on a preset storage structure to obtain a pre-calculated database capable of representing the coverage of the associated party;
[0113] The query unit is specifically used for:
[0114] In the pre-calculated database, the query data set is queried to obtain a first query result.
[0115] Optionally, the query unit specifically includes: a construction unit, an identification query unit;
[0116] The construction unit is used to construct a query identifier list of the query data set according to the node identifiers of the nodes within each node pair;
[0117] The identifier query unit is used to query the pre-calculated database based on the query identifier list to obtain node pairs in the pre-calculated database corresponding to each query identifier in the query identifier list as the first query result.
[0118] Optionally, the second determining unit is specifically configured to: calculate a difference set between the query data set and the first query result, and use node pairs included in the difference set as candidate node pairs.
[0119] Optionally, the prediction unit specifically includes: a correlation prediction unit and a generation unit;
[0120] The association prediction unit is used to predict whether the association of the node pair to be predicted is greater than a preset threshold based on a pre-trained association relationship prediction model;
[0121] The generating unit is configured to generate a second query result according to the pair of nodes to be predicted if the correlation degree of the pair of nodes to be predicted is greater than a preset threshold.
[0122] Optionally, the device further comprises: a collection unit, a sample generation unit, and a training unit;
[0123] The collecting unit is used to collect node pairs with an associated relationship as positive samples, and determine samples that do not belong to the coverage of the associated party based on the positive samples, so as to perform random sampling in the samples that do not belong to the coverage of the associated party to obtain negative samples;
[0124] The sample generating unit is used to use the positive sample and the negative sample as training samples of the association relationship prediction model;
[0125] The training unit is used to determine the sample features corresponding to the training samples, so as to train the graph neural network model based on the sample features and obtain a graph neural network model that meets the requirements as the association relationship prediction model.
[0126] Optionally, the training unit specifically includes: a feature construction unit, a model training unit, and a detection unit;
[0127] The feature construction unit is used to construct sample features of the training sample; wherein the sample features include: vector cosine similarity, node in-degree, and node out-degree;
[0128] The model training unit is used to train the graph neural network model based on the sample features and the preset training strategy to obtain a trained graph neural network model;
[0129] The detection unit is used to detect the accuracy and recall rate of the trained graph neural network model, and if it is determined that the graph neural network model meets the threshold set based on the performance curve, the graph neural network model is used as an association relationship prediction model.
[0130] Optionally, the device further comprises: a details acquisition unit;
[0131] The detail acquisition unit is used to acquire association details of the node pairs in the first query result and the second query result;
[0132] The detail acquisition unit specifically includes: a first acquisition unit and a second acquisition unit;
[0133] The first acquisition unit is configured to acquire association details of the node pair based on an association path corresponding to the node pair in the pre-calculated database if the node pair belongs to the first query result;
[0134] The second acquisition unit is used to return the association details of the node pair based on the preset graph database if the node pair belongs to the second query result.
[0135] Optionally, the second acquisition unit specifically includes: a judgment unit and a setting unit;
[0136] The judgment unit is used to obtain the second query result output by the association relationship prediction model, and judge whether the second query result is associated;
[0137] The setting unit is configured to, if there is an association and the association depth of the second query result is greater than a preset depth threshold, cause the threshold graph database to set the association details of the second query result as a suspected association.
[0138] Optionally, the device further comprises: a name acquisition unit, a matching unit, and a correction unit;
[0139] The name acquisition unit is used to acquire a standard name set determined offline in advance; wherein the standard name set stores the standard node name of each node;
[0140] The matching unit is used to match the node name of the node in each node pair with the standard node name in the standard name set;
[0141] The correction unit is used to correct the node name of the node in the node pair if the node name does not match the standard node name.
[0142] Optionally, the node pair is a node pair generated by automatically matching two nodes to be queried based on the nodes to be queried submitted by the user, or a node pair automatically determined based on an uploaded template.
[0143] Figure 7 A schematic diagram of the structure of an IP content library service processing device provided for one or more embodiments of this specification, wherein the device comprises:
[0144] at least one processor 701; and,
[0145] A memory 702 that is communicatively connected to the at least one processor 701; wherein,
[0146] The memory 702 stores instructions that can be executed by the at least one processor 701. The instructions are executed by the at least one processor 701 to enable the at least one processor 701 to:
[0147] Acquire a query data set comprising a plurality of node pairs submitted by a user, wherein the query data set is used to request whether there is an association relationship between nodes in each of the node pairs;
[0148] Determine the graph data to which the plurality of node pairs belong, and determine the coverage of associated parties in the graph data;
[0149] According to the associated party coverage, query the query data set to obtain a first query result;
[0150] According to the first query result, determining, among the multiple node pairs, a node pair that is outside the coverage of the associated party as a candidate node pair;
[0151] Determine whether the nodes in the candidate node pair belong to the same connected component in the graph data, and if so, determine the candidate node pair as the node pair to be predicted;
[0152] Using a pre-trained association relationship prediction model to predict the node pair to be predicted, to obtain a second query result;
[0153] The first query result and the second query result are returned to the user.
[0154] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0155] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0156] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device, equipment, and non-volatile computer storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0157] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0158] The above description is only one or more embodiments of this specification and is not intended to limit this specification. For those skilled in the art, one or more embodiments of this specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of one or more embodiments of this specification shall be included in the scope of the claims of this specification.
Claims
1. A batch query method for association relationships, comprising: Acquire a query data set comprising a plurality of node pairs submitted by a user, wherein the query data set is used to request whether there is an association relationship between nodes in each of the node pairs; Determine the graph data to which the plurality of node pairs belong, and determine the coverage of associated parties in the graph data; According to the associated party coverage, query the query data set to obtain a first query result; According to the first query result, determining, among the multiple node pairs, a node pair that is outside the coverage of the associated party as a candidate node pair; Determine whether the nodes in the candidate node pair belong to the same connected component in the graph data, and if so, determine the candidate node pair as the node pair to be predicted, the connected component being the subgraph data with associated paths in the graph data; Using a pre-trained association relationship prediction model to predict the node pair to be predicted, to obtain a second query result; Respond to the user according to the first query result and the second query result; Before obtaining the query data set comprising a plurality of node pairs submitted by the user, the method further comprises: Obtaining a strategy for determining important associations among associations; Determine a plurality of designated nodes belonging to the graph data as reference nodes; According to the determination strategy, determining a node that belongs to the graph data and has the important association relationship with the reference node as an associated node; The determining of the coverage of the associated parties in the graph data specifically includes: Each of the designated nodes and its associated nodes are taken as associated parties, and the associated party coverage in the graph data is determined according to the associated parties.
2. The method according to claim 1, wherein the associated node is determined offline in advance; After determining the node that belongs to the graph data and has the important association relationship with the reference node as the association node, the method further includes: Determine a corresponding query identifier according to the identifiers of the designated node and its associated node; Obtaining an association path between the specified node and its associated node; According to the query identifier and the association path, storage is performed based on a preset storage structure to obtain a pre-calculated database capable of representing the coverage of the associated parties; Querying the query data set according to the associated party coverage to obtain a first query result specifically includes: In the pre-calculated database, the query data set is queried to obtain a first query result.
3. The method according to claim 2, wherein querying the query data set in the pre-calculated database to obtain the first query result specifically comprises: Constructing a query identifier list of the query data set according to the node identifiers of the nodes within each of the node pairs; A query is performed in the pre-calculated database based on the query identifier list to obtain node pairs in the pre-calculated database corresponding to each query identifier in the query identifier list as a first query result.
4. The method according to claim 1, wherein determining, according to the first query result, a node pair among the plurality of node pairs that is outside the coverage of the associated party as a candidate node pair, specifically comprises: A difference set between the query data set and the first query result is calculated, and the node pairs included in the difference set are used as candidate node pairs.
5. The method according to claim 1, wherein the using a pre-trained association relationship prediction model to predict the pair of nodes to be predicted to obtain the second query result specifically comprises: Based on a pre-trained association relationship prediction model, predict whether the association degree of the node pair to be predicted is greater than a preset threshold; If the correlation degree of the to-be-predicted node pair is greater than a preset threshold, a second query result is generated according to the to-be-predicted node pair.
6. The method according to claim 1, wherein before using the pre-trained association relationship prediction model to predict the node pair to be predicted, the method further comprises: Collecting node pairs with an associated relationship as positive samples, and determining samples that do not belong to the coverage of the associated party based on the positive samples, so as to perform random sampling in the samples that do not belong to the coverage of the associated party to obtain negative samples; Using the positive sample and the negative sample as training samples of the association relationship prediction model; Determine the sample features corresponding to the training samples, train the graph neural network model based on the sample features, and obtain a graph neural network model that meets the requirements as the association relationship prediction model.
7. The method according to claim 6, wherein determining the sample features corresponding to the training samples, training a graph neural network model based on the sample features, and obtaining a graph neural network model that meets the requirements as the association relationship prediction model specifically comprises: Constructing sample features of the training samples; wherein the sample features include: vector cosine similarity, node in-degree, and node out-degree; Based on the sample features and the preset training strategy, the graph neural network model is trained to obtain a trained graph neural network model; If it is detected that the trained graph neural network model meets the threshold set based on the performance curve, the graph neural network model is used as an association relationship prediction model; wherein the performance curve is used to set the accuracy and recall rate of the graph neural network model.
8. The method according to claim 2, after responding to the user according to the first query result and the second query result, the method further comprises: Obtaining association details of node pairs in the first query result and the second query result; The obtaining of association details of the node pairs in the first query result and the second query result specifically includes: If the node pair belongs to the first query result, obtaining association details of the node pair based on the association path corresponding to the node pair in the pre-calculated database; If the node pair belongs to the second query result, the association details of the node pair are returned based on the preset graph database.
9. The method according to claim 8, wherein if the node pair belongs to the second query result, returning the association details of the node pair based on the preset graph database specifically comprises: Obtaining the second query result output by the association relationship prediction model, and determining whether the second query result is associated; If there is an association and the association depth of the second query result is greater than a preset depth threshold, the threshold graph database sets the association details of the second query result as a suspected association.
10. The method according to claim 1, after obtaining the query data set comprising a plurality of node pairs submitted by the user, the method further comprises: Obtaining a standard name set determined offline in advance; wherein the standard name set stores the standard node name of each node; Matching the node names of the nodes within each of the node pairs with the standard node names in the standard name set; If the node name does not match the standard node name, the node name of the node within the node pair is modified.
11. The method according to any one of claims 1 to 10, wherein the node pair is a node pair generated by automatically matching two nodes to be queried based on the nodes to be queried submitted by the user, or a node pair automatically determined based on an uploaded template.
12. A batch query device for association relationships, the device comprising: An acquisition unit, used to acquire a query data set comprising a plurality of node pairs submitted by a user, wherein the query data set is used to request whether there is an association relationship between nodes in each of the node pairs; A first determining unit, configured to determine the graph data to which the plurality of node pairs belong, and determine a coverage range of associated parties in the graph data; A query unit, configured to query the query data set according to the associated party coverage to obtain a first query result; A second determining unit, configured to determine, according to the first query result, a node pair among the multiple node pairs that is outside the coverage of the associated party as a candidate node pair; A judging unit, used to judge whether the nodes in the candidate node pair belong to the same connected component in the graph data, and if so, determine the candidate node pair as a node pair to be predicted, wherein the connected component is a subgraph data with an associated path in the graph data; A prediction unit, configured to use a pre-trained association relationship prediction model to predict the pair of nodes to be predicted, and obtain a second query result; a returning unit, configured to respond to the user according to the first query result and the second query result; The device further comprises: a strategy acquisition unit, a third determination unit, and a fourth determination unit; The strategy acquisition unit is used to acquire a determination strategy for an important association relationship among the association relationships; The third determining unit is used to determine a plurality of designated nodes belonging to the graph data as reference nodes; The fourth determining unit is used to determine, according to the determination strategy, a node belonging to the graph data and having the important association relationship with the reference node as an associated node; The determining of the coverage of the associated parties in the graph data specifically includes: Each of the designated nodes and its associated nodes are taken as associated parties, and the associated party coverage in the graph data is determined according to the associated parties.
13. The device according to claim 12, wherein the associated node is determined offline in advance; the device further comprises: A fifth determining unit, a path acquiring unit, and a storage unit; The fifth determining unit is used to determine a corresponding query identifier according to the identifiers of the designated node and its associated node; The path acquisition unit is used to acquire the associated path between the designated node and its associated node; The storage unit is used to store according to the query identifier and the association path based on a preset storage structure to obtain a pre-calculated database capable of representing the coverage of the associated party; The query unit is specifically used for: In the pre-calculated database, the query data set is queried to obtain a first query result.
14. The device according to claim 13, wherein the query unit specifically comprises: Construction unit, identification query unit; The construction unit is used to construct a query identifier list of the query data set according to the node identifiers of the nodes within each node pair; The identifier query unit is used to query the pre-calculated database based on the query identifier list to obtain node pairs in the pre-calculated database corresponding to each query identifier in the query identifier list as the first query result.
15. The device according to claim 12, wherein the second determining unit is specifically configured to: A difference set between the query data set and the first query result is calculated, and the node pairs included in the difference set are used as candidate node pairs.
16. The device according to claim 12, wherein the prediction unit specifically comprises: Relevance prediction unit and generation unit; The association prediction unit is used to predict whether the association of the node pair to be predicted is greater than a preset threshold based on a pre-trained association relationship prediction model; The generating unit is configured to generate a second query result according to the pair of nodes to be predicted if the correlation degree of the pair of nodes to be predicted is greater than a preset threshold.
17. The apparatus of claim 12, further comprising: Collection unit, sample generation unit, training unit; The collecting unit is used to collect node pairs with an associated relationship as positive samples, and determine samples that do not belong to the coverage of the associated party based on the positive samples, so as to perform random sampling in the samples that do not belong to the coverage of the associated party to obtain negative samples; The sample generating unit is used to use the positive sample and the negative sample as training samples of the association relationship prediction model; The training unit is used to determine the sample features corresponding to the training samples, so as to train the graph neural network model based on the sample features and obtain a graph neural network model that meets the requirements as the association relationship prediction model.
18. The apparatus according to claim 17, wherein the training unit specifically comprises: Feature construction unit, model training unit, detection unit; The feature construction unit is used to construct sample features of the training sample; wherein the sample features include: vector cosine similarity, node in-degree, and node out-degree; The model training unit is used to train the graph neural network model based on the sample features and the preset training strategy to obtain a trained graph neural network model; The detection unit is used to detect the accuracy and recall rate of the trained graph neural network model, and if it is determined that the graph neural network model meets the threshold set based on the performance curve, the graph neural network model is used as an association relationship prediction model.
19. The apparatus of claim 13, further comprising: Details acquisition unit; The detail acquisition unit is used to acquire association details of the node pairs in the first query result and the second query result; The detail acquisition unit specifically includes: a first acquisition unit and a second acquisition unit; The first acquisition unit is configured to acquire association details of the node pair based on an association path corresponding to the node pair in the pre-calculated database if the node pair belongs to the first query result; The second acquisition unit is used to return the association details of the node pair based on the preset graph database if the node pair belongs to the second query result.
20. The device according to claim 19, wherein the second acquisition unit specifically comprises: Judgment unit, setting unit; The judgment unit is used to obtain the second query result output by the association relationship prediction model, and judge whether the second query result is associated; The setting unit is configured to, if there is an association and the association depth of the second query result is greater than a preset depth threshold, cause the threshold graph database to set the association details of the second query result as a suspected association.
21. The apparatus of claim 12, further comprising: Name acquisition unit, matching unit, correction unit; The name acquisition unit is used to acquire a standard name set determined offline in advance; wherein the standard name set stores the standard node name of each node; The matching unit is used to match the node name of the node in each node pair with the standard node name in the standard name set; The correction unit is used to correct the node name of the node in the node pair if the node name does not match the standard node name.
22. The device as described in any one of claims 12-21, wherein the node pair is a node pair generated by automatically matching two nodes to be queried based on the nodes to be queried submitted by the user, or a node pair automatically determined based on an uploaded template.
23. A batch query device for association relationships, the device comprising: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to: Acquire a query data set comprising a plurality of node pairs submitted by a user, wherein the query data set is used to request whether there is an association relationship between nodes in each of the node pairs; Determine the graph data to which the plurality of node pairs belong, and determine the coverage of associated parties in the graph data; According to the associated party coverage, query the query data set to obtain a first query result; According to the first query result, determining, among the multiple node pairs, a node pair that is outside the coverage of the associated party as a candidate node pair; Determine whether the nodes in the candidate node pair belong to the same connected component in the graph data, and if so, determine the candidate node pair as the node pair to be predicted, the connected component being the subgraph data with associated paths in the graph data; Using a pre-trained association relationship prediction model to predict the node pair to be predicted, to obtain a second query result; Respond to the user according to the first query result and the second query result; Before obtaining the query data set containing multiple node pairs submitted by the user, the following is further performed: Obtaining a strategy for determining important associations among associations; Determine a plurality of designated nodes belonging to the graph data as reference nodes; According to the determination strategy, determining a node that belongs to the graph data and has the important association relationship with the reference node as an associated node; The determining of the coverage of the associated parties in the graph data specifically includes: Each of the designated nodes and its associated nodes are taken as associated parties, and the associated party coverage in the graph data is determined according to the associated parties.
Citation Information
Patent Citations
Method and terminal for inquiring associated information
CN109753590A
Enterprise map generation method and device, computer equipment and storage medium
CN109800335A