A person name disambiguation method and device based on equity relationship graph and personnel information
By utilizing equity relationship graphs and personnel information in financial data, the ambiguity problem of nodes with the same name is solved, efficient and accurate name disambiguation is achieved, the workload of training set screening is reduced, and more information is utilized.
Patent Information
- Application Number
- CN202410693583.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-31
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-05-31
AI Technical Summary
When dealing with the problem of name disambiguation in financial data, existing technologies have problems such as large workload in training set screening, limited model applicability, failure to effectively utilize personnel information and disconnected nodes in equity relationship graphs, resulting in difficulty in disambiguation.
A name disambiguation method based on equity relationship graph and personnel information is adopted. By reading the business license, resume and position data of the enterprise, an initial equity relationship graph is constructed. The similarity is calculated by using resume matching, connected component segmentation, graph neural network and personnel position data to realize the similarity judgment and merging of nodes with the same name.
It reduces the number of model training steps, uses personnel information to improve computing efficiency and accuracy, solves the problem of disambiguation of disconnected nodes, and achieves faster and more effective name disambiguation.
Smart Images

Figure CN118821755B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of personal name disambiguation, and particularly relates to a personal name disambiguation method and device based on a stock right relation graph and personnel information. BACKGROUND
[0002] Personal name disambiguation aims to eliminate the ambiguity of personal names in different environments, classify the same personal names according to different entities in the real world, and thus effectively organize and cluster information and provide the information to users. In financial data, the data of a person node often only contains a natural person's name and does not contain information for uniquely identifying an identity, which causes that when the same personal name appears in different enterprises, it cannot be determined whether the same personal name is the same person. The occurrence of this situation will cause many problems. For example, when drawing a correlation graph, if the natural persons in different enterprise information cannot be determined to be the same person, the graph will not be merged, and the factual correlation information between different enterprises cannot be established. For another example, in the case where it is not determined whether the two people with the same name are the same person, different data are merged recklessly, which may cause errors in the construction of a correlation network.
[0003] Existing methods often use a machine learning algorithm model to construct a training set and train an artificial intelligence model to determine the similarity of the same personal name. However, when facing a large amount of financial stock data, the workload of screening an effective training set is extremely large, which greatly causes resource waste; meanwhile, the artificial intelligence model obtained by training is often only applicable to the field data of the training set, and has limited performance in actual use; in addition, in actual financial field data, a company can not only obtain a stock right relation graph, but also provide more effective attribute information such as a person's resume and service condition for part of the person nodes, but since the attribute information cannot guarantee to cover each person node, and such data are often independent of the stock right relation graph, there are difficulties in matching to the person nodes, and the previous method does not effectively utilize this part of information. In addition, the current personal name disambiguation method based on enterprise relations is highly dependent on the comprehensiveness of enterprise relations, and when the nodes to be disambiguated are not connected in the stock right relation graph, that is, not connected, effective personal name disambiguation cannot be performed. Therefore, there is an urgent need for a personal name disambiguation method that does not require complex training, can utilize rich personnel information, and can still perform disambiguation when the nodes to be compared are not connected, to more quickly and effectively solve the problem of personal name disambiguation in financial data.
[0004] The nodes of the stock right relation graph are abstract representations of enterprise or individual entities, and the edges are abstract representations of stock right relations. The stock right relation graph represents a network composed of stock right relations.
[0005] Personnel information includes personnel resume information and personnel position information. Neither of them can cover all natural person node data. The resume information describes the resume content of the natural person, and the position information abstractly represents the natural person's position in the enterprise in the form of edge data. Summary of the Invention
[0006] The technical problem to be solved by the present invention is the ambiguity problem of multiple nodes with the same name in the enterprise equity relationship diagram. The similarity of two existing nodes is judged and based on the similarity results, it is decided whether to merge and optimize the individual nodes with the same name.
[0007] Based on the above technical problems, the present invention adopts the following technical solutions:
[0008] A method for disambiguating names based on an equity relationship diagram and personnel information includes the following steps:
[0009] Step 1: Read the business data, resume data, and employee position data, and store them in a resilient distributed data structure. Based on the business data, obtain the initial node and edge datasets and construct the initial equity relationship graph. Read the IDs of the natural person node pairs to be compared.
[0010] Step 2: Perform resume matching and resume feature similarity calculation on the resume data read in step 1, and match the resume information of personnel that does not cover all natural person nodes in the resume data with the natural person nodes of the initial equity relationship graph; perform connected component segmentation on the initial equity relationship graph constructed in step 1, and divide the initial equity relationship graph into multiple independent equity relationship subgraphs;
[0011] Step 3: Based on the equity relationship subgraph obtained in step 2, determine whether the natural person node pairs to be compared read in step 1 belong to the same connected component; if the natural person node pairs to be compared belong to the same connected component, execute step 4; otherwise, execute step 5;
[0012] Step 4: Calculate the natural person similarity based on the natural person nodes to be compared, the common neighbors in the equity relationship subgraph, and the structural similarity of the equity relationship subgraph, and perform name disambiguation based on the natural person similarity and resume similarity;
[0013] Step 5: Based on the personnel position data in step 1, calculate the proportion of duplicate names of senior executives of enterprises for comparison, and perform name disambiguation based on the duplicate name proportion of senior executives and the similarity of resumes.
[0014] Furthermore, in step 1, the node data set includes a node ID, a node name, and a node type, and the node data structure is: [node ID, node name, node type], where the node ID includes a company node ID and a natural person node ID, and the node type includes a company node and a natural person node;
[0015] The edge data set includes a source node ID, a destination node ID, and a holding ratio of the source node to the destination node; and the edge data structure is: [source node ID, destination node ID, holding ratio of the source node to the destination node];
[0016] The resume data includes a resume natural person ID, an employed company ID, a natural person name, and resume information; and the resume data structure is: [resume natural person ID, employed company ID, natural person name, resume information];
[0017] The personnel employment data includes a natural person ID, an employed company ID, and a position; and the personnel employment data structure is: [natural person node ID, employed company ID, position].
[0018] Further, in step 2, the resume matching and resume feature similarity calculation process is specifically:
[0019] The holder information of the company node in the equity relationship graph is obtained through the equity penetration algorithm, and the result data of the equity penetration is preprocessed as [company node ID, natural person node ID, natural person name]; the [employed company ID, natural person name] information extracted from the resume data is used to find the corresponding result matched therewith in the result data of the equity penetration;
[0020] If there is no matching item or the matching result is more than one item, the resume matching fails; if there is exactly one matching item, it is considered that the matching is successful, and [resume natural person ID, natural person node ID] is recorded as the basis for subsequent name disambiguation;
[0021] The resume data format is adjusted to [natural person node ID, natural person name, resume information] after the matching is successful, and the pre-trained natural language model is used to solve the embedding representation of the resume and calculate the cosine similarity as the resume similarity.
[0022] Further, the step 4 includes the following sub-steps:
[0023] Step 4.1: based on the constructed initial equity relationship graph, the equity relationship sub-graphs of the natural person nodes to be compared and their respective neighbors within the specified hop number are obtained, the common neighbors of the natural person nodes to be compared are obtained, and the common neighbor proportion is further calculated;
[0024] Step 4.2: for each node in the equity relationship sub-graph, the node initial feature value is assigned; for each node in the equity relationship graph, the node name in the equity relationship graph is used as the input of the language model, and a 1024-dimensional feature vector is obtained as the node feature value;
[0025] Step 4.3: After obtaining the node eigenvalues, use the graph neural network to aggregate the structural features of the nodes in the graph to obtain a new eigenvector for each node. The cosine similarity of the two eigenvectors corresponding to the natural person node pair to be compared is calculated as the feature similarity of the two nodes, that is, the structural similarity of the equity relationship subgraph;
[0026] Step 4.4: For the case where both of the two people have resumes, calculate the following similarities: the feature similarity of the two resumes calculated in step 2; the feature similarity of the two nodes after structural information aggregation in step 4.3; the similarity converted based on the proportion of common neighbors in step 4.1, and the weighted average of the three similarities according to different weights to obtain the final similarity; For the case where both of the two people do not have resumes, calculate the following similarities: the feature similarity of the two nodes after structural information aggregation in step 4.3; the similarity converted based on the proportion of common neighbors in step 4.1, and the weighted average of the two similarities according to different weights to obtain the final similarity;
[0027] Step 4.5: If the final similarity exceeds the set threshold, the natural person node pair to be compared is the same person, and the nodes are merged. If the similarity does not exceed the set threshold, the natural person node pair to be compared is not the same person.
[0028] Furthermore, obtaining the common neighbors of the natural person nodes to be compared in step 4.1 includes: adopting a node iteration method, aggregating the bidirectional neighbor information of adjacent nodes each time the nodes are aggregated, and when reaching the specified hop count loop, directly taking out the neighbor array data of the target node to obtain the neighbors within the specified hop count; solving the neighbors within the specified hop count for the natural person node pairs to be compared respectively, comparing them, and obtaining common neighbors. All neighbors within the specified hop count plus the node to be compared itself are the set of all nodes in the subgraph within the specified number of nodes to be compared, and the edge information in which both the target node and the source node are in the point set is the five-hop subgraph edge set.
[0029] Furthermore, in step 4.3, a three-layer GraphSAGE graph neural network structure is constructed, and the dimension changes from 1024 to 1024 to 768 and then to 512. In addition, nonlinear activation functions are not used in the neural network calculation to better retain the structural information around the nodes.
[0030] Furthermore, in step 4.4, the resume feature similarity and the feature vector similarity of the node to be compared after aggregation by the graph neural network are calculated using cosine similarity, and the common neighbor similarity is converted into similarity by the proportion of common neighbors to all neighbors.
[0031] Furthermore, step 5 includes the following sub-steps:
[0032] Step 5.1: Query the respective job companies of the to-be-compared node pairs in the personnel position data, and respectively summarize them into two sets;
[0033] Step 5.2: Based on the two sets in step 5.1, match the same-name executives in the job companies of the two persons in a combined form, and record the matching results;
[0034] Step 5.3: Based on the matching results in step 5.2, solve the proportion of the number of executives with the same name to the total number of executives in the company, and judge based on a threshold of 50%. If there are two job companies with the same name accounting for more than 50%, output a Boolean true value and directly enter step 5.4, otherwise, output a Boolean false value and continue the calculation of the remaining combination cases;
[0035] Step 5.4: For the resume feature similarity result, if the similarity is greater than 75%, output a Boolean true value, if the similarity is less than 75% or the resume matching fails, output a Boolean false value. Further, if one of the resume similarity Boolean value and the executive matching Boolean value is true, it is considered that the two natural person nodes given are the same person, and the node is merged, otherwise, it is considered that the two natural person nodes given are not the same person.
[0036] Further, in step 5.2, the combination is to traverse the two sets respectively and generate to-be-compared enterprise pairs. For each to-be-compared enterprise pair, directly enter step 5.3. If a Boolean true value has been output, end the traversal operation of the two sets, otherwise, repeat step 5.2 after calculation to continue calculating the next to-be-compared enterprise pair.
[0037] On the other hand, the present application provides a person name disambiguation device based on stock ownership relationship graph and personnel information, comprising: a data reading module, a resume matching module, a graph data construction module, a connected component segmentation module, a connected point pair disambiguation judgment module and a non-connected point pair disambiguation judgment module;
[0038] The data reading module is used to read enterprise business data, store the enterprise business data into an elastic distributed data structure, obtain initial vertex data and initial edge data, and read the IDs of the to-be-compared natural person node pairs;
[0039] The graph data construction module is used to construct initial stock ownership relationship graph data based on the initial vertex data and the initial edge data;
[0040] The resume matching module is used to read the resume data for resume matching and resume feature similarity calculation, and match the personnel resume information of all natural person nodes not covered in the resume data with the natural person nodes of the initial stock ownership relationship graph;
[0041] Connected component segmentation module: It is used to segment the constructed initial equity relationship graph into connected components, dividing the initial equity relationship graph into multiple independent equity relationship subgraphs; and judging whether the natural person node pairs to be compared read in belong to the same connected component;
[0042] Connected point pair disambiguation judgment module: It is used to calculate the natural person similarity based on the common neighbors in the equity relationship graph and the structural similarity of the equity relationship subgraph of the natural person nodes to be compared, and perform name disambiguation based on the natural person similarity and resume similarity;
[0043] The non-connected point pair disambiguation judgment module is used to calculate the proportion of duplicate names of corporate executives of natural persons based on personnel employment data, and to perform name disambiguation based on the proportion of duplicate names of corporate executives and the similarity of resumes.
[0044] Compared with the prior art, the present invention has the following beneficial effects:
[0045] The present invention adopts a name disambiguation method based on equity relationship graph and personnel information. When dealing with the name disambiguation problem, it reduces the targeted model training steps and applies more valuable personnel information, solving the pain point that name disambiguation cannot be performed due to the lack of connectivity of the nodes to be disambiguated in the equity relationship graph, thereby improving computing efficiency and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0047] Figure 1 It is a flowchart of the name disambiguation method based on the equity relationship diagram and personnel information proposed by the present invention.
[0048] Figure 2 This is a flowchart of resume matching in step 2.
[0049] Figure 3 This is a schematic diagram of splitting the connected subgraph in step 3.
[0050] Figure 4 is the pseudocode of the algorithm for calculating common neighbors in step 4.1.
[0051] Figure 5 This is a schematic diagram of aggregating graph structural features through graph neural network in step 4.3.
[0052] Figure 6Schematic diagram of data for executive matching in the process of name disambiguation for non-connected nodes.
[0053] Figure 7 It is a structural diagram of the name disambiguation device based on the equity relationship diagram and personnel information proposed by the present invention.
[0054] Figure 8 Schematic diagram of the structure of the connected point pair disambiguation judgment module proposed in the present invention.
[0055] Figure 9 Schematic diagram of the structure of the non-connected point pair disambiguation judgment module proposed in the present invention. DETAILED DESCRIPTION
[0056] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0057] Example 1
[0058] See also Figure 1 The name disambiguation method based on the equity relationship diagram and personnel information proposed in the present invention includes the following steps.
[0059] Step 1: Read the business data, resume data, and employee position data, and store them in a resilient distributed data structure. Based on the business data, obtain the initial node and edge datasets and construct the initial equity relationship graph data. Read the IDs of the natural person node pairs to be compared.
[0060] Step 2: Perform resume matching and resume feature similarity calculation on the resume data read in step 1, and match the resume information of personnel that does not cover all natural person nodes in the resume data with the natural person nodes in the equity relationship graph; perform connected component segmentation on the initial equity relationship graph constructed in step 1, and divide the initial equity relationship graph into multiple independent equity relationship subgraphs;
[0061] In step 1, the node dataset contains node ID, node name and node type. The node data structure is: [node ID, node name, node type]. The node ID includes company node ID and natural person node ID.
[0062] The edge data set contains the source node ID, the destination node ID, and the source node's shareholding ratio to the destination node; the edge data structure is: [source node ID, destination node ID, source node's shareholding ratio to the destination node];
[0063] The resume data includes the resume person ID, company ID, name of the person, and resume information. The resume data structure is: [resume person ID, company ID, name of the person, resume information];
[0064] Personnel employment data includes natural person ID, company ID, and position; the personnel employment data structure is: [natural person node ID, company ID, position].
[0065] See also Figure 2 , the resume matching in step 2 further includes the following steps.
[0066] Step 2.1: Obtain shareholder information for company nodes in the equity relationship graph using the equity penetration algorithm. Preprocess the resulting data into [company node ID, natural person node ID, natural person name]; extract [employee company ID, natural person name] from the resume data.
[0067] Step 2.2 uses the [company ID, natural person name] information extracted from the resume data to search for corresponding matching results in the equity penetration result data; note that based on actual data, it is found that the person ID in the resume information and the person ID in the equity relationship diagram are often difficult to unify, so they are represented by the resume natural person ID and the natural person node ID respectively.
[0068] Step 2.2: Information matching: Extract [Company ID, Name of Individual] from the resume and use this information to search for matching results in the equity relationship.
[0069] Step 2.3: Matching result judgment. If there is no matching item or the matching result is more than one, the resume information cannot be mapped to a specific natural person node, and the resume matching fails; if there is exactly one matching item, the match is considered successful.
[0070] At this point, after statistics on the company data, it was found that there were relatively few shareholders (larger shareholders) with the same name in the same company. Therefore, after matching the resumes, it was considered that the natural person ID in the resume and the natural person node ID in the equity relationship diagram actually described the same person. Therefore, [resume natural person ID, natural person node ID] was recorded as the basis for subsequent name disambiguation.
[0071] Step 2.4: Adjust the resume format. Adjust the resume format to [natural person node ID, natural person name, resume information]. Use the pre-trained natural language model to solve the embedding representation of each resume and calculate the cosine similarity as the resume similarity.
[0072] Step 2.5: If the match is successful, output the result.
[0073] See also Figure 3The equity connected component extraction method in step 2 can extract connected components from the complete equity graph data. The connected component sets are stored separately to form multiple equity relationship subgraphs. For example, the equity relationship graph structure on the left side of the figure is divided into connected component 1, connected component 2, and connected component 3 according to connectivity after subgraph segmentation.
[0074] Step 3: Based on the connected components solved in step 2, determine whether the natural person node pairs to be compared read in step 1 belong to the same connected component; if the natural person node pairs to be compared belong to the same connected component, execute step 4; otherwise, execute step 5;
[0075] Step 4: Calculate the natural person similarity based on the natural person nodes to be compared, the common neighbors in the equity relationship graph, the structural similarity of the equity relationship subgraph, and the resume similarity, and perform name disambiguation based on the natural person similarity;
[0076] See also Figure 4 The common neighbor solving algorithm proposed in this invention adopts a node iteration method, aggregating the bidirectional neighbor information of adjacent nodes each time a node is aggregated. When the specified hop count loop is reached, the neighbor array data of the target node is directly retrieved to obtain the neighbors within the specified hop count. The neighbors within the specified hop count are solved for the two target nodes respectively, and then compared to obtain the common neighbors. All neighbors within the specified hop count plus the node to be solved itself constitute the complete node set of the subgraph of the specified hop count of the node to be solved, and the edge information of both the target node and the source node in the point set constitutes the edge set of the subgraph of the specified hop count of the node to be solved. In this example, a five-hop subgraph is used.
[0077] Step 5: Based on the personnel position data in step 1, calculate the proportion of duplicate names of senior executives of enterprises for comparison, and perform name disambiguation based on the duplicate name proportion of senior executives and the similarity of resumes.
[0078] If the node pair to be compared belongs to the same connected component, the following steps are further performed:
[0079] Step 4.1: Solve the specified node equity relationship subgraph and common neighbors. For the constructed equity relationship graph, solve the equity subgraph of the specified two nodes and the respective neighbors within the specified number of hops, and further obtain the common neighbors of the two specified nodes and further calculate the common neighbor ratio. The common neighbor ratio takes the number of common neighbors of the compared node pair as the numerator and the total number of neighbors of the compared node pair as the denominator. The specific conversion process is as follows: for any one compared node pair, there are two cases for their common neighbors, one is that there is no common neighbor, at this time the common neighbor similarity is set to 0.5; the other is that there are common neighbors, at this time the specific similarity of the two compared nodes can be analyzed according to the proportion of the number of common neighbors to the total number of neighbors. After calculating the common neighbors, it is still necessary to quantitatively represent, after the actual test and analysis of the equity relationship graph, the proportion of common neighbors to all neighbors is divided into four intervals, which are 0% to 10%, 10% to 50%, 50% to 80% and 80% to 100%. The similarity of the above four intervals is respectively assigned as 0.7, 0.8, 0.9 and 1.
[0080] Step 4.2: Equity relationship subgraph node feature assignment. Assign initial node features to each node in the equity relationship subgraph. For each node in the equity relationship graph, use the node name [natural person name, company name] as the input of the language model to obtain a 1024-dimensional feature vector as the node feature value.
[0081] Step 4.3: Graph neural network aggregates subgraph structure information. After obtaining the node features, use graph neural network to aggregate the structure features of the nodes in the graph. A three-layer GraphSAGE graph neural network structure is constructed, with dimensions changing from 1024 to 1024 to 768 to 512. In addition, no nonlinear activation function is used in neural network calculation to better preserve the structure information around the node. Run the graph neural network algorithm on the equity relationship subgraph to aggregate the structure information, obtain a new feature vector for each node, and solve the cosine similarity of the two feature vectors corresponding to the compared natural person node pair as the feature similarity of the two nodes, that is, the equity relationship subgraph structure similarity.
[0082] Please refer to Figure 5The Graph Neural Network aggregates the graph structure features, adopts the GraphSAGE network layer as the basic structure of the constructed graph neural network, and first assigns an initial node feature value to each node in the equity relationship subgraph. For all nodes in the specified hop subgraph, the name of the node [name of a natural person, name of a company] is used as the input of the language model to obtain a 1024-dimensional feature vector as the node feature value. After obtaining the feature value of the node, the subgraph structure with the feature vector is input into a three-layer GraphSAGE graph neural network trained by structure data, and a new node feature vector with aggregated structure information is obtained. The graph neural network is composed of an input layer, a hidden layer, and an output layer. In this example, the prediction object is two nodes with the same name. Since the full equity relationship graph is often too large, and there is a lot of data information irrelevant to the prediction node, the initial equity relationship subgraph is used as the input, and the feature vector of each node after message aggregation is obtained after the neighbor node message aggregation in the hidden layer.
[0083] Step 4.4: Natural person similarity calculation. For the case where both given persons have resumes, the following similarities are calculated: the similarity Sim1 between the two resume feature vectors calculated in step 2; the similarity Sim2 between the two nodes after structure information aggregation in step 4.3; and the similarity Sim3 converted from the proportion of common neighbors of the two nodes in step 4.1. The three similarities are weighted and averaged according to the weights of 50%, 15%, and 35% to obtain the final similarity. For other cases, only Sim2 and Sim3 are calculated, and the two similarities are weighted and averaged according to the weights of 30% and 70% to obtain the final similarity.
[0084] Step 4.5: Connected point disambiguation judgment. For the final similarity, if the similarity exceeds the threshold value 0.8, it is considered that the two given natural person nodes are the same person, and the nodes are merged. If the similarity does not exceed the threshold value, it is considered that the two given natural person nodes are not the same person.
[0085] When the node pair to be compared does not belong to the same connected component, the following steps are performed:
[0086] Step 5.1: Inquiring the company where the person serves. For the read-in node pair information, the companies where the two nodes serve are inquired in the personnel service information, and the companies are respectively summarized into two sets.
[0087] Step 5.2: Matching the executives in the companies where the person serves. The same-name executives in the companies where the two persons serve are matched in a combined form, and the matching results are recorded.
[0088] Step 5.3: Determine the percentage of executives with duplicate names. Based on the matching results obtained previously, determine the percentage of executives with duplicate names relative to the total number of executives in the company. Use 50% as the threshold for this determination. If there are two companies where the percentage of duplicate names exceeds 50%, output a Boolean true value and proceed directly to step 4.4'. Otherwise, output a Boolean false value and continue the calculation for the remaining combinations.
[0089] Step 5.4: Disambiguation of Disconnected Nodes. For the resume feature similarity results, if the similarity is greater than 75%, a Boolean true value is output. If the similarity is less than 75% or the match fails, a Boolean false value is output. Furthermore, if either the resume similarity Boolean value or the executive match Boolean value is true, the two given natural person nodes are considered to be the same person and the nodes are merged. Otherwise, the two given natural person nodes are considered to be different persons.
[0090] See also Figure 6 The example of the judgment of the disambiguation judgment of the non-connected points proposed by the present invention is shown in the figure. If there are two companies A and B, and the two companies have 3 and 4 executives' employment information respectively, when performing name disambiguation on Huang Mouqing, when comparing the employment information of companies A and B, Company A has Huang Mouqing with id 2 and Huang Moushan with id 3. Company B has Huang Mouqing with id 6 and Huang Moushan with id 7. At this time, both companies have an executive named Huang Mouqing and an executive named Huang Moushan, and the number of executives with the same name is 2, which is more than half of the executives of Company B (1.5). At this time, it is believed that the main shareholders of the two companies are similar, and then it is judged that Huang Mouqing with id 2 and Huang Mouqing with id 6 are the same natural person.
[0091] Example 2
[0092] See also Figure 7 The name disambiguation device based on the equity relationship graph and personnel information proposed in the present invention includes a data reading module M1, a resume matching module M2, a graph data construction module M3, a connected component segmentation module M4, a connected point pair disambiguation judgment module M6, and a non-connected point pair disambiguation judgment module M5.
[0093] The data reading module M1 can read in the cleaned and pre-processed business data of enterprises and some characteristic data of natural persons, and generate node data, edge data and natural person characteristic data for constructing graph data respectively; the node data must include enterprise nodes and shareholder nodes, enterprise nodes include sole proprietorships, limited liability companies or joint-stock companies, and shareholder nodes include natural persons and others; resume data includes natural person ID, natural person name and natural person resume content; personnel position data includes natural person ID, company ID and position status.
[0094] The partial feature matching module M2 can use the company shareholder data and resume data to map the resume data of the personnel containing only the name of the natural person to a specific natural person node.
[0095] The graph data construction module M3 uses node data and edge data to construct an equity relationship graph.
[0096] The connected component segmentation module M4 takes the equity relationship graph as input, divides the equity relationship graph into multiple mutually disconnected equity relationship connected subgraphs through connected component extraction, and at the same time determines whether the node pairs to be compared belong to the same connected component.
[0097] The non-connected node pair disambiguation judgment module M5 is used to solve the situation when the node pairs to be compared do not belong to the same connected component.
[0098] The connected point pair disambiguation judgment module M6 is used to resolve the situation when the node pairs to be compared belong to the same connected component.
[0099] See also Figure 8 The connected point pair disambiguation judgment module M6 proposed in the present invention includes 5 sub-modules, among which the five-hop subgraph and common neighbor calculation module M6-1 takes the equity relationship graph as input, and outputs the point data, edge data and common neighbors of the five-hop subgraph of the specified two nodes through a program written in Spark. The node feature module M6-2 applies a pre-trained natural language processing module to convert the feature information into a corresponding 1024-dimensional feature vector and assigns it to the matching natural person node. For nodes without feature information, the name is assigned to the node as a feature vector. The graph structure information aggregation module M6-3 uses a graph neural network model, takes the initial equity relationship graph as input, aggregates the graph structure information, and generates a feature vector that aggregates the result information. The similarity calculation module M6-4 performs cosine similarity calculation on the feature representation vector and feature vector of the two natural person nodes to be compared, and converts the common neighbors into similarity. Finally, similarity aggregation is performed to obtain the similarity between the two specified nodes. Name disambiguation judgment module M6-5 determines whether to perform name disambiguation based on the final similarity. If the similarity exceeds a threshold, the two natural person nodes are considered to be the same person and the nodes are merged. If the similarity does not exceed the threshold, the two natural person nodes are considered to be different persons.
[0100] See also Figure 9The non-connected point pair disambiguation judgment module M5 contains three sub-modules, wherein the incumbent enterprise query module M5-1 queries the respective incumbent enterprises of the read-in node pair information in the personnel incumbent information, and respectively induces the respective incumbent enterprises into two sets. The incumbent enterprise senior manager matching and judgment module M5-2 matches the senior managers with the same name in the respective incumbent enterprises of the two persons in a combined form, and solves the matching result, solves the proportion of the number of senior managers with the same name in the total number of senior managers of the enterprise, and judges whether the proportion is more than 50% as a threshold. If there are two incumbent enterprises with the same name and the proportion is more than 50%, a Boolean true value is output, otherwise, a Boolean false value is output. The non-connected point pair disambiguation judgment module M5-3. For the resume feature similarity result, if the similarity is greater than 75%, a Boolean true value is output, if the similarity is less than 75% or the matching fails, a Boolean false value is output. Further, if the resume similarity Boolean value and the enterprise senior manager matching Boolean value are true, it is considered that the two natural person nodes are the same person, and the node is merged. Otherwise, return to the module M5-2 to continue to calculate the next group of enterprises. If all the enterprise combinations have been traversed, it is considered that the two natural person nodes are not the same person.
[0101] The present application proposes a person name disambiguation method based on a stock ownership relationship graph and personnel information. The method does not need to collect a data set for training when deployed, and an effective person name disambiguation module is designed for the difficult-to-judge case of non-connected nodes in the relationship graph. In addition, personnel information is applied to simplify the deployment steps, increase the dimension of consideration for natural persons, simplify the process, and improve the calculation efficiency. The device can be deployed in a single machine environment for simple and small volume person name disambiguation tasks, or in a distributed environment for massive data person name disambiguation. The device exhibits good performance in single machine and cluster environments.
[0102] It should be understood that parts not described in detail in the specification are all prior art.
[0103] It should be understood that the above description of the preferred embodiments is more detailed, and therefore should not be considered as limiting the scope of patent protection of the present application. It is not necessary or possible to enumerate all embodiments here. Those skilled in the art can make substitutions or modifications without departing from the scope of the claims of the present application, and all fall within the scope of protection of the present application. The scope of protection of the present application should be subject to the appended claims.
Claims
1. A name disambiguation method based on equity relationship diagram and personnel information, characterized in that: The steps include: Step 1: Read the business data, resume data, and employee position data, and store them in a resilient distributed data structure. Based on the business data, obtain the initial node and edge datasets and construct the initial equity relationship graph. Read the IDs of the natural person node pairs to be compared. Step 2: Perform resume matching and resume feature similarity calculation on the resume data read in step 1, and match the resume information of personnel that does not cover all natural person nodes in the resume data with the natural person nodes of the initial equity relationship graph; perform connected component segmentation on the initial equity relationship graph constructed in step 1, and divide the initial equity relationship graph into multiple independent equity relationship subgraphs; Step 3: Based on the equity relationship subgraph obtained in step 2, determine whether the natural person node pairs to be compared read in step 1 belong to the same connected component; if the natural person node pairs to be compared belong to the same connected component, execute step 4; Otherwise, go to step 5; Step 4: Calculate the natural person similarity based on the natural person nodes to be compared, the common neighbors in the equity relationship subgraph, and the structural similarity of the equity relationship subgraph, and perform name disambiguation based on the natural person similarity and resume similarity; Step 5: Based on the personnel position data in step 1, calculate the proportion of duplicate names of senior executives of enterprises for comparison, and perform name disambiguation based on the duplicate name proportion of senior executives and the similarity of resumes.
2. The name disambiguation method based on the equity relationship diagram and personnel information according to claim 1 is characterized in that: In step 1, the node data set includes node ID, node name and node type. The node data structure is: [node ID, node name, node type]. Node ID includes company node ID and natural person node ID. Node type includes company node and natural person node. The edge data set contains the source node ID, the destination node ID, and the source node's shareholding ratio to the destination node; the edge data structure is: [source node ID, destination node ID, source node's shareholding ratio to the destination node]; The resume data includes the resume person ID, company ID, name of the person, and resume information. The resume data structure is: [resume person ID, company ID, name of the person, resume information]; Personnel employment data includes natural person ID, company ID, and position; the personnel employment data structure is: [natural person node ID, company ID, position].
3. The name disambiguation method based on the equity relationship diagram and personnel information according to claim 1 is characterized in that: In step 2, the resume matching and resume feature similarity calculation process is specifically as follows: The shareholder information of the company node in the equity relationship diagram is obtained through the equity penetration algorithm, and the result data of the equity penetration is preprocessed into [company node ID, natural person node ID, natural person name]. [Working company ID, natural person name] is extracted from the resume data, and the [working company ID, natural person name] information extracted from the resume data is used to find the corresponding matching results in the equity penetration result data. If there is no match or more than one match is found, the resume matching fails. If there is exactly one match, the match is considered successful, and [resume natural person ID, natural person node ID] is recorded as the basis for subsequent name disambiguation. If the match is successful, the resume data format is adjusted to [natural person node ID, natural person name, resume information], and the pre-trained natural language model is used to solve the embedding representation of the resume and calculate the cosine similarity as the resume similarity.
4. The name disambiguation method based on the equity relationship diagram and personnel information according to claim 1 is characterized in that: The step 4 includes the following sub-steps: Step 4.1: Based on the constructed initial equity relationship graph, obtain the equity relationship subgraph of the natural person node to be compared and their respective neighbors within the specified number of hops, obtain the common neighbors of the natural person node to be compared and further calculate the common neighbor ratio; Step 4.2: Assign initial features to each node in the equity relationship subgraph. For each node in the equity relationship graph, use the node name in the equity relationship graph as the input of the language model to obtain a 1024-dimensional feature vector as the node feature value; Step 4.3: After obtaining the node eigenvalues, use the graph neural network to aggregate the structural features of the nodes in the graph to obtain a new eigenvector for each node. The cosine similarity of the two eigenvectors corresponding to the natural person node pair to be compared is calculated as the feature similarity of the two nodes, that is, the structural similarity of the equity relationship subgraph; Step 4.4: For the case where both of the two people have resumes, calculate the following similarities: the similarity of the two resume features calculated in step 2; the similarity of the two node features after structural information aggregation in step 4.3; based on the conversion of the common neighbor ratio in step 4.1 into similarity, and weighted average the three similarities according to different weights to obtain the final similarity; for the case where both of the two people do not have resumes, calculate the following similarities: the similarity of the two node features after structural information aggregation in step 4.3; Based on the conversion of the common neighbor ratio in step 4.1 into similarity, the two similarities are weighted averaged according to different weights to obtain the final similarity; Step 4.5: If the final similarity exceeds the set threshold, the natural person node pair to be compared is the same person, and the nodes are merged. If the similarity does not exceed the set threshold, the natural person node pair to be compared is not the same person.
5. The method for disambiguating names based on equity relationship diagram and personnel information according to claim 4 is characterized in that: Obtaining the common neighbors of the natural person nodes to be compared in step 4.1 includes: adopting a node iteration method, aggregating the bidirectional neighbor information of adjacent nodes each time the nodes are aggregated, and when reaching the specified hop count loop, directly taking out the neighbor array data of the target node to obtain the neighbors within the specified hop count; solving the neighbors within the specified hop count for the natural person node pairs to be compared respectively, comparing them, and obtaining common neighbors. All neighbors within the specified hop count plus the node to be compared itself are the set of all nodes in the subgraph within the specified number of nodes to be compared, and the edge information of both the target node and the source node in the point set is the edge set of the five-hop subgraph.
6. The method for disambiguating names based on equity relationship diagram and personnel information according to claim 4 is characterized in that: In step 4.3, a three-layer GraphSAGE graph neural network structure is constructed, and the dimension changes from 1024 to 1024 to 768 and then to 512. In addition, no nonlinear activation function is used in the neural network calculation to better retain the structural information around the node.
7. The method for disambiguating names based on equity relationship diagram and personnel information according to claim 4 is characterized in that: In step 4.4, the resume feature similarity and the feature vector similarity of the node to be compared after aggregation of the graph neural network are calculated using cosine similarity, and the common neighbor similarity is converted into similarity by the proportion of common neighbors to all neighbors.
8. The method for disambiguating names based on equity relationship diagram and personnel information according to claim 1, characterized in that: The step 5 includes the following sub-steps: Step 5.1: Query the respective companies of the nodes to be compared in the personnel employment data, and summarize them into two sets respectively; Step 5.2: Based on the two sets in step 5.1, match the executives with the same name in the companies where the two people work in a combined manner and record the matching results; Step 5.3: Based on the matching results from Step 5.2, calculate the ratio of executives with duplicate names to the total number of executives in the company, using 50% as the threshold. If there are two companies where the percentage of duplicate names exceeds 50%, output a Boolean true value and proceed directly to Step 5.
4. Otherwise, output a Boolean false value and continue the calculation for the remaining combinations. Step 5.4: For the resume feature similarity results, if the similarity is greater than 75%, a Boolean true value is output; if the similarity is less than 75% or the resume matching fails, a Boolean false value is output. Furthermore, if one of the resume similarity Boolean value and the corporate executive matching Boolean value is true, the given two natural person nodes are considered to be the same person, and the nodes are merged; otherwise, the given two natural person nodes are considered to be not the same person.
9. The name disambiguation method based on the equity relationship diagram and personnel information according to claim 8 is characterized in that: In step 5.2, the combination method is to traverse the two sets separately and generate enterprise pairs to be compared. For each group of enterprise pairs to be compared, step 5.3 is directly entered. If a Boolean true value has been output, the traversal operation of the two sets is terminated. Otherwise, step 5.2 is repeated after calculation to continue calculating the next group of enterprise pairs to be compared.
10. A name disambiguation device based on equity relationship diagram and personnel information, characterized in that: include: Data reading module, resume matching module, graph data construction module, connected component segmentation module, connected point pair disambiguation judgment module and disconnected point pair disambiguation judgment module; Data reading module: It is used to read the business data of enterprises, store the business data in a resilient distributed data structure, and obtain the initial vertex data and initial edge data; Read in the ID of the natural person node pair to be compared; A graph data construction module, which is used to construct initial equity relationship graph data based on the initial vertex data and initial edge data; Resume matching module: It is used to read resume data for resume matching and resume feature similarity calculation, and match the resume information of personnel that does not cover all natural person nodes in the resume data with the natural person nodes in the initial equity relationship diagram; Connected component segmentation module: It is used to segment the constructed initial equity relationship graph into connected components, dividing the initial equity relationship graph into multiple independent equity relationship subgraphs; and judging whether the natural person node pairs to be compared read in belong to the same connected component; Connected point pair disambiguation judgment module: It is used to calculate the natural person similarity based on the common neighbors in the equity relationship graph and the structural similarity of the equity relationship subgraph of the natural person nodes to be compared, and perform name disambiguation based on the natural person similarity and resume similarity; A non-connected point pair disambiguation judgment module is used to calculate the percentage of duplicate names of senior executives of companies that are compared with natural persons based on their employment data, and to perform name disambiguation based on the percentage of duplicate names of senior executives and the similarity of their resumes; The device for disambiguating names based on a shareholding relationship diagram and personnel information is used to execute the steps of the method for disambiguating names based on a shareholding relationship diagram and personnel information as described in any one of claims 1-9.
Citation Information
Patent Citations
Industrial and commercial senior management name disambiguation method based on enterprise association relationship
CN110020433A
Name disambiguation method and system based on enterprise association relationship
CN113326377A