People-enterprise relationship mining method based on graph representation technology and word segmentation technology

Through the human-enterprise relationship mining method based on chart representation technology and word segmentation technology, the problem of difficulty in building a human-enterprise relationship database in the existing technology is solved, efficient and accurate disambiguation of personal names of industrial and commercial executives is achieved, and the construction and analysis of enterprise relationship databases in multiple application scenarios is supported.

CN120336398APending Publication Date: 2025-07-18ANHUI CNBI SOFTWARE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510399000.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The lack of personnel identity information in the existing enterprise and industrial and commercial public data has led to the inability to accurately identify human-enterprise relationships, which hinders the construction of the human-enterprise relationship database and the effective management of commercial business.

Method used

The human-enterprise relationship mining method based on chart characterization technology and word segmentation technology is adopted to mine enterprise relationship relationships through the information of equity chains, branch chains and guarantee chains. Combined with chart characterization learning and natural language processing technology, the enterprise relationship and brand information are extracted, and the duplicate persons are disambiguated. The graph attention model and the data structure are used to collect the results.

Benefits of technology

It has achieved efficient and accurate disambiguation of personal names of industrial and commercial executives, and has a high comprehensive and accurate detection rate. It can process large-scale data in limited resources and in a short period of time. It supports the construction of human-enterprise relationship databases and correlation map analysis in multiple application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336398A_ABST
    Figure CN120336398A_ABST
Patent Text Reader

Abstract

The invention discloses a human-enterprise relationship mining method based on a graph representation technology and a word segmentation technology, and the method comprises the following steps: S1, carrying out the mining of an enterprise association relationship based on the information of an equity chain, a branch chain and a guarantee chain, and mining the enterprise association relationship through the features of the structure of an enterprise association graph, carrying out name disambiguation on the persons with the duplication names based on the mined association relationship to obtain a name disambiguation result R1 of the persons with the duplication names; s2, based on company names in a preset data source, brand names are extracted by adopting a natural language processing technology, enterprise association relationships are mined by using the brand names, name disambiguation is carried out, and a name disambiguation result R2 is obtained. The method can efficiently disambiguate the business high-level management names, has a high disambiguation recall ratio and precision ratio, can define a similarity function and a weight function in a personalized manner, has certain flexibility, can meet business high-level management name disambiguation in multiple application scenes, and provides support for subsequent enterprise relationship library construction and enterprise association map analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of personal name disambiguation, and particularly to a method for mining the relationship between individuals and enterprises based on graph representation technology and word segmentation technology. Background Art

[0002] Personal name disambiguation aims to eliminate the ambiguity of personal names in different environments, and its main task is to identify the real person corresponding to the personal name based on the personal name and other information.

[0003] Existing enterprise and industrial and commercial publicity data often only involve personal names for personnel, without the identity information of the corresponding personnel. This makes it impossible for us to accurately identify the relationship between individuals and enterprises, thus hindering the construction of the person-enterprise relationship database and further posing challenges to commercial operations such as supply chain management and risk access screening.

[0004] Due to the huge number of enterprises and personnel involved, existing personal name disambiguation methods are difficult to efficiently and accurately complete the task of personnel identification. Summary of the Invention

[0005] The technical problem to be solved by the present invention is: how to solve the problem that the existing security protection system has a single protection type, resulting in poor protection effect and having a certain impact on the use of the security protection system, and provides a method for mining the relationship between individuals and enterprises based on graph representation technology and word segmentation technology.

[0006] The present invention solves the above technical problem through the following technical solutions. The present invention includes the following steps:

[0007] S1, mining enterprise association relationships based on the information of equity chains, branch chains, and guarantee chains, mainly using features of the structure of enterprise association graphs such as weighted path distance, the number of common neighbors, and the similarity of embedding vectors based on graph representation learning to mine enterprise association relationships, and performing personal name disambiguation for homonymous personnel based on the discovered association relationships to obtain the personal name disambiguation result R1 of homonymous personnel.

[0008] The equity chain refers to the chain relationship formed by the shareholding relationships among enterprises. For example, if enterprise A holds shares in enterprise B, and enterprise B holds shares in enterprise C and enterprise D, then there are two investment chains: enterprise A -> enterprise B -> enterprise C, and enterprise A -> enterprise B -> enterprise D. The equity chain contains information on a company's external investments and information on parent companies / subsidiaries, etc. The theoretical basis for using the equity chain to help disambiguate people with the same name is that the same person is more likely to hold positions or be a shareholder in two enterprises with a shareholding relationship. For example, a middle-level manager in the parent company also serves as a senior executive in the subsidiary; indirectly holding / controlling a company for accounting or other reasons; or an individual following an investment in an industry. According to Bayes' principle, under the condition that other conditions remain unchanged, a higher prior probability means a higher posterior probability, that is, the people with the same name in two enterprises with a shareholding relationship are more likely to correspond to the same person.

[0009] The branch chain refers to the chain relationship between an enterprise and its branches. For example, the Anhui branch of xx company is a branch of the xx head office, and the Hefei branch of xx company is a branch of the Anhui branch of xx company. Then the head office -> Anhui branch -> Hefei branch constitutes a branch chain. The main difference between a branch company and a subsidiary is that it has no independent legal personality and does not involve equity, which can be used as a supplement to the equity chain.

[0010] The guarantee chain refers to the guarantee relationship among enterprises. When an enterprise conducts financing, such as applying for a bank loan, a guarantee behavior often occurs. There are three involved parties in the guarantee behavior, namely the creditor, the debtor, and the guarantor. When the debtor is unable to fulfill the debt, the guarantor should fulfill the debt. Among them, the creditor is often a financial institution such as a bank or an industrial investment fund, the debtor is often the borrowing company, and the guarantee company undertakes the risk that the borrowing company cannot fulfill the debt and often has a relatively close relationship with the borrowing company. Therefore, the information on the guarantee / guaranteed chain is helpful for us to disambiguate people with the same name.

[0011] A heterogeneous graph refers to a graph that contains different types of nodes and edges. A heterogeneous graph can be defined in mathematical terms as: G = (V, E, φ, φ), where V is the set of nodes, E is the set of edges, φ is the mapping from nodes to their attributes, and φ is the mapping from edges to their attributes. Specifically, the nodes here represent companies, and the edges represent the relationships between companies, including shareholding / being shareheld, branch company / superior company, and guarantee / being guaranteed relationships.

[0012] The weighted path distance refers to the weights of all the edges on the path between two nodes on a heterogeneous graph. Here, the weight of an edge is the mapping φ from the edge to its attributes. For those with a shorter weighted distance, their relationship is closer in the external relationship graph of the company. For the people with the same name in these companies, there is a greater possibility that they are the same person.

[0013] Common neighbors: For a node, its neighbors can be defined as all the points within a distance of d. Having more common neighbors between two nodes means a closer connection between the corresponding two companies. Then, for the people with the same name in these two companies, there is a greater possibility that they are the same person.

[0014] Graph representation learning and embedding vectors. Graph representation learning refers to the process of obtaining the embedding vectors (representation vectors) of nodes / edges in a graph according to the information of the graph structure and corresponding algorithms based on model assumptions. Then, classification, clustering, link prediction, etc. tasks are performed on the nodes in the graph according to the representation vectors. Here, we use the node2vec algorithm to obtain the embedding representation of the graph structure in the enterprise relationship graph, and through feature engineering methods, extract multi-dimensional features of the company, and combine with the graph attention model (GAT, Graph Attention Networks) to fuse the graph structure features of the enterprise relationship graph and the features of the enterprise itself. Finally, we perform clustering based on the fused feature representation to mine the implicit relationships between enterprises and perform name disambiguation based on this implicit relationship.

[0015] The core formula of the graph attention model is as follows

[0016] e ij = LeakReLU(a T [Wh i ||Wh j ) ;

[0017]

[0018] h′ i = σ(∑ j∈N(i) α ij Wh j ) ;

[0019] hi and hj represent the hidden variables input to the model, and W is the mapping matrix to be trained, which is used to map the input hidden variables to the space for attention weighting;

[0020] Whi and Whj are the representations of companies i and j in this space respectively, and a T [Whi||Whj] is the product weighting of the concatenation result of vector a and Whi, Whj, which is used to fuse the information of these two aspects;

[0021] αij represents the similarity between i and j, and it is used for attention weighting to obtain h′ i , that is, the hidden representation of the second layer of the model. Similarly, the hidden representation of the nth layer is obtained as the output of the model;

[0022] Taking into account the information of the above three dimensions, that is, e ij 、α ij and h′i , disambiguate the persons with the same name;

[0023] For two related companies, a linear weighted scoring function f is designed:

[0024] f(u, v) = w1x1(u, v) + w2x2(u, v) + w3x3(u, v)

[0025] where x1 is used to measure the weighted distance. The shorter the weighted distance, the larger x1;

[0026] x1 is used to measure the common neighbors. The more common neighbors, the larger x2. x3 is used to measure the similarity of the feature vectors. The more similar the feature vectors, the larger x3. Here, the similarity is based on cosine similarity ;

[0027] u represents the hidden vector representation of the first company, v represents the hidden vector representation of the second company, wi represents the importance weight of the i-th feature. u and v are obtained through model training, and wi is a preset value; when f is greater than the threshold, it is determined that the persons with the same name involved in the two companies are the same.

[0028] S2. Based on the company names in the preset data source, use natural language processing technology to extract the brand names, use the brand names to mine the enterprise association relationships, perform name disambiguation, and obtain the name disambiguation result R2.

[0029] Adopt the method of word segmentation + filtering words + rule design to extract the brand information from the company names. The brand information here refers to the identifying words in the company names. For example, the brand name of "Xiaomi Technology Co., Ltd." is "Xiaomi", and the brand name of "Anhui Jingbang Software Technology Co., Ltd." is "Jingbang". For companies with similar brand names, we believe that the persons with the same name in these companies are more likely to be the same person.

[0030] Word segmentation means splitting a paragraph into multiple independent words. For example, the word segmentation result of "Xiaomi Technology Co., Ltd." may be ["Xiaomi", "Technology", "Limited Liability", "Company"], and the word segmentation result of "Anhui Jingbang Software Technology Co., Ltd." may be ["Anhui", "Jingbang", "Software Technology", "Limited Company"]. Then, through heuristic rules, we select the words that are most likely to be the company brand names from the word segmentation results.

[0031] In particular, for self-employed individuals, considering that the possibility of cross-regional operation is smaller, under the condition of considering similar brand names, it is further required that the geographical locations of the enterprises are the same.

[0032] S3. Merge the previous name disambiguation results R1, R2, and R3 to form the final name disambiguation result R3.

[0033] The merging here means that in the disambiguation result R1, (p1, p2) is regarded as the same person, in the disambiguation result R2, (p2, p3) is regarded as the same person, and in the disambiguation result R3, (p3, p4) is regarded as the same person. Then we should regard (p1, p2, p3, p4) as the same person.

[0034] Use the data structure of the union-find set to help us efficiently complete the merging task.

[0035] The present invention has the following advantages compared with the prior art: The method for mining the person-enterprise relationship based on the graph representation technology and the word segmentation technology obtains the enterprise (node) representation vectors based on the graph structure and the representation vectors based on the enterprise attributes in the enterprise shareholding, branch, and guarantee graphs through the graph representation technology, and uses a graph neural network to fuse the two to obtain new representation vectors. Using the representation vectors and other attributes in the graph, calculate the similarity of the company, and perform name disambiguation based on the similarity. Use the word segmentation technology in natural language processing to design heuristic rules, extract the corresponding brand information according to the company name, and perform disambiguation based on the brand information. In addition, supplement the disambiguation results with publicly available information such as company email addresses and phone numbers. At the same time, an efficient implementation method of the disambiguation algorithm is designed, which can complete the disambiguation task of large-scale data within limited computing resources and a short time, realize efficient disambiguation of the names of industrial and commercial executives, have a high recall rate and precision rate for disambiguation, and can customize the similarity function and weight function, with a certain degree of flexibility, and can meet the disambiguation of the names of industrial and commercial executives in multiple application scenarios, providing support for the subsequent construction of the person-enterprise relationship database and the analysis of the person-enterprise association graph. Brief Description of the Drawings

[0036] Figure 1 It is the overall flowchart of the present invention. Detailed Embodiment

[0037] The following details the embodiments of the present invention. These embodiments are implemented on the premise of the technical solution of the present invention, and provide detailed implementation manners and specific operation processes. However, the protection scope of the present invention is not limited to the following embodiments.

[0038] As Figure 1 shown, this embodiment provides a technical solution: A method for mining the person-enterprise relationship based on the graph representation technology and the word segmentation technology, including:

[0039] S1. Mine the enterprise association relationships based on the information of the equity chain, branch chain, and guarantee chain. Mainly use the features of the structure of the enterprise association graph, such as the weighted path distance, the number of common neighbors, and the similarity of the embedding vectors based on graph representation learning, on heterogeneous graphs to mine the enterprise association relationships, and perform disambiguation of the names of people with the same name based on the discovered association relationships to obtain the disambiguation result R1 of the names of people with the same name.

[0040] S1.1 Extract the data of equity relationships, branch relationships, and guarantee relationships from the database and clean out the valid data.

[0041] S1.2 Perform feature encoding on information such as company names and addresses.

[0042] S1.3 Establish the structure of the enterprise association graph and use the node2vec algorithm to encode the graph structure.

[0043] Process the enterprise association graph. Based on the in-degree, out-degree, node attributes, and information of common neighbors of the nodes in the graph, reconstruct a new enterprise association graph for the input of the subsequent graph attention model.

[0044] The processing method based on the in-degree information of the nodes in the graph is mainly to exclude the nodes with a large in-degree, which corresponds to the situation where a company is held by many other companies. The possibility of accidental associations in these associations is relatively large, which may lead to more misidentifications. Therefore, we exclude these situations. In particular, we set the threshold in_degree_threhold and only retain the nodes with an in-degree less than or equal to in_degree_threhold.

[0045] The processing method based on the out-degree information of the nodes in the graph is mainly to remove the nodes with a large in-out degree, which corresponds to the situation where a company invests in many other companies. The possibility of accidental associations in these associations is relatively large, which may lead to more misidentifications. Therefore, we exclude these situations. In particular, we set the threshold out_degree_threhold and only retain the nodes with an out-degree less than or equal to out_degree_threhold.

[0046] The processing method based on the node attribute information in the graph mainly excludes some special types of companies. For companies whose names include keywords such as "bank", "securities", "asset management", "insurance", etc., and companies whose business scope includes "fund sales", etc., since they have many weak associated investments with other companies and are not suitable for using the corresponding information for name disambiguation of individuals, we exclude the outgoing edges of the nodes corresponding to these companies in the enterprise association graph. For companies of the partnership type, we perform penetration graph construction (for example, if there is a chain in the original enterprise association graph: ordinary company A -> partnership company B -> partnership company C -> ordinary company D, then there is a corresponding chain in the new graph: ordinary company A -> ordinary company D). This is because in practice, there is a high possibility of controlling through multiple layers of partnership companies, and penetration graph construction is conducive to discovering such strong association relationships.

[0047] The processing method based on common neighbor information is mainly to exclude some accidental associations. For an investment relationship (i.e., a directed edge in the enterprise association graph), let the investing company be A. If the invested company B is also invested by another company C, unless there is another company D such that both A and B invest in D, we do not retain this investment relationship.

[0048] S1.4 Fuse the information feature encoding and structure encoding as the input of the graph attention model, and train the graph attention model based on the annotation information. Obtain the company representation vector based on the graph attention model;

[0049] The core formula of the graph attention model is as follows:

[0050] e ij =LeakReLU(a T [Wh i ||Wh j );

[0051]

[0052] h′ i =σ(∑ j∈N(i) α ij Wh j );

[0053] hi and hj represent the hidden variables input to the model, and W is the mapping matrix to be trained, which is used to map the input hidden variables to the space for attention weighting;

[0054] Whi and Whj are the representations of companies i and j in this space respectively, and a T [Whi||Whj] is the product weighting of the concatenation result of vector a and Whi, Whj, which is used to fuse the information of these two aspects;

[0055] αij represents the similarity between i and j, and it is used for attention weighting to obtain h′ i , that is, the hidden representation of the second layer of the model. Similarly, the hidden representation of the nth layer is obtained as the output of the model;

[0056] Taking into account the information in the above three dimensions, that is, e ij 、α ij and h′ i , to disambiguate the people with the same name;

[0057] For two related companies, a linear weighted scoring function f is designed:

[0058] f(u, v) = w1x1(u, v) + w2x2(u, v) + w3x3(u, v)

[0059] where x1 is used to measure the weighted distance. The shorter the weighted distance, the larger x1 is;

[0060] x1 is used to measure the common neighbors. The more common neighbors, the larger x2 is. x3 is used to measure the similarity of the representation vectors. The more similar the representation vectors are, the larger x3 is. The similarity here is based on cosine similarity ;

[0061] u represents the hidden vector representation of the first company, v represents the hidden vector representation of the second company, wi represents the importance weight of the ith feature, u and v are obtained through model training, and wi is a preset value;

[0062] When f is greater than the threshold, it is determined that the people with the same name involved in the two companies are the same one.

[0063] S1.5 Establish an enterprise personnel association graph. Select the central node u on the graph and perform a breadth-first search starting from this central node. Let the searched node be v. When f(u, v) reaches the set threshold, traverse all the people with the same name corresponding to u and v, and disambiguate the people with the same name among them.

[0064] S2. Based on the company names in the preset data source, use natural language processing technology to extract brand names, use the brand names to mine enterprise association relationships, perform name disambiguation, and obtain the name disambiguation result R2.

[0065] S2.1 Extract the company names and company category data from the database and clean the valid data;

[0066] The screening of company valid data is multi-dimensional and is divided into the following aspects:

[0067] There is a use_flag column in the data source, and only the data with use_flag = 0 can be valid data.

[0068] There is an "is_personal" column in the data source, which represents whether the shareholder is a company or a natural person. We only extract the data of companies.

[0069] There is a "stock_name_id" column in the data source. If the shareholder is a company, its value is the corresponding company ID. We extract the non-empty data of this column.

[0070] There is a "use_flag" column in the data source. Only when "use_flag = 0" can the data be valid.

[0071] There may be outliers in the company names. We filter out those with a company name length within a certain range and without special characters as valid data.

[0072] S2.2, Use the word segmentation technology to segment the company names and design a heuristic algorithm to extract the brand names.

[0073] S2.2.1, The heuristic algorithm mainly includes the following steps: First, establish a stop word list, such as countries like "China" and "UK", and common place names like "Beijing" and "Shanghai". Second, scan the word segmentation results from left to right and skip when a stop word is scanned. Third, when the current word length is greater than or equal to 2, take the current word as the brand name. Fourth, when the current word length is 1, concatenate the current word and the next word as the brand name.

[0074] S2.2.2, For brand / company names that are difficult to identify by the heuristic rules, use the large language model to identify and extract them.

[0075] S2.3, For non-self-employed individuals, merge the companies with similar brands to obtain the merge result.

[0076] S2.4, For self-employed individuals, merge the companies with similar brands and close regions to obtain the merge result.

[0077] S2.5, For the merged companies (related companies), traverse the involved duplicate names and disambiguate the duplicate names among them.

[0078] S4, Merge the previous name disambiguation results R1, R2, R3.1, and R3.2 to form the final name disambiguation result R4.

[0079] S4.1, Merging each disambiguation result is a problem of calculating the connected components on the graph. We use the data structure of the union-find set to efficiently complete this task.

[0080] S4.2, The union-find set is a data structure used to handle the merging and query problems of some disjoint sets. It mainly supports two operations: Union: Merge two sets into a new set.

[0081] Find: Determine whether two elements are in the same set. Both of these operations are very efficient (close to constant time complexity).

[0082] Due to the huge data scale, we cannot read all the data into memory for processing at once. Therefore, we choose to process the data in batches according to the names of people. The correctness of this approach lies in the fact that different names of people can definitely not be disambiguated as the same person.

[0083] S4.3. Establish a union-find set, read the disambiguation records of duplicate names in R1, R2, and R3 in sequence, and merge the sets where the corresponding duplicate names are located.

[0084] S4.4. Re-output the merged union-find set in the format of disambiguation records of duplicate names.

[0085] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0086] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0087] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for mining the human-enterprise relationship based on graph representation technology and word segmentation technology, characterized in that, It includes the following steps: S1: Mine the enterprise association relationships based on the information of the equity chain, branch chain, and guarantee chain, use the characteristics of the structure of the enterprise association graph to mine the enterprise association relationships, and perform disambiguation of the names of people with the same name based on the discovered association relationships to obtain the disambiguation result R1 of the names of people with the same name; S2. Based on the company names in the preset data source, use natural language processing technology to extract the brand names, use the brand names to mine the enterprise association relationships, and perform disambiguation of names to obtain the disambiguation result R2 of names; S4. Merge the previous disambiguation results R1 and R2 of names to form the final disambiguation result R3 of names.

2. The human-enterprise relationship mining method based on graph characterization technology and word segmentation technology according to claim 1, wherein: The characteristics of the structure of the enterprise association graph include: the weighted path distance on the heterogeneous graph, the number of common neighbors, and the similarity of the embedding vectors based on graph representation learning; The heterogeneous graph is a graph containing different types of nodes and edges. The mathematical language definition of the heterogeneous graph is: G = (V, E, φ, φ), where V is the set of nodes, E is the set of edges, φ is the mapping from nodes to their attributes, and φ is the mapping from edges to their attributes; V represents companies, and E represents the relationships between companies, including shareholding / being shareheld, branch company / superior company, and guarantee / being guaranteed relationships. The weighted path distance refers to the weights of all the edges on the path between two nodes on the heterogeneous graph; For a node, its neighbors are defined as all the nodes within a distance d from it, where d is a preset value; Graph representation learning refers to the process of obtaining the embedding vectors of nodes / edges in the graph according to the information of the graph structure and the model assumptions using a preset algorithm; Then, perform classification, clustering, and link prediction tasks on the nodes in the graph according to the representation vectors; That is, use the node2vec algorithm to obtain the embedding representation of the graph structure in the enterprise relationship graph, extract the multi-dimensional features of the company through feature engineering methods, and combine the graph attention model to fuse the structural features of the enterprise relationship graph and the characteristics of the enterprise itself; Finally, perform clustering based on the fused feature representation, mine the implicit relationships between enterprises, and perform disambiguation of names based on the implicit relationships; The core formula of the graph attention model is as follows: e ij = LeakReLU(a T [Wh i || Wh j ); h' i = σ(∑ j∈N(i) α ij Wh j ); hi and hj represent the input hidden variables of the model, and W is the mapping matrix to be trained, which is used to map the input hidden variables to the space for attention weighting; Whi and Whj are the representations of companies i and j in this space, respectively, a T [Whi||Whj] is the product weighting of the concatenation result of vector a and Whi, Whj; αij represents the similarity between i and j, and it is used for attention weighting to obtain h′ i , that is, the hidden representation of the second layer of the model. Similarly, the hidden representation of the nth layer is obtained as the output of the model; Taking into account the information in the above three dimensions, namely e ij , α ij and h′ i , disambiguate the persons with the same name; For two related companies, a linear weighted scoring function f is designed: f(u, v) = w1x1(u, v) + w2x2(u, v) + w3x3(u, v) Among them, x1 is used to measure the weighted distance. The shorter the weighted distance, the larger x1; x1 is used to measure the common neighbors. The more common neighbors there are, the larger x2 is. x3 is used to measure the similarity of the representation vectors. The more similar the representation vectors are, the larger x3 is. The similarity here is based on cosine similarity ; u represents the hidden vector representation of the first company, v represents the hidden vector representation of the second company, wi represents the importance weight of the i-th feature, u and v are obtained through model training, and wi is a preset value; When f is greater than the threshold, it is determined that the people with the same name involved in the two companies are the same.

3. The human-enterprise relationship mining method based on graph representation technology and word segmentation technology according to claim 1, wherein: The specific process of S1 is as follows: S1.1: Extract the data of equity relationships, branch relationships, and guarantee relationships from the database and clean out the valid data; S1.2: Perform feature encoding on the valid data; S1.3: Establish the enterprise association graph structure and encode the original enterprise graph structure using the node2vec algorithm; Process the enterprise association graph. Based on the in-degree, out-degree, node attributes, and information of common neighbors of the nodes in the graph, reconstruct a new enterprise association graph for the input of the subsequent graph attention model. S1.4: Integrate the information feature encoding and structural encoding as the input of the graph attention model. Based on the annotation information, train the graph attention model to obtain the company representation vector based on the graph attention model. S1.5: Establish an enterprise personnel association graph. Select the central node u on the graph and perform a breadth-first search starting from this central node. Let the searched node be v. When f(u, v) reaches the set threshold, traverse all the people with the same name corresponding to u and v, and disambiguate the people with the same name among them.

4. A method for mining the human-enterprise relationship based on graph representation technology and word segmentation technology according to claim 1, characterized in that: The specific process of S2 is as follows: S2.1 Extract the company names and company category data from the database and clean the valid data. S2.2 Use the word segmentation technology to segment the company names, design heuristic rules to filter out the regional and industry information, and extract the brand names. S2.3 For non-self-employed individuals, merge the companies with similar brands to obtain the merging result. S2.4 For self-employed individuals, merge the companies with similar brands and close regions to obtain the merging result. S2.5 For the merged companies, traverse the people with the same name involved and disambiguate the people with the same name among them.

5. A method for mining the human-enterprise relationship based on graph representation technology and word segmentation technology according to claim 1, characterized in that: The specific process of S4 is as follows: S4.1 Establish a union-find set, sequentially read the disambiguation records of people with the same name in R1, R2, and R3, and merge the sets where the corresponding people with the same name are located. S4.2 Re-output the merged union-find set in the format of disambiguation records of people with the same name.

6. The human-enterprise relationship mining method based on graph representation technology and word segmentation technology according to claim 5, wherein: The union-find set is used for merging and querying. Among them, merging means merging two sets into a new set. Querying means determining whether two elements are in the same set. When using the union-find set, select to process in batches according to people's names.