A method for disambiguating the names of paper data based on graph neural networks

By applying a graph-based neural network method in the paper data, learning the paper representation vector and enhancing the representation, the problem of low disambiguation accuracy of authors of the same name in the prior art is solved, and a more efficient disambiguation effect is achieved.

CN116578708BActive Publication Date: 2025-07-01ZHEJIANG SCI-TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310584872.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-23
Publication Date
2025-07-01
Estimated Expiration
2043-05-23

AI Technical Summary

Technical Problem

When processing paper data of the author of the same name, it is difficult to effectively capture the rich semantic and structural information in the paper data, resulting in a low disambiguation accuracy.

Method used

Using a graph neural network-based method, each paper is used as a node of a heterogeneous network, edges are established through the paper attribute features, and an unsupervised graph autoencoder is used to learn paper representation vectors, and augment vector representation is enhanced through hierarchical attention mechanism networks, and finally the disambiguation of the author of the same name is achieved through hierarchical clustering.

Benefits of technology

It improves the accuracy of disambiguation of the author of the same name, can more effectively utilize the correlation information between papers, avoids the dependence on a large amount of labeled data in traditional methods, and enhances the accuracy of paper vector representation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116578708B_ABST
    Figure CN116578708B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for name disambiguation of paper data based on graph neural networks. This algorithm takes each paper as a node of a heterogeneous network, establishes edges through the strong correlation between paper attribute features, and uses an unsupervised graph autoencoder to learn the representation vector of each paper. At the same time, a hierarchical attention mechanism network is also adopted to enhance the vector representation of the paper. Finally, the hierarchical clustering algorithm is used to achieve name disambiguation for authors with the same name. Compared with traditional methods, the present invention uses graph neural networks to represent the nodes in the heterogeneous network, which can make full use of the correlation information between nodes and improve the accuracy of disambiguation. The present invention uses an unsupervised graph autoencoder to learn the representation vector of the paper, avoiding the problem of requiring a large amount of labeled data in traditional disambiguation methods. The present invention adopts a hierarchical attention mechanism network to learn the weight relationship between nodes and meta-paths, further enhancing the vector representation of the paper and the accuracy of disambiguation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of entity disambiguation, and particularly relates to a method for disambiguating the names of paper data based on a graph neural network. Background Art

[0002] The emergence of digital libraries has provided scholars with high-quality academic information resources, enabling them to conveniently access a vast amount of academic journals, papers, and scholar information, thus facilitating their academic research. With the continuous in-depth development of scientific research, researchers increasingly need high-quality academic resources to support their research work. Therefore, ensuring the accuracy of data in digital libraries has become particularly important. However, due to the widespread existence of the phenomenon of authors having the same name and the inconsistent recording methods caused by cultural differences, there are a large number of scholars with the same name in academic databases, which brings great trouble to users' information retrieval. Users need to spend a lot of time screening the retrieval results, increasing the difficulty of users' retrieval and hindering the development of scientific research activities. In addition, the existence of scholars with the same name may also lead to the misattribution of the research results of scholars to other scholars with the same name, which may affect the popularity and reputation of scholars, and even result in confusion and incorrect citations. At the same time, the citation counts of papers by scholars with the same name may also be wrongly included in the citation count of a specific scholar, thus affecting their academic ranking and evaluation and having an impact on scientometrics. Therefore, disambiguating authors with the same name has become an urgent problem to be solved in the literature database.

[0003] Regarding the problem of disambiguating authors with the same name, the existing solutions are mainly divided into the following several types:

[0004] 1. Supervised disambiguation methods mainly use a manually labeled training set to train a classification model to classify the literature of authors with the same name; however, supervised disambiguation methods need to be trained on a pre-labeled data set, and the cost of manually labeling the training set is too high and not suitable for disambiguating a large amount of data. Therefore, it has certain limitations.

[0005] 2. Unsupervised disambiguation methods mainly use the attribute features of the literature to calculate the similarity and use clustering algorithms for disambiguation; unsupervised methods do not require prior annotation of the data set, but it is difficult to select a suitable similarity determination threshold when calculating the similarity of the literature; at the same time, when clustering, since the number of authors with the same name cannot be determined in advance, the number of clustering results, that is, the number of clusters of authors with the same name, cannot be determined. Therefore, the disambiguation accuracy is relatively low.

[0006] 3. Semi-supervised disambiguation methods lie between supervised and unsupervised methods. They can use a small amount of labeled data information to train a classifier to classify a large amount of unlabeled data, thereby improving the accuracy of disambiguation results. However, this method often has a more complex structure, and its overall performance is more dependent on the integrity of manually labeled information, has high requirements for data quality, and there is a possibility of artificially generated noise.

[0007] 4. Graph-based disambiguation methods usually take authors or papers as nodes of the network, then construct a graph based on the relationships between papers or between authors and papers, and finally perform disambiguation through calculating the similarity between nodes or clustering algorithms. Usually, the disambiguation effect of this method is good, but existing graph-based disambiguation methods usually only consider simple relationships such as co-authorship relationships and citation relationships between papers, and the networks constructed by these simple relationships cannot effectively capture the rich semantic and structural information in paper data. Summary of the Invention

[0008] In view of the above, the present invention proposes a method for name disambiguation of paper data based on graph neural networks. This algorithm takes each paper as a node of a heterogeneous network, establishes edges through the strong correlation between paper attribute features, and uses an unsupervised graph autoencoder to learn the representation vectors of each paper. At the same time, a hierarchical attention mechanism network is adopted to enhance the vector representation of papers, and finally, the disambiguation of authors with the same name is achieved through a hierarchical clustering algorithm.

[0009] A method for name disambiguation of paper data based on graph neural networks includes the following steps:

[0010] (1) Use feature engineering to extract the paper features of each paper in the paper dataset as metadata for name disambiguation, and take each paper as a node in the heterogeneous network;

[0011] (2) Divide the paper dataset into several clusters of authors with the same name based on the conversion method of initial consonants of pinyin to solve the problem that the same author name has multiple different spellings;

[0012] (3) Use Word2Vec to perform word vector embedding representation on the paper features and generate the feature vectors of each paper, and then use a triplet loss model to adjust the feature vectors, and finally perform preliminary clustering based on the feature vectors;

[0013] (4) Construct an academic relationship network according to the co-corresponding authors of papers, and perform secondary clustering on authors with the same name in the same relationship network based on strong rules;

[0014] (5) Use a graph autoencoder to learn the distributed representation of nodes in the academic relationship network, so as to obtain the representation vectors of each node containing paper attribute information and paper relationship information;

[0015] (6) Use a node-level and semantic-level hierarchical attention network to learn the weight relationship between different nodes on the same meta-path and the weight relationship between different meta-paths, and then enhance the representation vector of the paper node through weighted fusion;

[0016] (7) The enhanced paper representation vectors are clustered using a hierarchical clustering algorithm to achieve name disambiguation.

[0017] Furthermore, the paper features extracted in step (1) are composed of two parts: paper attribute features and paper relationship features, wherein the paper attribute features include the author's name (first author), email address, address, institution name, and title, and the paper relationship features include co-authors, keywords, and publications.

[0018] Furthermore, the specific implementation process of step (2) is as follows:

[0019] Step 1: The names of authors of all papers are considered as classes, forming a class set A = {a1, a2, ..., a n};

[0020] Step 2: All author names are converted to lowercase and special symbols (such as commas, semicolons, hyphens, etc.) are removed;

[0021] Step 3: Use a unique Chinese character to correspond to the full pinyin of the author's name (e.g. Zeng corresponds to Zeng, Zheng corresponds to Zheng);

[0022] Step 4: Analyze whether the author's name is the full pinyin or the abbreviation of the initial consonants, and parse the full pinyin into the pinyin, the initial consonants corresponding to the pinyin, and the Chinese characters corresponding to the pinyin;

[0023] Step 5: If the author names of any two classes a1 and a2 in set A are both written in full pinyin and the corresponding Chinese characters are the same, or the author names of classes a1 and a2 contain initial consonant abbreviations and the corresponding initial consonants are the same, then a1 and a2 are merged into class a. 12 , and class a 12 Add to set A and remove a1 and a2;

[0024] Step 6: Repeat Step 5 until there are no more classes in set A that can be merged, and then the clustering ends.

[0025] Furthermore, in step (3), the word vector of each paper feature is first generated by Word2Vec, and then the weight of each paper feature is calculated by TF-IDF. Finally, the feature vector of each paper is obtained by weighted summing of all word vectors. The specific calculation formula is as follows:

[0026]

[0027] Among them: x m represents the paper feature, D i represents the feature set of paper i, x i represents the feature vector of paper i, represents the paper feature x m of the word vector, f m represents the paper feature x m of the weight coefficient.

[0028] Furthermore, in step (3), the triple loss model is used to adjust the feature vector, that is, a large number of positive and negative sample pairs are used as training data. The positive sample pair is two papers belonging to the same author, and the negative sample pair is two papers belonging to different authors. Then, according to the following loss function ζ d the triple loss model is trained. After the training is completed, the Word2Vec in the model is taken to recalculate and generate the feature vector of each paper;

[0029]

[0030] Among them: y ij =1 indicates that paper i and paper j belong to the same author, that is, the positive sample pair, y ik =0 indicates that paper i and paper k belong to different authors, that is, the negative sample pair, d ij represents the Euclidean distance between the feature vectors of paper i and paper j, d ik represents the Euclidean distance between the feature vectors of paper i and paper k, m is a fixed boundary distance constant, + is the hinge loss function.

[0031] Furthermore, in step (3), according to the adjusted feature vector, the similarity between the feature vectors of any two paper nodes is calculated by traversing in the heterogeneous network through cosine similarity. If the similarity is high enough (i.e., greater than the threshold), an edge is constructed between these two nodes.

[0032] Furthermore, since the email address is unique, in the case of complete email information, if two authors with the same name have the same email, these two authors are considered to be the same person. In step (4), the scholars who have co-authored with the same corresponding author are in the same academic relationship network.

[0033] Furthermore, the strong rules in step (4) include:

[0034] ① If the author names of two papers are the same, the address information is the same and they have the same co-authors, then it can be considered that these two papers belong to the same author;

[0035] ② If the author names, address information of two papers are the same and they are published in the same publication, then these two papers can be considered to belong to the same author;

[0036] ③ If the author names, address information of two papers are the same and they contain the same keywords, then these two papers can be considered to belong to the same author.

[0037] Furthermore, in step (6), first, the graph attention network is used to perform weighted fusion on the neighbor nodes on the same meta-path to obtain the node-level paper representation vector; then, the semantic-level attention mechanism is used to learn the importance of different meta-paths, and the semantics of each meta-path are fused to obtain the final paper representation vector; the meta-path is a path composed of nodes connected based on the same paper relationship features.

[0038] Based on the above technical solutions, the present invention has the following beneficial technical effects:

[0039] 1. The present invention uses a graph neural network to represent the nodes in a heterogeneous network, which can make full use of the correlation information between nodes and improve the accuracy of disambiguation.

[0040] 2. The present invention uses an unsupervised graph autoencoder to learn the paper representation vector, avoiding the problem of requiring a large amount of labeled data in traditional disambiguation methods.

[0041] 3. The present invention adopts a hierarchical attention mechanism network to learn the weight relationship between nodes and meta-paths, further enhancing the vector representation of papers and the accuracy of disambiguation. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 is a schematic flow chart of the method for disambiguating the names of papers in the present invention.

[0043] Figure 2 is a schematic network structure diagram of the triple loss model.

[0044] Figure 3 is a schematic flow chart of the method for disambiguating authors with the same name in the present invention.

[0045] Figure 4 is a schematic network structure diagram of the graph autoencoder.

[0046] Figure 5 is a schematic structure diagram of the hierarchical attention mechanism network.

[0047] Figure 6 is a schematic diagram of node-level weight calculation. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] In order to describe the present invention more specifically, the technical solution of the present invention is described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0049] like Figure 1 As shown, the paper data name disambiguation method based on graph neural network of the present invention includes the following steps:

[0050] (1) Data preprocessing.

[0051] The present invention requires preprocessing of the original data, including data cleaning and normalization.

[0052] First, the original data may contain missing information, abnormal characters, etc., and the special identifier "null" is used to fill them. Then, feature engineering is used to extract paper attribute features and paper relationship features, including author name, email address, co-author, address, institution name, title, keywords and publications, as metadata for disambiguation. Next, the acquired text data is denoised and segmented, including removing punctuation and special symbols, removing extra spaces and line breaks, removing stop words and lowercasing strings, and removing useless words. After denoising, the NTLK tool is used for segmentation and root extraction. Finally, an attribute label is added before each feature to facilitate the subsequent calculation of feature weights using TF-IDF.

[0053] (2) Solve the problem of one person having multiple roles.

[0054] The present invention uses a method based on pinyin initials to solve the problem of multiple different writing methods for the same author's name, and the specific implementation method is as follows:

[0055] Step 1: Treat the author name of each paper as a category, forming a set A = {a1, a2, ..., a n};

[0056] Step 2: Remove the capitalization and special symbols (such as commas, semicolons, hyphens, etc.) of the author's name;

[0057] Step 3: Match the pinyin of each name with a unique Chinese character, for example, "Zeng" matches "曾" and "Zheng" matches "郑";

[0058] Step 4: Analyze whether the author's name is a full pinyin name or an abbreviation of initial consonants, and parse the full pinyin name into pinyin, the initial consonants corresponding to the pinyin, and the Chinese characters corresponding to the pinyin;

[0059] Step 5: If the author names of class a1 and class a2 are both full names and the corresponding Chinese characters are the same, or if the author names of class a1 and class a2 contain initial consonant abbreviations and the corresponding initial consonants are the same, then merge a1 and a2 into a. 12 , and put a12 Add it to set A, and at the same time remove a1 and a2, otherwise jump to Step7;

[0060] Step6: If the number of classes in the class set is greater than 1, repeat Step4 and Step5;

[0061] Step7: End clustering.

[0062] (3) Feature content embedding.

[0063] The present invention uses the Word2Vec model to generate word vectors for each feature, then calculates the weight value of each feature through TF-IDF, and finally weighted-fuses and averages all word vectors to obtain the feature vector x of each paper i , and the specific calculation formula is as follows:

[0064]

[0065] where: x m represents the paper feature, D i represents the paper feature set, represents the word vector corresponding to each feature, f m represents the weight coefficient corresponding to each feature.

[0066] After obtaining the feature vector of the paper, the vector is adjusted through the triplet loss model to obtain a more accurate result. The structure of the triplet loss model is as Figure 2 shown. Given two papers D i and D j , if they belong to the same author, they form a pair of positive sample pairs; conversely, if they belong to different authors, they form a pair of negative sample pairs. The purpose of the triplet loss model is to find an accurate distance threshold m to distinguish positive sample pairs and negative sample pairs. It can make the distance between positive sample pairs closer and the distance between negative sample pairs farther away. Its loss function ζ d is as follows:

[0067]

[0068] where: d ij represents the distance between paper node i and node j, usually calculated using the Euclidean distance, that is, d ij =‖d i -d j ‖; y ij =1 indicates that the two papers belong to the same author, that is, a pair of positive sample pairs; y ik =0 indicates that the two papers belong to different authors, that is, a pair of negative sample pairs; [] + is the hinge loss function, which can be understood as [x]+ = max(0, x), where m is a fixed boundary distance constant.

[0069] Finally, the similarity between the paper feature vectors is calculated by cosine similarity. If the similarity between two papers is high enough, an edge is constructed between the corresponding nodes of the two papers.

[0070] (4) Relationship network construction.

[0071] Since email addresses are unique, in the case of complete email information, if two authors with the same name have the same email, they are considered the same person. Therefore, scholars who have co-authored with the same corresponding author are considered to be in the same academic relationship network. The present invention constructs a heterogeneous academic relationship network using the co-corresponding authors of the literature and disambiguates the authors with the same name in the same academic relationship network according to the following algorithm. The process of the algorithm is as follows:

[0072]

[0073] As Figure 3 shown, first, the email information of papers a1 and a2 is compared for the first clustering to reduce the complexity of subsequent clustering and improve efficiency; then, secondary clustering is performed according to the address institution. If the address institutions are the same, the two papers are considered to belong to the same author; if the primary institutions are the same but the secondary institutions are different, it can be judged again through citation and co-authorship relationships; if the co-authorship relationship and citation relationship match, clustering is performed; when the primary institutions cannot be matched, disambiguation can be performed by matching co-authorship relationships, citation relationships, and disciplines; if all three of these features can be matched, they are considered the same author and clustered into the specified cluster.

[0074] (5) Relationship network learning.

[0075] The present invention uses an unsupervised graph autoencoder to learn the distribution representation of nodes in the heterogeneous network, and then predicts the link relationship between nodes to obtain a new paper vector representation. The model structure of the graph autoencoder is as Figure 4 shown. The graph autoencoder consists of a node encoder model Z = g1(Y, A) and an edge decoder model which are composed of two parts. Among them, is the embedding matrix of node D, A ∈ R N×N is the adjacency matrix of graph G, which is mainly used to represent the relationship between nodes. Z = is the node embedding matrix, is the adjacency matrix predicted by the model. The goal is to minimize the reconstruction error between the predicted adjacency matrix and the original adjacency matrix A.

[0076] Coding part: The graph autoencoder uses a two-layer graph convolutional neural network (GCN) as the encoder to obtain the embedded representation of nodes. The calculation formula of the encoder g1 is as follows:

[0077]

[0078] Where: is the symmetrically normalized adjacency matrix, that is D is the node degree matrix of graph G, Relu(.) = max(0,.), and W0 and W1 are the parameters of the first and second layers of the graph neural network.

[0079] Decoding part: The graph autoencoder uses the inner product method to reconstruct the structural information of the original graph. The calculation formula of the decoder g2 is as follows:

[0080] g2(Z) = sigmoid(Z T Z)

[0081] Node D i and D j The probability of an edge existing between them is as follows:

[0082]

[0083] The cross-entropy is used as the loss function, and the specific formula is as follows:

[0084]

[0085] Finally, based on the graph autoencoder, a latent variable Z = [z1, z2,..., z n containing the paper attribute feature information and the relationship information between papers can be obtained, and it is used as the new vector representation of the paper.

[0086] (6) Relationship network enhancement.

[0087] The present invention uses a hierarchical attention mechanism network including node level and semantic level to enhance the vector representation of paper nodes. The structure of the network is as Figure 5 shown. In the weight calculation at the node level, the neighbor nodes on the same meta-path are weighted and fused through the graph attention network, so as to obtain a better node embedded representation. The node-level calculation process is as Figure 6 shown. Since the vector representation of each paper has been obtained using the graph autoencoder, only the weights of each neighbor node need to be calculated. Given the central node i, the weight calculation formula of its neighbor node j is as follows:

[0088] N ij = att node (n i , n j ) = σ(nT ·[n i ‖n j )

[0089] Where: N ij represents the importance of node j to node i. It should be noted that since the heterogeneous network is asymmetric, the weight coefficient N ij is also asymmetric; att node represents the node-level weight network model for generating weights. For nodes on the same meta-path, the weight network model is consistent; σ represents the sigmoid activation function, n represents the node-level attention vector, which is obtained through a single-layer feed-forward neural network training; n i n j represents the embedding vector of the corresponding node.

[0090] Normalizing the attention value of the node can obtain the weight coefficient M ij , and the calculation formula is as follows:

[0091]

[0092] Where: i represents the neighbor nodes (including node i itself) on the same meta-path.

[0093] By aggregating the neighbor nodes on the meta-path, the embedding representation of the central node i can be obtained, and its calculation formula is as follows:

[0094]

[0095] After obtaining the node representations under each meta-path, then use the semantic-level attention mechanism to learn the importance of different meta-paths and fuse the semantics of each meta-path to obtain the final vector representation; given the meta-path r i , r i 's weight is calculated as follows:

[0096]

[0097] Where: att sem represents the semantic-level weight network model, W is the weight matrix, q is the semantic-level attention vector, which is obtained through a feed-forward neural network; ν represents the set of attribute features, b is the bias vector; the weight coefficients S i of different types of meta-paths can be calculated by the following formula:

[0098]

[0099] By performing a weighted calculation on the weight coefficient of the meta-path and the node embedding, the final representation of node i can be obtained as follows:

[0100]

[0101] The model uses loss entropy as the loss function, and the specific calculation formula is as follows:

[0102]

[0103] Where: C is the parameter of the classifier, and y l represents the labeled node, and Y l and Z l are the label value and the predicted value of the label data.

[0104] (7) Clustering.

[0105] Finally, clustering is performed on the paper representation vectors obtained after enhancement through the hierarchical clustering algorithm, so as to achieve name disambiguation.

[0106] The above description of the embodiments is to enable those of ordinary skill in the art to understand and apply the present invention. It is obvious that those skilled in the art can easily make various modifications to the above embodiments and apply the general principles described herein to other embodiments without creative efforts. Therefore, the present invention is not limited to the above embodiments, and the improvements and modifications made by those skilled in the art to the present invention should be within the protection scope of the present invention.

Claims

1. A method for name disambiguation of paper data based on graph neural network, comprising the following steps: (1) Use feature engineering to extract the paper features of each paper in the paper dataset as metadata for name disambiguation, and regard each paper as a node in the heterogeneous network; (2) Divide the paper dataset into several clusters of authors with the same name based on the conversion method of initials of Chinese pinyin to solve the problem that the same author name has multiple different spellings; (3) Use Word2Vec to perform word vector embedding representation on the paper features and generate the feature vectors of each paper. Then, use the triplet loss model to adjust the feature vectors. That is, use a large number of positive and negative sample pairs as training data. The positive sample pairs are two papers belonging to the same author, and the negative sample pairs are two papers belonging to different authors. Then, according to the following loss function Train the triplet loss model. After training is completed, take the Word2Vec in the model to recalculate and generate the feature vectors of each paper. Finally, perform preliminary clustering based on the feature vectors; Wherein: y ij = 1 indicates that paper i and paper j belong to the same author, i.e., a positive sample pair, y ik = 0 indicates that paper i and paper k belong to different authors, i.e., a negative sample pair, d ij represents the Euclidean distance between the feature vectors of paper i and paper j, d ik represents the Euclidean distance between the feature vectors of paper i and paper k, and m is a fixed boundary distance constant, + is the hinge loss function; (4) Construct an academic relationship network according to the co-corresponding authors of the papers, and perform secondary clustering on the authors with the same name in the same relationship network based on strong rules; (5) Use a graph autoencoder to learn the distributed representation of nodes in the academic relationship network, so as to obtain the representation vectors of each node containing paper attribute information and relationship information between papers; (6) Use a hierarchical attention mechanism network including node level and semantic level to learn the weight relationship between different nodes on the same meta-path and the weight relationship between different meta-paths, and then enhance the representation vectors of paper nodes through weighted fusion; (7) Perform clustering on the enhanced paper representation vectors through a hierarchical clustering algorithm to achieve name disambiguation.

2. The method for disambiguating the names of paper data according to claim 1, wherein: The paper features extracted in step (1) consist of two parts: paper attribute features and paper relationship features, where the paper attribute features include author name, email, address institution name, title, and the paper relationship features include co-authors, keywords, publications.

3. The method for disambiguating the names of paper data according to claim 1, characterized in that: The specific implementation process of step (2) is as follows: Step1: Consider the author names of all papers as classes, forming a class set A = {a1, a2, …, a n}; Step2: Convert all author names to lowercase and remove special symbols; Step3: Replace the full spelling of pinyin in the author name with a unique Chinese character; Step4: Analyze whether the author name is the full pinyin or the initials abbreviation, and parse the full pinyin into pinyin, the initials corresponding to the pinyin, and the Chinese characters corresponding to the pinyin; Step 5: If the full pinyin names of the authors of any two classes a1 and a2 in set A are the same and the corresponding Chinese characters are identical, or the authors' names of classes a1 and a2 contain abbreviated initials and the corresponding initials are the same, then merge a1 and a2 into class a 12 , and add class a 12 to set A, and at the same time remove a1 and a2; Step6: Repeat Step5 until there are no more classes to merge in set A, and end the clustering.

4. The method for disambiguating the names of paper data according to claim 1, characterized in that: In step (3), first generate word vectors for each paper feature through Word2Vec, then calculate the weights of each paper feature through TF-IDF, and finally obtain the feature vectors of each paper by weighted summation of all word vectors. The specific calculation formula is as follows: Among them: x m represents the paper feature, D i represents the feature set of paper i, x i represents the feature vector of paper i, represents the paper feature x m is the word vector of, f m represents the paper feature x m is the weight coefficient of.

5. The method for disambiguating the names in the thesis data according to claim 1, wherein: In step (3), according to the adjusted feature vectors, calculate the similarity between the feature vectors of any two paper nodes by traversing in the heterogeneous network through cosine similarity. If the similarity is high enough, build an edge between these two nodes.

6. The method for disambiguating the names in paper data according to claim 1, characterized in that: Since the email address is unique, in the case of complete email information, if two authors with the same name have the same email, it is considered that these two authors are the same person. In step (4), scholars who have a co-author relationship with the same corresponding author are in the same academic relationship network.

7. The method for disambiguating the names of paper data according to claim 1, wherein: The strong rules in step (4) include: ① If two papers have the same author name, the same address information, and the same co-authors, then it can be considered that these two papers belong to the same author; ② If two papers have the same author name, the same address information, and are published in the same publication, then it can be considered that these two papers belong to the same author; ③ If the author names of two papers are the same, the address information is the same, and they contain the same keywords, then it can be considered that these two papers belong to the same author.

8. The method for disambiguating the names of paper data according to claim 1, wherein: In the step (6), first, the graph attention network is used to perform weighted fusion on the neighbor nodes on the same meta-path to obtain the node-level paper representation vector; then, the semantic-level attention mechanism is used to learn the importance of different meta-paths, and the semantics of each meta-path are fused to obtain the final paper representation vector; the meta-path is a path composed of nodes connected based on the same paper relationship characteristics.