Author Name Disambiguation and Misattribution Detection Method Based on Large Language Models

By using large language models to perform semantic analysis and feature encoding of academic literature, combined with the maximum entropy Markov field model and dynamic threshold filtering strategy, the accuracy and stability of author disambiguation methods in the existing technology in large-scale data environments is solved, and more efficient author identity matching and disambiguation effects are achieved.

CN119990109BActive Publication Date: 2025-06-20SOUTHWEST PETROLEUM UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510460617.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-06-20
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

Existing author disambiguation methods are difficult to ensure the accuracy and stability of matching in large-scale data environments, especially when faced with different people of the same name, attribution variants and changes in research directions, it is difficult to effectively identify and process them.

Method used

Using a large language model-based method, the title, abstract and author's signature in academic literature are deeply analyzed and vector encoding through the BERT model to generate high-dimensional semantic eigenvectors, and combined with the maximum entropy Markov field model and dynamic threshold filtering strategy, the author's identity is probabilistically determined and optimized.

Benefits of technology

It improves the ability to identify signature variants, avoids identity misjudgment caused by different expressions, enhances the ability to identify interdisciplinary and cross-institutional cooperative authors, and improves the accuracy and stability of identity matching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990109B_ABST
    Figure CN119990109B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for author name disambiguation and detection of incorrect assignments based on large language models, which relates to the technical field of assignment detection. The method includes the following steps: S1. Form a standardized academic literature data set; S2. Generate high-dimensional semantic feature vector representations for characterizing the context semantic relationships of each text information; S3. Form a unified author identity characterization vector; S4. Use the original academic literature data set to construct a multi-relationship graph of academic literature citations, authors and academic literature, authors and institutions, and co-authorship relationships; S5. Generate preliminary author identity matching results; S6. Generate author identity probability distribution prediction results; S7. Generate optimized author identity matching data; S8. Use the reverse inference strategy to automatically correct the author identity matching data to form the final author identity attribution result. The present invention improves the ability to identify signature variants and avoids misjudgment of identity caused by different expressions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of distribution detection, and particularly relates to a method for author name disambiguation and misattribution detection based on a large language model. Background Art

[0002] With the continuous growth of the number of academic literature and the increasing frequency of cross-institutional cooperation, author name disambiguation has become one of the key technologies in academic database management and scientific research evaluation systems. Academic literature databases need to judge identity attribution based on author signature information to support core applications such as literature retrieval, scholar profiling, and scientific research influence assessment. However, relying solely on author names for identity matching has significant limitations. Problems such as different people with the same name, signature variants, and changes in the research directions of authors make it difficult for existing author disambiguation methods to ensure the accuracy and stability of matching in a large-scale data environment.

[0003] Currently, traditional author disambiguation methods mainly include three types: rule-based matching, text similarity calculation, and relationship inference based on social network graphs. Rule-based matching methods rely on an author's institution, email address, or resume information for identity attribution. However, in cases where an author changes institutions or publishes anonymous academic literature, this method is prone to misjudgment. Text similarity calculation methods match different academic papers of the same author by analyzing the similarity of academic paper titles, abstracts, and keywords. However, this method cannot effectively handle situations where research directions change, resulting in a low accuracy rate for identity matching of scholars with long-term cross-disciplinary research. Relationship inference methods based on social network graphs use an author's co-authorship relationships, institutional affiliations, and academic paper citation information to construct an academic cooperation network and perform identity reasoning through the network structure. However, existing graph network-based methods usually rely on static features and are difficult to adapt to complex identity changes when dealing with dynamic author information.

[0004] In addition, machine learning and deep learning methods introduced in recent years have attempted to improve the accuracy of author disambiguation through feature engineering or neural network models. For example, using topic models to analyze the similarity of an author's research interests or using graph neural networks to mine the structural information of an author in an academic cooperation network. However, machine learning still faces the following technical challenges: on the one hand, traditional text matching methods rely too much on explicit features and cannot deeply understand the semantic styles of authors in different academic papers, resulting in the disambiguation effect being limited by the consistency of the corpus; on the other hand, network inference methods based on structural features are difficult to solve the problems of different people with the same name and signature variants, especially in the case of insufficient training data, which is prone to incorrect attribution.

[0005] Therefore, there is an urgent need for a method that combines semantic understanding, graph structure reasoning, and a dynamic reasoning mechanism to improve the accuracy, scalability, and adaptability of author name disambiguation. Summary of the Invention

[0006] An object of the present invention is to propose a method for author name disambiguation and detection of incorrect assignments based on a large language model. The present invention improves the ability to recognize signature variants and avoids misjudgment of identity caused by different expressions.

[0007] A method for author name disambiguation and detection of incorrect assignments based on a large language model according to an embodiment of the present invention includes the following steps:

[0008] S1. Collect raw data from a database and construct a raw academic literature dataset, and preprocess the raw academic literature dataset to form a standardized academic literature dataset;

[0009] S2. Use the BERT large model to perform in-depth semantic parsing and vector encoding on the standardized academic literature dataset to generate high-dimensional semantic feature vectors;

[0010] S3. Based on the high-dimensional semantic feature vectors, fuse the author signatures and their context information in each academic literature to form a unified author identity representation vector;

[0011] S4. Use the raw dataset to construct a multi-relational graph of academic literature citations, authors and academic literature, authors and institutions, and co-authorship relationships, and extract structural features reflecting academic cooperation networks and author attribution information based on the structural feature extraction method;

[0012] S5. Combine the author identity representation vector and the structural features to perform preliminary matching and clustering on the author signatures in the academic literature, identify multiple signature variants and homonym phenomena of the same author, and generate a preliminary author identity matching result;

[0013] S6. Based on the preliminary author identity matching result, high-dimensional semantic features and structural features, combine the maximum entropy Markov field model, and use the maximum entropy principle to probabilistically determine the author identity, and generate a predicted result of the author identity probability distribution;

[0014] S7. Perform multiple rounds of inference processing on the matching result according to the predicted result of the author identity probability distribution, use the dynamic threshold filtering strategy to distinguish between high-confidence matching results and low-confidence matching results, and perform secondary verification and optimization on the low-confidence matching results to generate optimized author identity matching data;

[0015] S8. Use the reverse inference strategy to automatically correct the author identity matching data to form the final author identity attribution result.

[0016] The beneficial effects of the present invention are:

[0017] (1) The present invention uses the BERT language model to perform high-dimensional semantic encoding on the titles, abstracts, and author signatures in academic literature, extracts the context relationships in the text through the self-attention mechanism, and generates unified semantic feature vectors. Author disambiguation usually relies on keyword matching or text similarity calculation based on TF-IDF, and it is difficult to accurately handle variants of the same author's different signatures in academic literature. The constructed deep semantic parsing model can effectively capture the context information of author signatures, improve the recognition ability of signature variants, and avoid misjudgment of identity caused by different expressions.

[0018] (2) The present invention proposes a method for modeling multi-relational graphs of academic literature citations, authors and academic literature, authors and institutions, and co-authorship relationships, incorporates the structural information of the academic cooperation network into the disambiguation calculation, and the structured feature extraction can effectively characterize the long-term academic behavior patterns of authors, improve the recognition ability of authors who cooperate across disciplines and institutions. By calculating the author identity representation vector, factors such as the author's academic literature publication situation, co-authorship relationship density, and institutional affiliation information are considered simultaneously during identity matching, thereby constructing a more stable identity attribution strategy.

[0019] (3) The present invention is based on the identity attribution probability modeling method of the maximum entropy Markov field, breaking through the limitations of traditional rule-based or static classification identity matching methods. Through the maximum entropy principle, when calculating the author identity matching probability, it not only depends on the text information of a single academic literature, but also combines the structural features in the entire academic cooperation network, and introduces a dynamic conditional probability modeling strategy, thereby making the identity matching result more stable. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the drawings:

[0021] Figure 1 is a flowchart of a method for author name disambiguation and misassignment detection based on a large language model proposed by the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0022] Now, the present invention will be further described in detail with reference to the drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.

[0023] Refer to Figure 1 , a method for author name disambiguation and misassignment detection based on a large language model, includes the following steps:

[0024] S1. Collect the original data from the database and construct the original academic literature dataset, and preprocess the original academic literature dataset to form a standardized academic literature dataset;

[0025] Specifically, in this embodiment, S1 includes the following sub-steps:

[0026] S11. Extract the academic literature dataset D from the original data raw , where the academic literature dataset consists of multiple academic literatures, and each academic literature d i in the academic literature dataset contains the academic literature title, abstract, and author signature:

[0027] , , i = 1, 2, , M;

[0028] In the formula, d i represents the i-th academic literature data, T i is the title text of the i-th academic literature, A i is the abstract text of the i-th academic literature, S i is the author signature text of the i-th academic literature, and M is the number of academic literatures;

[0029] In fact, the original data not only includes the academic literature title, abstract, and author signature, but also includes other information, such as paper citation information, author affiliation information, etc. Here, only the title, abstract, and author signature are used, and the other information will be used in subsequent steps.

[0030] S12. Standardize the format and perform unified encoding on d i ;

[0031] In this process, it mainly includes two aspects of work:

[0032] Standardization of data format. The structure of d i is made consistent according to the predetermined format standard. For example, the format of the academic literature is unified as: "Title: Literature A; Author: Zhang San; Abstract: Abstract 1". Of course, those skilled in the art can choose other formats to unify it.

[0033] Character encoding conversion. That is, no matter which system or database the original data comes from, all the text included in d i (including the academic literature title, abstract, and author signature) needs to be converted into the same character encoding format. Existing encoding methods can be used, such as the commonly used UTF-8. This can ensure that data from different sources will not have problems such as garbled characters or information loss due to different encoding methods during subsequent processing.

[0034] The purpose of doing this is to ensure that all literature data have the same encoding and format standards during subsequent preprocessing, semantic parsing, and model input processes, thereby improving the consistency and stability of the entire system in processing data.

[0035] S13. Perform noise data filtering on the uniformly encoded academic literature dataset: Remove academic literature containing invalid characters, garbled codes, and abnormal formats to generate a noise-filtered academic literature dataset;

[0036] In this step, invalid characters mainly refer to unconventional symbols or illegal characters appearing in the text, such as characters outside the predefined legal character range (e.g., legal Chinese characters, letters, punctuation marks, etc.). During implementation, a "whitelist" or set of legal characters can be set, and characters not in this set (such as control characters, garbled symbols, or other meaningless symbols) are determined to be invalid characters.

[0037] Garbled codes are display anomalies caused by character encoding conversion errors, such as common replacement symbols (e.g., " ") or dense sequences of garbled characters. By detecting whether these typical garbled marks exist in the text, it can be determined that there is a garbled code problem in the literature record.

[0038] Abnormal format means that the literature record fails to be organized according to the predefined format, such as missing necessary fields (e.g., one of the academic literature title, abstract, author signature), or incorrect field order, incorrect delimiter, etc. The determination of abnormal format depends on the verification of the standardization requirements in S12. If a record does not match the standard format in structure, it can be considered to have an abnormal format.

[0039] Of course, during the actual operation process, those skilled in the art can also design other filtering methods according to the actual situation.

[0040] When it is detected that a certain academic literature has the above-mentioned invalid characters, garbled codes, and abnormal formats, it is directly deleted.

[0041] After the operation in S13, the encoding may be disrupted. Therefore, if this situation is detected, it can be encoded again.

[0042] S14. Calculate the text similarity between any two literatures based on the academic literature title text, identify and remove duplicate literatures, and then construct a standardized academic literature data set D using the remaining literatures final ;

[0043] In this step, first calculate based on the academic literature title text T jCalculate the text similarity between documents: There are many calculation methods for this similarity, such as the conventional edit distance, Jaccard similarity coefficient, TF-IDF vector, and embedding model methods currently. Those skilled in the art can select a suitable calculation method according to the actual situation. These methods are all conventional methods in this field, so the specific operation steps are not elaborated here.

[0044] After identifying duplicate documents, delete the redundant duplicate documents and keep only one of them. After operating on all documents, the remaining documents are unique, and then construct a standard academic document data set D. final During the disambiguation process, each academic document is analyzed as a unique entity only.

[0045] S2. Use the BERT large model to perform in-depth semantic parsing and vector encoding on the standard academic document data set to generate high-dimensional semantic feature vectors; this step includes the following sub-steps:

[0046] S21. Extract the academic document title, abstract, and author signature data from the standard academic document data set D final and construct a text data set X input :

[0047] j = 1, 2,..., N;

[0048] In the formula, d j represents the jth standard academic document data, T j is the title text of the jth standard academic document, A j is the abstract text of the jth standard academic document, S j is the author signature text of the jth standard academic document, and N is the number of standard academic documents;

[0049] S22. Use the BERT language model to perform semantic parsing on X input : Perform word segmentation on the academic document title text, academic document abstract text, and author signature text respectively, and convert the segmented text sequence into the corresponding word vector representation X bert :

[0050]

[0051] Among them, , and respectively represent the word vector representations of the title text, abstract text, and author signature text of the jth standard academic document after being processed by the BERT language model;

[0052] S23. Use the BERT model to perform in-depth semantic parsing on each text sequence in X bert to generate a high-dimensional text feature matrix H bert :

[0053]

[0054] where t represents the text type; L t is the number of words of text type t; x t,n is the word vector of the nth word in text type t; α t,n is the self-attention weight; f(x t,n , θ) represents the encoding function of the BERT model for x t,n , and θ is the parameter set:

[0055]

[0056] where H is the number of attention heads; ω h is the learnable scaling factor of the hth attention head; GAttn h (x t,n , PE(x t,n )) represents the output of the hth gated multi-head attention module; PE(x t,n ) is the dynamic position encoding function; LN( ) represents the layer normalization function, and FFN( ) is the feed-forward neural network;

[0057] S24. Fuse the high-dimensional text features output by BERT to generate a comprehensive semantic representation V fusion :

[0058]

[0059] where , and are the high-dimensional text feature vectors calculated by BERT for the title, abstract, and author signature text of the academic literature respectively; λ T , λ A and λ S are the importance calibration coefficients of the title, abstract, and author respectively; the symbol concat( ) represents the vector concatenation operation; W f is the learnable weight matrix for linear fusion; b f is the bias vector; σ( ) is the activation function; the symbol ⊙ represents the element-wise multiplication; W int is the learnable weight matrix for calculating the interaction term; tanh( ) is the hyperbolic tangent activation function; the symbol ⊕ represents the vector addition operation;

[0060] S25. For the comprehensive semantic feature vector V fusionPerform normalization processing to obtain a high-dimensional semantic feature vector .

[0061] S3. Based on the high-dimensional semantic feature vector representation, fuse the author signatures and their context information in each academic literature to form a unified author identity representation vector;

[0062] S31. Use the high-dimensional semantic feature vector as the input, and for each academic literature d final in the standard academic literature data set D i extract the author signature text S i and its corresponding context information, and construct an input feature set containing multi-level author identity representations:

[0063]

[0064] Among them, is the comprehensive semantic feature vector of the i-th academic literature, and S i is the author signature text of the i-th academic literature;

[0065] S32. Calculate the bidirectional attention weights between the author signature text S i and the context information of its corresponding academic literature, and calculate the attention of the author signature to the academic literature content:

[0066]

[0067] Calculate the attention of the academic literature content to the author signature:

[0068]

[0069] Among them, A context is the attention matrix of the author signature to the academic literature content, and A fusion is the attention matrix of the academic literature content to the author signature, Q S is the query vector of the author signature, K fusion is the key vector of the academic literature title and abstract, Q fusion is the query vector of the academic literature content, K S is the key vector of the author signature, and d k is the vector dimension scaling factor;

[0070] S33. Calculate the weighted feature vector through the bidirectional attention weights between the author signature and the academic literature content:

[0071]

[0072]

[0073] Among them, represents the author signature fusion feature of the th academic literature, represents the content fusion feature of the academic literature;

[0074] S34. Concatenate the features after bidirectional attention processing to form the final semantic fusion feature:

[0075]

[0076] Among them, As the final fusion feature, it represents the semantic relevance of the signature text S i in its affiliated academic literature text;

[0077] S35. Calculate the signature variants of the same author in different academic literatures and construct a signature consistency discrimination matrix:

[0078]

[0079] Among them, D variant (S i , S j ) represents the semantic similarity measure between the author signatures S i and S j , and sim( ) is a vector similarity calculation function, which is used to measure the signature variants that the same author may adopt in different academic literatures;

[0080] S36. Cluster the similarity of all calculated author signature variants and construct an author identity normalization mapping for normalizing different signature variants of the same author to a unified identity identifier:

[0081]

[0082] Among them, S normalized is the set of normalized author signatures, is the grouped signature variants after clustering, and τ is the signature similarity threshold;

[0083] S37. Adopt a fusion mechanism to construct the final author identity representation vector, and the author identity representation vector represents the signature consistency and academic style characteristics of the same author in different academic literatures:

[0084]

[0085] Among them, V identity is the identity representation vector; MLP( ) represents the Multi-Layer Perceptron.

[0086] S4. Construct a multi-relational graph of academic literature citations, authors and academic literature, authors and institutions, and co-authorship relationships using the original academic literature dataset, and extract the structural features reflecting the academic cooperation network and author affiliation information; this step includes the following sub-steps:

[0087] S41. Extract academic literature citation information, author information, institutional information, and co-authorship information among authors from the original data;

[0088] S42. Based on the academic literature citation information and combined with the standardized academic literature dataset, regard each academic literature as a node, construct a directed edge based on the citation relationship, and thus construct an academic literature citation structure, which reflects the knowledge transfer path between academic literatures;

[0089] S43. Combine the relationship between authors and academic literature and the standardized academic literature dataset, regard all authors as independent entity nodes, and establish a connection relationship between authors and academic literature according to the signature situation of academic literature;

[0090] S44. Construct a mapping network between authors and institutions based on the institutional information of authors, bind the authors to their affiliated institutions, and establish corresponding edge connection relationships to represent the professional affiliation of authors;

[0091] S45. Construct an academic cooperation network according to the co-authorship information. Regard each author as a node in the network. If two authors sign the same academic literature, establish an undirected edge in the academic cooperation network, indicating that there is a direct academic cooperation relationship between the two authors;

[0092] S46. On the basis of the constructed academic cooperation network, calculate the affiliation weights of each author in different network relationships: For authors with a higher number of academic literature publications and citation frequencies than the preset values, assign a high academic influence weight; For authors who have been affiliated with a certain institution for a long time, adjust the affiliation weight in combination with the institutional stability; For authors with a close cooperation network, calculate their social influence through the co-occurrence frequency and the number of cooperative academic literatures; Here, the preset values of the number of academic literature publications and citation frequencies can be set according to the actual situation;

[0093] S47. Adopt a structural feature extraction method, combined with the affiliation weights calculated in S46, extract the global structural features of the academic cooperation network from the academic literature citation relationship graph, author and academic literature relationship graph, author and institution relationship graph, and co-authorship relationship graph, and generate a structural feature vector V structure (A k )。

[0094] For the structural feature extraction method of the above steps, conventional methods in the art can be adopted, such as random walk and graph embedding method, graph convolutional network method, and these methods can all extract their structural features and finally generate structural feature vectors.

[0095] S5. Combine the author identity representation vector and the structural features to perform preliminary matching and clustering on the author signatures in academic literature, identify multiple signature variants and homonym phenomena of the same author, and generate a preliminary author identity matching result; this step includes the following sub-steps:

[0096] S51. Fuse the author identity representation vector and the structural features into a comprehensive feature representation; the method of fusing the two is a conventional method, so it will not be elaborated here.

[0097] S52. Calculate the comprehensive feature similarity between any two authors. When the similarity is greater than the threshold, these two authors can be initially regarded as the same person; in this step, common similarity calculation methods can be selected, such as cosine similarity, Euclidean distance, etc., and calculated through the fused comprehensive feature vectors of the two authors, and the threshold can be set according to the actual situation.

[0098] S53. Use a clustering algorithm to match the results of S52. Each clustering cluster represents different signature variants of the same author, and output a preliminary author identity matching set M initial . In this step, common clustering algorithms in the art can be used, such as hierarchical clustering, K-Means clustering, etc. In the clustering results, homonyms or mis-matches will form different clustering clusters.

[0099] S6. Construct a maximum entropy Markov field model based on the preliminary author identity matching result, and input the high-dimensional semantic features and structural features as observation variables into the maximum entropy Markov field model, and use the maximum entropy principle to probabilistically determine the author identity and generate an author identity probability distribution prediction result; this step includes the following sub-steps:

[0100] S61. Construct the observation information of each author in the preliminary author identity matching set M initial into a serialized input, and define the observation vector corresponding to each author:

[0101] k = 1, 2,..., W;

[0102] where, concat(·,·) represents the vector concatenation operation, X k is used as the comprehensive observation feature of author A k , which contains both text semantic information and structural features; V structure (A k ) represents the structural feature vector; V identityRepresents the identity characterization vector; W is the set M of author identity matches initial The number of authors in it, k represents M initial The author number in it;

[0103] S62. Construct a maximum entropy Markov field model to probabilistically determine the author's identity, define this model as a sequence model, where the state sequence s = (s1, s2,..., s k ) represents the identity attribution identifier of each author, and the local conditional probability of the maximum entropy Markov field model is expressed in logarithmic linear form:

[0104] ;

[0105] Among them, P(s k |s k-1 , X k ) is the local conditional probability of the k-th author; s k is the selected author identity state at the k-th position; f i (s k-1 , s k , X k ) represents the i-th feature function, and the input is the previous state s k-1 , the current state s k and the corresponding observation vector X k ; λ i is the learnable weight parameter of the i-th feature; n is the number of feature functions;

[0106] S63. Construct a global probability distribution based on the local conditional probability, and calculate the joint probability for the entire author identity state sequence s:

[0107]

[0108] Among them, X = {X1, X2,..., X K} is the set of observation vectors of all authors;

[0109] S64. Use the global probability distribution to calculate the final identity probability of author A k , and through marginalization of the state sequence, obtain the author identity probability prediction result:

[0110]

[0111] Among them, P(A k ) represents the final probability that author A k is classified into a specific identity;

[0112] By summing and inducting over possible state sequences, select the identity attribution with the highest probability as the final identity determination result of the author.

[0113] S7. Perform multiple rounds of reasoning on the matching results based on the author identity probability distribution prediction results, use a dynamic threshold filtering strategy to distinguish high-confidence matching results from low-confidence matching results, and perform secondary verification and optimization on the low-confidence matching results to generate optimized author identity matching data; this step includes the following sub-steps:

[0114] S71. According to the current matching result, adjust the author identity representation vector of S3, then repeat step S5, and adjust the learnable weight parameter in S62, and repeat the above operation multiple times until the preset number of repetitions is reached or the prediction result of S6 is stable;

[0115] In this step, for the number of repetitions, that is, the number of reasoning rounds, you can set a threshold value for the rate of change after each round of reasoning, you can directly set a fixed number of rounds, and you can also use a dynamic stop strategy to set the number of reasoning rounds.

[0116] The operation based on the change rate threshold after each round of reasoning is as follows: Set the change rate of the matching results after each round of reasoning, such as the proportion of matches added or removed. When the change rate of several consecutive rounds (such as 2 to 3 rounds) is lower than the preset threshold (such as 1% or lower), it can be considered that the result is close to convergence and no further iteration is required.

[0117] Fixed number of rounds means that, based on experience, a fixed number of rounds (such as 3 to 5 rounds) can be set in advance, and the process stops after reaching a predetermined number of iterations.

[0118] The dynamic stopping strategy refers to: coordinating the real-time update of the matching results and the confidence distribution, dynamically adjusting the number of inferences, and automatically ending the iteration when the overall matching results are stable.

[0119] S72. According to the final result of S71, the distribution characteristics of the overall confidence are calculated; the matching results with confidence higher than the dynamic threshold are regarded as "high confidence", and those below the threshold are regarded as "low confidence"; the high confidence part is directly retained for use, and the low confidence part enters S73; in this step, firstly, according to the author identity matching probability distribution obtained from the final result of S71, the distribution characteristics of the overall confidence are calculated; then according to the data distribution, a dynamic threshold is set, and those above the threshold are regarded as high confidence, and those below the dynamic threshold are regarded as low confidence; for the high confidence part, we can directly retain it for use, but for the low confidence part, in order to prevent errors in the method of this embodiment, it is necessary to further test it.

[0120] In S73, fine-grained features are added to S1 and S4, and the calculation of S1 - S6 is redone for the low-confidence part. Then, the recalculated results are compared with the original results, and then it proceeds to S74; alternatively, an auxiliary model including a graph neural network is used to recalculate the results of the low-confidence part, and the newly calculated results are compared with the original results, and then it proceeds to S74.

[0121] This step proposes two different methods for verifying the low-confidence part. One is to add fine-grained features for recalculation to prevent errors in the calculation results caused by overly rough original features. For fine-grained features, they include signature details, cooperation depth, and domain-related knowledge. The signature details include institutional abbreviations and email addresses, the cooperation depth includes co-signature frequency and cooperation time sequence, and the domain-related knowledge includes keywords. Among them, each fine-grained feature is added to the corresponding part of the feature. For example, for signature details, they are added to the author identity representation vector of S3; for cooperation depth, they are added to the co-authorship information to construct an academic cooperation network in S4; for domain-related knowledge, they are added to the standardized academic literature data set.

[0122] In S74, if the newly calculated results match the original results, it is considered that the original results are calculated correctly; if the newly calculated results conflict with the original results, manual review is introduced.

[0123] Through secondary verification, not only are incorrect matches eliminated, but also some originally low-confidence matches that should actually be retained can be corrected and improved, further optimizing the finally generated author identity matching data.

[0124] Generally speaking, in this step, based on this multi-round inference mechanism, the matching results are divided into high-confidence and low-confidence parts using a dynamic threshold strategy. Secondary verification and optimization are implemented for low-confidence data, so as to ensure that the finally generated matching data can reach a relatively high level in terms of both accuracy and robustness.

[0125] In S8, an automatic correction is performed on the author identity matching data using a reverse inference strategy to form the final author identity attribution result.

[0126] The reverse inference strategy means that after the preliminary author identity matching is completed, the matching results are then used to reverse back to the original data and model predictions to check and verify the matching rationality, thereby automatically correcting the preliminary results to form the final author identity attribution result.

[0127] The specific operations are as follows:

[0128] Taking the preliminary matching results as the basis: In the forward inference process, the system obtains a preliminary author identity matching data set based on the input data and model predictions. This set contains the confidence levels and prediction probabilities corresponding to each matching candidate.

[0129] Reverse Verification and Consistency Check: Using the reverse reasoning strategy, the system traces the preliminary matching results back to the original features and data distribution, that is, evaluates whether the prediction results are consistent with the data features and the original input.

[0130] For each preliminarily matched author identity, the system checks its rationality in other associated features (such as signature conventions, collaboration relationships, historical records, etc.).

[0131] If it is found that some matching results do not conform to the expected features or statistical laws during the reverse verification process, it is considered that there may be biases or errors in this matching.

[0132] Automatic Correction and Adjustment: According to the verification results of the reverse reasoning, the system automatically corrects the detected inconsistent matches. The correction process includes:

[0133] Adjust the confidence distribution of the match, recalculate the probability distribution of the incorrect match, and then update the matching results;

[0134] Use pre-set rules or additional auxiliary models to re-evaluate the suspicious matches;

[0135] For matches that are extremely error-prone or uncertain, a secondary correction mechanism can be triggered to further narrow the error range.

[0136] Generate the Final Result: After the comprehensive inspection and correction of the preliminary matches by the reverse reasoning strategy, the system filters out the incorrect matches, corrects the potential biases, and finally summarizes the author identity attribution results that meet all the constraints and feature verification conditions. These results have higher accuracy and consistency than before.

[0137] To further illustrate the superiority of the method of the embodiment of the present invention, specific test cases are given below.

[0138] Example 1: In July 2023, the technical team of an international academic database found during routine maintenance that the number of academic literatures related to the scholar "J.**" in the database had increased significantly, from 280 in 2022 to 745 in 2023. However, the database analysts noticed that the academic literatures covered an unusually wide range of disciplines, including computer science, medicine, biology, environmental science, etc., far exceeding the research coverage of general scholars. This abnormal situation indicates that the database may have grouped multiple different "J.**" as the same scholar, resulting in incorrect matching of author identities and thus affecting the accuracy of academic evaluation.

[0139] The technical team then conducted a sampling analysis of the academic literature data related to "J.**", selected 50 academic literature in the field of computer science, and manually checked the institutional information of the authors. The results showed that the author "J.**" of at least 17 academic literature actually belonged to different individuals, from University A, University B, and University C respectively. In addition, in 5 academic literature, the institutional information of this author changed. In an academic literature in 2022 in the embodiment, the author's affiliated institution was "Institution M", but in 2023, the author's institution changed to "Institution U". However, the research direction of this author did not change significantly and still focused on the field of machine learning, indicating that the existing disambiguation system in this database failed to correctly identify the situation of institutional changes.

[0140] To solve this problem, the technical team decided to apply the author disambiguation method of the present invention and verified its effectiveness through multiple rounds of tests.

[0141] First, the technical team extracted all the academic literature data related to "J.**" from the database, including academic literature titles, abstracts, and author signature information, a total of 745 academic literature, involving 103 different institutions, and standardized this data, removing garbled characters, duplicate data, and unifying the character encoding format.

[0142] Next, the team used the BERT language model to perform semantic parsing on the academic literature titles, abstracts, and author signatures, and extracted high-dimensional semantic feature vectors. In the embodiment, in an academic literature named "D**" published in May 2023, the semantic features of the author "J.**" were highly similar to those of another academic literature named "C**" published in March 2023, indicating that the authors of these two academic literature might be the same person. However, compared with an academic literature named "G**" published in April 2023, the semantic feature similarity was extremely low, indicating that the author "J.**" of this academic literature might be another different scholar.

[0143] Subsequently, the technical team constructed a multi-relational graph based on academic literature citation relationships, author-academic literature relationships, author-institution relationships, and co-authorship networks, and calculated the structural features. Through the PageRank algorithm, the system found that a certain "J.**" mainly co-authored academic literature with researchers from Institution M and Institution B, while another "J.**" mainly collaborated with researchers from University A and University B. This finding further supported the previous hypothesis that the database misclassified the two "J.**" as the same scholar.

[0144] To further improve the matching accuracy, the technical team uses the maximum entropy Markov field model to calculate the author identity matching probability. In the embodiment, for the academic literature published by "J.**" in Institution M and Institution B, the system calculates the identity matching probability to be 92.5%, while for the academic literature of "J.**" in University A and University B, the matching probability is only 14.3%. This indicates that these two groups of "J.**" may be different people.

[0145] Finally, based on the above analysis, the system successfully identifies that there are actually 9 different "J.**" involved in 745 academic documents. The database team further verifies the author information of these academic documents and communicates with some scholars to confirm the accuracy of the disambiguation results. In the embodiment, a researcher "J.**" (Department of Computer Science) said that his research direction is artificial intelligence. However, his personal page in the database shows that he published an academic document on "cancer biology" in 2023, which is obviously a mis-match. After the application of the method of the present invention, this error is corrected, making the attribution of the scholar's academic documents accurate again. The comparative experimental data is shown in Table 1 below:

[0146] Table 1 Comparative data of author name disambiguation experiment

[0147] Method Disambiguation accuracy rate (%) Recall rate (%) F1-score (%) False matching rate (%) Computation time (seconds) Rule matching (TF-IDF) 76.8 62.5 68.9 13.2 250 Machine learning (Random Forest) 83.4 71.2 76.8 9.5 480 The present invention (BERT + MEMRF) 94.2 91.8 93.0 3.2 310

[0148] As can be seen from Table 1, the method of the present invention has high accuracy, can effectively disambiguate author names, and correct misassigned information at the same time.

[0149] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution of the present invention and its inventive concept, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.

Claims

1. A method for author name disambiguation and misassignment detection based on a large language model, characterized in that: The steps include: S1. Collect original data from the database and construct an original academic literature data set, preprocess the original academic literature data set to form a standardized academic literature data set; S2. Use the BERT model to perform deep semantic analysis and vector encoding on the standardized academic literature data set to generate high-dimensional semantic feature vectors; S3. Based on the high-dimensional semantic feature vector, the author signatures and their context information in each academic document are fused to form a unified author identity representation vector; S4. Use the original data set to construct a multivariate relationship graph of academic literature citations, authors and academic literature, authors and institutions, and co-authorship relationships, and extract structural features that reflect the academic collaboration network based on the structural feature extraction method; S5. Combine author identity representation vectors and structural features to perform preliminary matching and clustering of author signatures in academic literature, identify multiple signature variants and duplicate names of the same author, and generate preliminary author identity matching results; S6. Based on the preliminary author identity matching results, identity characteristics and structural characteristics, combined with the maximum entropy Markov field model, the maximum entropy principle is used to make a probabilistic determination of the author identity and generate a prediction result of the author identity probability distribution; S7. Perform multiple rounds of reasoning on the matching results based on the author identity probability distribution prediction results, use a dynamic threshold filtering strategy to distinguish high confidence matching results from low confidence matching results, and perform secondary verification and optimization on the low confidence matching results to generate optimized author identity matching data; this step includes the following sub-steps: S71. According to the current matching result, the author identity representation vector of S3 is adjusted, and then step S5 is repeated, and the learnable weight parameters of the maximum entropy Markov field model are adjusted, and the above operation is repeated multiple times until a preset number of repetitions is reached or the prediction result of S6 is stable; S72. According to the final result of S71, the distribution characteristics of the overall confidence are calculated; the matching results with confidence higher than the dynamic threshold are regarded as "high confidence", and those below the threshold are regarded as "low confidence"; the high confidence part is directly retained and used, and the low confidence part enters S73; S73. Add fine-grained features to S1 and S4, recalculate S6 for the low-confidence part, and compare the recalculated result with the original result, and then enter S74; or, use an auxiliary model including a graph neural network to recalculate the result of the low-confidence part, and compare the new calculated result with the original result, and then enter S74; S74. If the new calculation result matches the original result, the original result is considered to be calculated correctly; if the new calculation result conflicts with the original result, manual review is introduced; S8. Use reverse reasoning strategy to automatically correct author identity matching data to form the final author identity attribution result.

2. The method according to claim 1, characterized in that The original data includes academic document titles, academic document abstracts, author information, academic document citation information, author institution information and author co-author information.

3. The method according to claim 1, characterized in that The S1 comprises the following steps: S11. Collect raw data from academic literature database and establish extracted academic literature dataset D raw : ; Where, d i represents the data of the i-th academic document, T i is the title text of the i-th academic document, A i is the abstract text of the i-th academic document, S i is the author signature text of the i-th academic document, and M is the number of academic documents; S12. i Carry out format standardization and unified coding; S13. Perform noise data filtering on the uniformly coded academic literature dataset: remove academic literature containing invalid characters, garbled characters and abnormal formats, and generate a noise-filtered academic literature dataset; S14. Calculate the text similarity between any two documents based on the academic document title text, identify and remove duplicate documents, and then use the remaining documents to construct a standardized academic document dataset D final .

4. The method according to claim 1, characterized in that: The S2 comprises the following steps: S21. From the standardized academic literature data set D final Extract academic literature titles, abstracts, and author signature data from the dataset and construct a text dataset X input : ; Where, d j represents the jth standardized academic literature data, T j is the title text of the jth standardized academic document, A j is the abstract text of the jth standardized academic document, S j is the author signature text of the jth standardized academic document, and N is the number of standardized academic documents; S22. Use BERT language model to input Semantic parsing: Perform word segmentation on the academic document title text, academic document abstract text and author signature text respectively, and convert the word segmented text sequence into the corresponding word vector representation X bert : ; in, , and Respectively represent The title text, abstract text and author signature text of a standardized academic document are represented by word vectors after being processed by the BERT language model; S23. Use the BERT model to bert Perform deep semantic analysis on each text sequence in the text to generate a high-dimensional text feature matrix H bert : ; Where t represents the text type; L t is the number of words of text type t; x t,n is the word vector of the nth word in text type t; α t,n is the self-attention weight; f(x t,n ,θ) represents the BERT model for x t,n The encoding function of , θ is the parameter set: ; Where H is the number of attention heads; ω h is the learnable scaling factor of the h-th attention head; GAttn h (x t,n ,PE(x t,n )) represents the output of the h-th gated multi-head attention module; PE(x t,n ) is a dynamic position encoding function; LN( ) represents a layer normalization function, and FFN( ) is a feed-forward neural network; S24. Fuse the high-dimensional features of each text output by BERT to generate a comprehensive semantic representation V fusion : ; in, , and are high-dimensional text feature vectors calculated by BERT for academic document titles, abstracts, and author signature texts; the symbol concat() represents a vector concatenation operation; W f is the learnable weight matrix for linear fusion; b f is the bias vector; σ( ) is the activation function; the symbol ⊙ represents element-wise multiplication; W int is a learnable weight matrix used to calculate the interaction term; tanh( ) is the hyperbolic tangent activation function; the symbol ⊕ represents the vector addition operation; S25. Comprehensive semantic feature vector V fusion Normalize it to get a high-dimensional semantic feature vector .

5. The method according to claim 1, characterized in that The S3 comprises the following steps: S31. Using high-dimensional semantic feature vector As input, for the standardized academic literature dataset D final Each academic article in i Extract author signature text S i and its corresponding context information, constructing an input feature set containing multi-level author identity representation: ; in, is the high-dimensional semantic feature vector of the i-th academic document, S i The author signature text of the i-th academic document; S32. Calculate the author signature text S i The bidirectional attention weight between the contextual information of the corresponding academic document is used to calculate the attention of the author's signature to the content of the academic document: ; Calculate the attention of academic literature content to author signatures: ; Among them, A context is the attention matrix of the author's signature on the content of academic literature, A fusion is the attention matrix of academic literature content to author signatures, Q S is the query vector of the author’s signature, K fusion is the key vector of academic literature title and abstract, Q fusion is the query vector of academic literature content, K S is the key vector of the author's signature, d k is the vector dimension scaling factor; S33. Calculate the weighted feature vector by the bidirectional attention weights of the author's signature and the academic document content: ; ; in, Indicates The author signature fusion features of academic papers, Indicates the content integration characteristics of academic literature; S34. Concatenate the features processed by the bidirectional attention to form the final semantic fusion features: ; Among them, concat() is a vector concatenation operation. As the final fusion feature, characterize the signature text S i its semantic relevance within the context of the scholarly literature to which it belongs; S35. Calculate the signature variants of the same author in different academic documents and construct a signature consistency discrimination matrix: ; Among them, D variant (S i , S j ) indicates the author's signature S i and S j The semantic similarity measure between them, sim( ) is a vector similarity calculation function, which is used to measure the signature variants that the same author may use in different academic papers; S36. Cluster all calculated similarities of author signature variants and construct an author identity normalization mapping to normalize different signature variants of the same author to a unified identity: ; Among them, S normalized is the normalized set of author signatures, is the signature variant group after clustering, τ is the signature similarity threshold; S37. The fusion mechanism is used to construct the final author identity representation vector, which represents the signature consistency and academic style characteristics of the same author in different academic documents: ; Among them, V identity is the identity representation vector; MLP( ) represents a multi-layer perceptron.

6. The method for author name disambiguation and misassignment detection based on a large language model according to claim 1, characterized in that: The S4 comprises the following steps: S41. Extract academic literature citation information, author information, institution information, and co-author information from raw data; S42. Based on the academic literature citation information and the standardized academic literature data set, each academic literature is regarded as a node, and directed edges are constructed based on the citation relationship, thereby constructing the academic literature citation structure; S43. Combined with the relationship between authors and academic documents, combined with the standardized academic document data set, all authors are treated as independent entity nodes, and the connection relationship between authors and academic documents is established according to the signature of academic documents; S44. Construct a mapping network between authors and institutions based on the author's institutional information, bind the author to his or her institution, and establish corresponding edge connection relationships to represent the author's professional affiliation; S45. Construct an academic cooperation network based on the co-author information, with each author as a node in the network. If two authors sign the same academic document, an undirected edge is established in the academic cooperation network, indicating that the two authors have a direct academic cooperation relationship; S46. Based on the constructed academic cooperation network, calculate the belonging weight of each author in different network relationships: for authors whose academic papers published are higher than the preset value and whose citation frequency is higher than the preset value, assign high academic influence weight; for authors who have been affiliated with a certain institution for a long time, adjust the belonging weight in consideration of the stability of the institution; for authors with close cooperation networks, calculate their social influence by co-occurrence frequency and the number of cooperative academic papers; S47. Using the structural feature extraction method, combined with the attribution weights calculated in S46, the global structural features of the academic collaboration network are extracted from the academic literature citation relationship graph, the author and academic literature relationship graph, the author and institution relationship graph, and the co-author relationship graph to generate the structural feature vector V structure (A k ); The structural feature extraction method includes random walk and graph embedding method, and graph convolutional network method.

7. The method according to claim 1, characterized in that S5 includes the following sub-steps: S51. Fuse the author identity representation vector and the structural features into a comprehensive feature representation; S52. Calculate the comprehensive feature similarity of any two authors. When the similarity is greater than a threshold, the two authors can be preliminarily considered to be the same person. S53. Use a clustering algorithm to match the results of S52. Each cluster represents a different signature variant of the same author, and output a preliminary author identity matching set M. initial .

8. The method according to claim 1, characterized in that The S6 comprises the following steps: S61. Matching the preliminary author identity to the set M initial The observation information of each author in is constructed into serialized input, and the observation vector corresponding to each author is defined: ; Among them, concat() represents the vector concatenation operation, X k As the author A k The comprehensive observation features of V include both text semantic information and structural features; structure (A k ) represents the structural feature vector; V identity represents the identity representation vector; W is the author identity matching set M initial The number of authors, k represents M initial The author number in it; S62. Construct a maximum entropy Markov field model to determine the author's identity probabilistically. Define the model as a sequence model, where the state sequence s=(s1,s2,...,s k ) represents the identity of each author, and the local conditional probability of the maximum entropy Markov field model is expressed in log-linear form: ; Among them, P(s k |s k-1 ,X k ) is the local conditional probability of the kth author; s k is the author status selected at the kth position; f i (s k-1 ,s k ,X k ) represents the i-th feature function, and the input is the previous state s k-1 , current status k And the corresponding observation vector X k ; i is the learnable weight parameter of the i-th feature; n is the number of feature functions; S63. Construct a global probability distribution based on the local conditional probability and calculate the joint probability for the entire author status sequence s: ; Where X={X1,X2,...,X K } is the set of observation vectors of all authors; S64. Calculate author A using global probability distribution k The final identity probability of is obtained by marginalizing the state sequence to obtain the author identity probability prediction result: ; Among them, P(A k ) indicates author A k the final probability of being assigned a particular identity; By summing up and summing up the possible state sequences, the identity attribution with the highest probability is selected as the final identity determination result of the author.

9. The method according to claim 8, characterized in that In S73, the fine-grained features include signature details, cooperation depth, and field-related knowledge. The signature details include the abbreviation of the organization and the email address. The cooperation depth includes the frequency of joint signatures and the timing of cooperation. The field-related knowledge includes keywords.

Citation Information

Patent Citations

  • Literature author name duplication disambiguation method and literature author name duplication disambiguation construction system

    CN112131872A