Large language model-based author name disambiguation and error allocation detection method

By using large language models to conduct in-depth semantic analysis of academic literature, combined with the maximum entropy Markov field model and dynamic threshold filtering strategy, identifying the identities of attribution variants and interdisciplinary co-authors, the problem of difficulty in guaranteeing the accuracy and stability of author disambiguation methods in the existing technology in large-scale data environments is achieved, and higher identity matching accuracy and stability are achieved.

CN119990109AActive Publication Date: 2025-05-13SOUTHWEST PETROLEUM UNIV

Patent Information

Application Number
CN202510460617.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-05-13
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

Existing author disambiguation methods are difficult to ensure the accuracy and stability of matching in large-scale data environments, especially when faced with changes in the same name, attribution variants and research directions, it is difficult to effectively identify the identities of attribution variants and interdisciplinary co-authors.

Method used

Using a large language model-based approach, the academic literature is deeply analyzed through the BERT model, high-dimensional semantic eigenvectors are generated, and author identity representation and matching are performed to identify the identity of attribution variants and interdisciplinary co-authors.

Benefits of technology

It improves the ability to identify signature variants, avoids identity misjudgment caused by different expressions, enhances the ability to identify interdisciplinary collaborators, and improves the accuracy and stability of identity matching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990109A_ABST
    Figure CN119990109A_ABST
Patent Text Reader

Abstract

The invention discloses an author name disambiguation and error allocation detection method based on a large language model, and relates to the technical field of allocation detection. The method comprises the following steps: S1, forming a standardized academic literature data set; s2, generating a high-dimensional semantic feature vector representation used for representing a context semantic relationship of each piece of text information; s3, forming a uniform author identity representation vector; s4, utilizing the original academic literature data set to construct a multivariate relation graph of academic literature reference, authors and academic literatures, authors and institutions and copyright relation; s5, generating a preliminary author identity matching result; s6, generating an author identity probability distribution prediction result; s7, generating optimized author identity matching data; and S8, automatically correcting the author identity matching data by using a backward reasoning strategy to form a final author identity affiliation result. The signature variant identification capability is improved, and identity misjudgment caused by different expressions is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of assignment detection, and in particular to a method for author name disambiguation and misassignment detection based on a large language model. Background Art

[0002] With the continuous growth in the number of academic documents and the increasing frequency of cross-institutional cooperation, author name disambiguation has become one of the key technologies in academic database management and scientific research evaluation systems. Academic document databases need to make identity attribution judgments based on author signature information to support core applications such as document retrieval, scholar portrait construction, and scientific research impact assessment. However, relying solely on author names for identity matching has significant limitations. Problems such as the same name but different people, signature variants, and changes in author research directions make it difficult for existing author disambiguation methods to ensure the accuracy and stability of matching in large-scale data environments.

[0003] At present, traditional author disambiguation methods mainly include three types of methods: rule-based matching, text similarity calculation, and relationship inference based on social network graphs. The rule-based matching method relies on the author's institution, email address or resume information for identity attribution, but this method is prone to misjudgment when the author changes institutions or publishes anonymous academic papers. The method based on text similarity calculation matches different academic papers of the same author by analyzing the similarity of academic paper titles, abstracts and keywords, but this method cannot effectively cope with changes in research directions, resulting in a low accuracy rate in identity matching for scholars doing long-term cross-field research. The relationship inference method based on social network graphs uses the author's co-authorship, institutional affiliation and academic paper citation information to construct an academic cooperation network, and performs identity inference through the network structure. However, existing graph network-based methods usually rely on static features and are difficult to adapt to complex identity changes when dealing with dynamically changing author information.

[0004] In addition, the machine learning and deep learning methods introduced in recent years attempt to improve the accuracy of author disambiguation through feature engineering or neural network models, such as using topic models to analyze the similarity of authors' research interests or using graph neural networks to mine the structural information of authors in academic collaboration networks. However, machine learning still faces the following technical challenges: on the one hand, traditional text matching methods rely too much on explicit features and cannot deeply understand the semantic style of authors in different academic documents, resulting in the disambiguation effect being limited by the consistency of the corpus; on the other hand, network inference methods based on structural features are difficult to solve the problem of homonymous people and signature variants, especially in the absence of sufficient training data, which can easily lead to incorrect attribution.

[0005] Therefore, there is an urgent need for a method that combines semantic understanding, graph structure reasoning, and dynamic reasoning mechanism to improve the accuracy, scalability, and adaptability of author name disambiguation. Summary of the invention

[0006] One object of the present invention is to propose a method for author name disambiguation and misassignment detection based on a large language model, which improves the ability to recognize signature variants and avoids identity misjudgment due to different expressions.

[0007] According to an embodiment of the present invention, a method for author name disambiguation and misassignment detection based on a large language model comprises the following steps: S1. Collect original data from the database and construct an original academic literature data set, preprocess the original academic literature data set to form a standardized academic literature data set; S2. Use the BERT model to perform deep semantic analysis and vector encoding on the standardized academic literature data set to generate high-dimensional semantic feature vectors; S3. Based on the high-dimensional semantic feature vector, the author signatures and their context information in each academic document are fused to form a unified author identity representation vector; S4. Use the original data set to construct a multivariate relationship graph of academic literature citations, authors and academic literature, authors and institutions, and co-authorship relationships, and extract structural features that reflect the academic cooperation network and author attribution information based on the structural feature extraction method; S5. Combine author identity representation vectors and structural features to perform preliminary matching and clustering of author signatures in academic literature, identify multiple signature variants and duplicate names of the same author, and generate preliminary author identity matching results; S6. Based on the preliminary author identity matching results, high-dimensional semantic features and structural features, combined with the maximum entropy Markov field model, the maximum entropy principle is used to make a probabilistic determination of the author identity and generate the author identity probability distribution prediction results; S7. Perform multiple rounds of reasoning on the matching results based on the author identity probability distribution prediction results, use a dynamic threshold filtering strategy to distinguish high-confidence matching results from low-confidence matching results, and perform secondary verification and optimization on the low-confidence matching results to generate optimized author identity matching data; S8. Use reverse reasoning strategy to automatically correct author identity matching data to form the final author identity attribution result.

[0008] The beneficial effects of the present invention are: (1) The present invention uses the BERT language model to perform high-dimensional semantic encoding on academic document titles, abstracts, and author signatures in academic documents, extracts contextual relationships in the text through a self-attention mechanism, and generates a unified semantic feature vector. Author disambiguation usually relies on keyword matching or text similarity calculation based on TF-IDF, which makes it difficult to accurately process variants of different signatures of the same author in academic documents. The constructed deep semantic parsing model can effectively capture the contextual information of the author's signature, improve the ability to recognize signature variants, and avoid identity misjudgment due to different expressions.

[0009] (2) The present invention proposes a multivariate relationship graph modeling method for academic literature citations, authors and academic literature, authors and institutions, and co-authorship relationships. The structural information of the academic cooperation network is incorporated into the disambiguation calculation. The structured feature extraction can effectively characterize the author's long-term academic behavior pattern and improve the ability to identify cross-disciplinary and cross-institutional co-authors. By calculating the author's identity representation vector, the academic literature publication status, co-authorship density, and institutional affiliation information are considered simultaneously during identity matching, thereby constructing a more stable identity attribution strategy.

[0010] (3) The identity attribution probability modeling method based on the maximum entropy Markov field in the present invention breaks through the limitations of traditional identity matching methods based on rules or static classification. Through the maximum entropy principle, when calculating the author identity matching probability, it not only relies on the text information of a single academic document, but also combines the structural characteristics of the entire academic cooperation network and introduces a dynamic conditional probability modeling strategy, thereby making the identity matching results more stable. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings: Figure 1 The present invention provides a flowchart of a method for author name disambiguation and misassignment detection based on a large language model. DETAILED DESCRIPTION

[0012] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, which only illustrate the basic structure of the present invention in a schematic manner, and therefore only show the components related to the present invention.

[0013] refer to Figure 1 , a method for author name disambiguation and misassignment detection based on a large language model, comprising the following steps: S1. Collect original data from the database and construct an original academic literature data set, preprocess the original academic literature data set to form a standardized academic literature data set; Specifically, in this implementation, S1 includes the following sub-steps: S11. Extracting academic literature dataset D from the original data raw , the academic literature dataset consists of multiple academic documents. Each academic document in the academic literature dataset is d i Contains academic article titles, abstracts, and author signatures: , , i=1,2, ,M; Where, d i represents the data of the i-th academic document, T i is the title text of the i-th academic document, A i is the abstract text of the i-th academic document, S i is the author signature text of the i-th academic document, and M is the number of academic documents; In fact, the original data not only includes the title, abstract and author signature of the academic document, but also includes other information, such as paper citation information, author institution information, etc. Here, only the title, abstract and author signature are used, and the rest of the information will be used in subsequent steps.

[0014] S12. i Carry out format standardization and unified coding; In this process, there are two main aspects of work: Standardization of data formats. i For example, the format of academic documents is unified as follows: "Title: Document A; Author: Zhang San; Abstract: Abstract 1". Of course, those skilled in the art can choose other formats to unify them.

[0015] Character encoding conversion. That is, no matter which system or database the original data comes from, it is necessary to convert the character encoding to i All texts (including academic document titles, abstracts, and author signatures) are converted into the same character encoding format, which can use existing encoding methods, such as the commonly used UTF-8. This ensures that data from different sources will not be garbled or lose information during subsequent processing due to different encoding methods.

[0016] The purpose of this is to ensure that all literature data have the same encoding and format standards in the subsequent preprocessing, semantic parsing and model input processes, thereby improving the consistency and stability of data processing in the entire system.

[0017] S13. Perform noise data filtering on the uniformly coded academic literature dataset: remove academic literature containing invalid characters, garbled characters and abnormal formats, and generate a noise-filtered academic literature dataset; In this step, invalid characters mainly refer to unconventional symbols or illegal characters that appear in the text, such as characters that exceed the range of predefined legal characters (such as legal Chinese characters, letters, punctuation marks, etc.). During implementation, a "white list" or legal character set can be set, and characters that are not in this set (such as control characters, garbled symbols or other meaningless symbols) are judged as invalid characters.

[0018] Garbled characters are display anomalies caused by character encoding conversion errors, such as common replacement symbols (such as " ") or dense sequences of garbled characters. By detecting whether these typical garbled characters exist in the text, it can be determined that the document record has garbled characters.

[0019] Format anomaly refers to the failure of a document record to be organized in a predefined format, such as the lack of necessary fields (e.g., one of the academic document title, abstract, or author signature), or the field order not meeting expectations, or incorrect separators, etc. The determination of format anomaly depends on the verification of the standardization requirements in S12. If a record does not match the standard format in structure, it can be considered to have a format anomaly.

[0020] Of course, in actual operation, those skilled in the art can also design other filtering methods according to actual conditions.

[0021] When an academic document is detected to contain the above-mentioned invalid characters, garbled characters and abnormal formats, it will be deleted directly.

[0022] After the operation of S13, the encoding may be disrupted, so if this is detected, it can be encoded again.

[0023] S14. Calculate the text similarity between any two documents based on the academic document title text, identify and remove duplicate documents, and then use the remaining documents to construct a standardized academic document dataset D final ; In this step, we first calculate the academic literature title text T j Calculate the text similarity between documents: There are many methods for calculating the similarity, such as the currently conventional edit distance, Jaccard similarity coefficient, TF-IDF vector and embedding model. Those skilled in the art can select a suitable calculation method according to the actual situation. These methods are conventional methods in the field, so their specific operation steps are not described in detail here.

[0024] After identifying duplicate documents, the redundant duplicate documents are deleted and only one is retained. After all documents are processed, the remaining documents are unique, and then these documents are used to construct a standardized academic document data set D final ,In the disambiguation process, each academic document is only analyzed as a unique entity; S2. Use the BERT model to perform deep semantic analysis and vector encoding on the standardized academic literature data set to generate a high-dimensional semantic feature vector; this step includes the following sub-steps: S21. From the standardized academic literature data set D final Extract academic literature titles, abstracts, and author signature data from the dataset and construct a text dataset X input : j=1,2,...,N; Where, d j represents the jth standardized academic literature data, T j is the title text of the jth standardized academic document, A j is the abstract text of the jth standardized academic document, S j is the author signature text of the jth standardized academic document, and N is the number of standardized academic documents; S22. Use BERT language model to input Semantic parsing: Perform word segmentation on the academic document title text, academic document abstract text and author signature text respectively, and convert the word segmented text sequence into the corresponding word vector representation X bert :

[0025] in, , and Respectively represent The title text, abstract text and author signature text of a standardized academic document are represented by word vectors after being processed by the BERT language model; S23. Use the BERT model to bert Perform deep semantic analysis on each text sequence in the text to generate a high-dimensional text feature matrix H bert :

[0026] Among them, t represents the text type; L t is the number of words in text type t; x t,n is the word vector of the nth word in text type t; α t,n is the self-attention weight; f(x t,n ,θ) represents the BERT model for x t,nThe encoding function of , θ is the parameter set:

[0027] Where H is the number of attention heads; ω h is the learnable scaling factor of the h-th attention head; GAttn h (x t,n ,PE(x t,n )) represents the output of the h-th gated multi-head attention module; PE(x t,n ) is a dynamic position encoding function; LN( ) represents a layer normalization function, and FFN( ) is a feed-forward neural network; S24. Fuse the high-dimensional features of each text output by BERT to generate a comprehensive semantic representation V fusion :

[0028] in, , and are the high-dimensional text feature vectors calculated by BERT for the academic document title, abstract, and author signature text; T , A and λ S are the importance calibration coefficients of title, abstract and author respectively; the symbol concat() represents the vector concatenation operation; W f is the learnable weight matrix for linear fusion; b f is the bias vector; σ( ) is the activation function; the symbol ⊙ represents element-wise multiplication; W int is a learnable weight matrix used to calculate the interaction term; tanh( ) is the hyperbolic tangent activation function; the symbol ⊕ represents the vector addition operation; S25. Comprehensive semantic feature vector V fusion Normalize it to get a high-dimensional semantic feature vector .

[0029] S3. Based on the high-dimensional semantic feature vector representation, the author signatures and their context information in each academic document are fused to form a unified author identity representation vector; S31. Using high-dimensional semantic feature vector As input, for the standardized academic literature dataset D final Each academic article in i Extract author signature text S i and its corresponding context information, constructing an input feature set containing multi-level author identity representation:

[0030] in, is the comprehensive semantic feature vector of the i-th academic document, S i The author signature text of the i-th academic document; S32. Calculate the author signature text S i The bidirectional attention weight between the contextual information of the corresponding academic document is used to calculate the attention of the author's signature to the content of the academic document:

[0031] Calculate the attention of academic literature content to author signatures:

[0032] Among them, A context is the attention matrix of the author's signature on the content of academic literature, A fusion is the attention matrix of academic literature content to author signatures, Q S is the query vector of the author’s signature, K fusion is the key vector of academic literature title and abstract, Q fusion is the query vector of academic literature content, K S is the key vector of the author's signature, d k is the vector dimension scaling factor; S33. Calculate the weighted feature vector by the bidirectional attention weights of the author's signature and the academic document content:

[0033]

[0034] in, Indicates The author signature fusion features of academic papers, Indicates the content integration characteristics of academic literature; S34. Concatenate the features processed by the bidirectional attention to form the final semantic fusion features:

[0035] in, As the final fusion feature, characterize the signature text S i its semantic relevance within the context of the scholarly literature to which it belongs; S35. Calculate the signature variants of the same author in different academic documents and construct a signature consistency discrimination matrix:

[0036] Among them, D variant (S i , S j ) indicates the author's signature Si and S j The semantic similarity measure between them, sim( ) is a vector similarity calculation function, which is used to measure the signature variants that the same author may use in different academic papers; S36. Cluster all calculated similarities of author signature variants and construct an author identity normalization mapping to normalize different signature variants of the same author to a unified identity:

[0037] Among them, S normalized is the normalized set of author signatures, is the signature variant group after clustering, τ is the signature similarity threshold; S37. The fusion mechanism is used to construct the final author identity representation vector, which represents the signature consistency and academic style characteristics of the same author in different academic documents:

[0038] Among them, V identity is the identity representation vector; MLP( ) represents a multi-layer perceptron.

[0039] S4. Use the original academic literature dataset to construct a multivariate relationship graph of academic literature citations, authors and academic literature, authors and institutions, and co-authorship relationships, and extract structural features that reflect the academic cooperation network and author attribution information; this step includes the following sub-steps: S41. Extract academic literature citation information, author information, institution information, and co-author information from raw data; S42. Based on the academic literature citation information and the standardized academic literature data set, each academic literature is regarded as a node, and directed edges are constructed based on the citation relationship, thereby constructing an academic literature citation structure, which reflects the knowledge transfer path between academic literatures; S43. Combined with the relationship between authors and academic documents, combined with the standardized academic document data set, all authors are treated as independent entity nodes, and the connection relationship between authors and academic documents is established according to the signature of academic documents; S44. Construct a mapping network between authors and institutions based on the author's institutional information, bind the author to his or her institution, and establish corresponding edge connection relationships to represent the author's professional affiliation; S45. Construct an academic cooperation network based on the co-author information, with each author as a node in the network. If two authors sign the same academic document, an undirected edge is established in the academic cooperation network, indicating that the two authors have a direct academic cooperation relationship; S46. Based on the constructed academic cooperation network, calculate the belonging weight of each author in different network relationships: for authors whose academic papers published are higher than the preset value and whose citation frequency is higher than the preset value, assign high academic influence weight; for authors who have been affiliated with a certain institution for a long time, adjust the belonging weight in combination with the stability of the institution; for authors with close cooperation networks, calculate their social influence by co-occurrence frequency and the number of cooperative academic papers; the preset value of academic papers published and the preset value of citation frequency can be set according to actual conditions; S47. Using the structural feature extraction method, combined with the attribution weights calculated in S46, the global structural features of the academic collaboration network are extracted from the academic literature citation relationship graph, the author and academic literature relationship graph, the author and institution relationship graph, and the co-author relationship graph to generate the structural feature vector V structure (A k ).

[0040] The structural feature extraction method in the above steps can adopt conventional methods in the art, such as random walk and graph embedding method, graph convolutional network method, these methods can extract its structural features and finally generate a structural feature vector.

[0041] S5. Combine the author identity representation vector and structural features to perform preliminary matching and clustering of author signatures in academic documents, identify multiple signature variants and duplicate names of the same author, and generate preliminary author identity matching results; this step includes the following sub-steps: S51. The author identity representation vector and the structural feature are fused into a comprehensive feature representation; the method of fusing the two is a conventional method, so it will not be described in detail here.

[0042] S52. Calculate the comprehensive feature similarity of any two authors. When the similarity is greater than a threshold, the two authors can be preliminarily regarded as the same person. In this step, common similarity calculation methods can be selected, such as cosine similarity, Euclidean distance, etc., and the calculation is performed through the comprehensive feature vector after the fusion of the two authors. The threshold can be set according to the actual situation.

[0043] S53. Use a clustering algorithm to match the results of S52. Each cluster represents a different signature variant of the same author, and output a preliminary author identity matching set M. initial In this step, a common clustering algorithm in the art may be used, such as hierarchical clustering, K-Means clustering, etc. In the clustering results, duplicate names or mismatches will form different clusters.

[0044] S6. Based on the preliminary author identity matching results, a maximum entropy Markov field model is constructed, and high-dimensional semantic features and structural features are input into the maximum entropy Markov field model as observation variables. The maximum entropy principle is used to make a probabilistic judgment on the author identity and generate a prediction result of the author identity probability distribution. This step includes the following sub-steps: S61. Matching the preliminary author identity to the set M initial The observation information of each author in is constructed into serialized input, and the observation vector corresponding to each author is defined: k=1,2,...,W; Among them, concat(·,·) represents the vector concatenation operation, X k As the author A k The comprehensive observation features of V include both text semantic information and structural features; structure (A k ) represents the structural feature vector; V identity represents the identity representation vector; W is the author identity matching set M initial The number of authors, k represents M initial The author number in it; S62. Construct a maximum entropy Markov field model to determine the author's identity probabilistically. Define the model as a sequence model, where the state sequence s=(s1,s2,...,s k ) represents the identity of each author, and the local conditional probability of the maximum entropy Markov field model is expressed in log-linear form: ; Among them, P(s k |s k-1 ,X k ) is the local conditional probability of the kth author; s k is the author status selected at the kth position; f i (s k-1 ,s k ,X k ) represents the i-th feature function, and the input is the previous state s k-1 , current status k And the corresponding observation vector X k ; i is the learnable weight parameter of the i-th feature; n is the number of feature functions; S63. Construct a global probability distribution based on the local conditional probability and calculate the joint probability for the entire author status sequence s:

[0045] Where X={X1,X2,...,X K} is the set of observation vectors of all authors; S64. Calculate author A using global probability distribution k The final identity probability of is obtained by marginalizing the state sequence to obtain the author identity probability prediction result:

[0046] Among them, P(A k ) indicates author A k the final probability of being assigned a particular identity; By summing up and summing up the possible state sequences, the identity attribution with the highest probability is selected as the final identity determination result of the author.

[0047] S7. Perform multiple rounds of reasoning on the matching results based on the author identity probability distribution prediction results, use a dynamic threshold filtering strategy to distinguish high-confidence matching results from low-confidence matching results, and perform secondary verification and optimization on the low-confidence matching results to generate optimized author identity matching data; this step includes the following sub-steps: S71. According to the current matching result, adjust the author identity representation vector of S3, then repeat step S5, and adjust the learnable weight parameter in S62, and repeat the above operation multiple times until the preset number of repetitions is reached or the prediction result of S6 is stable; In this step, for the number of repetitions, that is, the number of reasoning rounds, you can set a threshold value for the rate of change after each round of reasoning, you can directly set a fixed number of rounds, and you can also use a dynamic stop strategy to set the number of reasoning rounds.

[0048] The operation based on the change rate threshold after each round of reasoning is as follows: Set the change rate of the matching results after each round of reasoning, such as the proportion of matches added or removed. When the change rate of several consecutive rounds (such as 2 to 3 rounds) is lower than the preset threshold (such as 1% or lower), it can be considered that the result is close to convergence and no further iteration is required.

[0049] Fixed number of rounds means that, based on experience, a fixed number of rounds (such as 3 to 5 rounds) can be set in advance, and the process stops after reaching a predetermined number of iterations.

[0050] The dynamic stopping strategy refers to: coordinating the real-time update of the matching results and the confidence distribution, dynamically adjusting the number of inferences, and automatically ending the iteration when the overall matching results are stable.

[0051] S72. According to the final result of S71, the distribution characteristics of the overall confidence are calculated; the matching results with confidence higher than the dynamic threshold are regarded as "high confidence", and those below the threshold are regarded as "low confidence"; the high confidence part is directly retained for use, and the low confidence part enters S73; in this step, firstly, according to the author identity matching probability distribution obtained from the final result of S71, the distribution characteristics of the overall confidence are calculated; then according to the data distribution, a dynamic threshold is set, and those above the threshold are regarded as high confidence, and those below the dynamic threshold are regarded as low confidence; for the high confidence part, we can directly retain it for use, but for the low confidence part, in order to prevent errors in the method of this embodiment, it is necessary to further test it.

[0052] S73. Add fine-grained features to S1 and S4, recalculate S1 to S6 for the low-confidence part, and compare the recalculated results with the original results, and then enter S74; or, use auxiliary models including graph neural networks to recalculate the results of the low-confidence part, and compare the new calculated results with the original results, and then enter S74.

[0053] This step proposes two different methods for testing the low-confidence part. One of them is to add fine-grained features for recalculation to prevent the original features from being too rough and causing errors in the calculation results. For fine-grained features, they include signature details, cooperation depth, and field-related knowledge. The signature details include institutional abbreviations and email addresses, and the cooperation depth includes the frequency of joint signatures and the timing of cooperation. The field-related knowledge includes keywords. Among them, each fine-grained feature is added to the part corresponding to the feature. For example, for signature details, it is added to the author identity representation vector of S3, for cooperation depth, it is added to the co-author information of S4 to construct an academic cooperation network, and for field-related knowledge, it is added to the standardized academic literature data set.

[0054] S74. If the new calculation result matches the original result, the original result is considered to be calculated correctly; if the new calculation result conflicts with the original result, manual review is introduced.

[0055] Through secondary verification, not only can false matches be eliminated, but some matches that originally had low confidence but should actually be retained can also be corrected and improved, further optimizing the final author identity matching data.

[0056] In general, in this step, based on the multi-round reasoning mechanism, a dynamic threshold strategy is used to divide the matching results into two parts: high and low confidence. Secondary verification and optimization are performed on the low confidence data to ensure that the final generated matching data can reach a high level in accuracy and robustness.

[0057] S8. Use reverse reasoning strategy to automatically correct author identity matching data to form the final author identity attribution result.

[0058] The reverse reasoning strategy means that after completing the preliminary author identity match, the matching results are used to work back to the original data and model predictions to check and verify the rationality of the match, thereby automatically correcting the preliminary results and forming the final author identity attribution results.

[0059] The specific operations are as follows: Preliminary matching results as a basis: During the forward reasoning process, the system obtains a preliminary author identity matching data set based on the input data and model predictions. This set contains the confidence and prediction probability corresponding to each matching candidate.

[0060] Reverse verification and consistency check: Using the reverse reasoning strategy, the system pushes the preliminary matching results back to the original features and data distribution, that is, to evaluate whether the prediction results are consistent with the data features and the original input.

[0061] For each preliminary matched author identity, the system checks its plausibility among other associated features (such as authorship conventions, collaborations, historical records, etc.).

[0062] If it is found that certain matching results do not conform to the expected characteristics or statistical laws during the reverse verification process, it is considered that the match may be biased or erroneous.

[0063] Automatic correction and adjustment: Based on the verification results of reverse reasoning, the system automatically corrects the detected inconsistent matches. The correction process includes: Adjust the matching confidence distribution, recalculate the probability distribution of incorrect matches, and then update the matching results; Re-evaluate suspicious matches using pre-set rules or additional auxiliary models; For matches that are extremely error-prone or uncertain, a secondary correction mechanism can be triggered to further narrow the error range.

[0064] Generate final results: After a comprehensive check and correction of preliminary matches using the reverse reasoning strategy, the system will filter out false matches, correct potential biases, and finally summarize the authorship attribution results that meet all constraints and feature verification conditions, which are more accurate and consistent than the previous ones.

[0065] In order to further illustrate the advantages of the method of the embodiment of the present invention, a specific test example is given below.

[0066] Example 1: In July 2023, the technical team of an international academic database discovered during routine maintenance that the number of academic papers related to scholar "J.**" in the database surged, from 280 in 2022 to 745 in 2023. However, database analysts noticed that the range of disciplines covered by these academic papers was unusually wide, including computer science, medicine, biology, and environmental science, far exceeding the research coverage of ordinary scholars. This abnormal situation indicates that the database may have classified multiple different "J.**" as the same scholar, resulting in mismatching of author identities, thereby affecting the accuracy of academic evaluation.

[0067] The technical team then sampled and analyzed the academic literature data related to "J.**", selected 50 academic papers in the field of computer science, and manually checked the author's institutional information. The results showed that the author "J.**" of at least 17 academic papers actually belonged to different individuals, from University A, University B, and University C. In addition, in 5 academic papers, the author's institutional information changed. In the example, the institution signed by an academic paper in 2022 was "Institution M", but in 2023, the author's institution was changed to "Institution U", but the author's research direction did not change significantly and was still concentrated in the field of machine learning, which showed that the existing disambiguation system of the database failed to correctly identify institutional changes.

[0068] To solve this problem, the technical team decided to apply the author disambiguation method of the present invention and verified its effectiveness through multiple rounds of testing.

[0069] First, the technical team extracted all academic literature data related to "J.**" from the database, including academic literature titles, abstracts and author signature information, a total of 745 academic papers involving 103 different institutions, and standardized these data to remove garbled characters and duplicate data, and unify the character encoding format.

[0070] Next, the team used the BERT language model to perform semantic analysis on the academic document titles, abstracts, and author signatures to extract high-dimensional semantic feature vectors. In the embodiment, in an academic document titled "D**" published in May 2023, the semantic features of the author "J.**" have a high degree of similarity with another academic document titled "C**" published in March 2023, indicating that the authors of these two academic documents may be the same person. However, compared with an academic document titled "G**" published in April 2023, its semantic feature similarity is extremely low, indicating that the author of the academic document "J.**" may be another different scholar.

[0071] Subsequently, the technical team constructed a multivariate relationship graph based on academic literature citation relationships, author-academic literature relationships, author-institution relationships, and co-authorship networks, and calculated structural features. Through the PageRank algorithm, the system found that a certain "J.**" mainly co-authored academic literature with researchers from institutions M and B, while another "J.**" mainly collaborated with researchers from universities A and B. This finding further supports the previous assumption that the database mistakenly classified the two "J.**" as the same scholar.

[0072] In order to further improve the matching accuracy, the technical team used the maximum entropy Markov field model to calculate the probability of author identity matching. In the embodiment, for the academic papers published by "J.**" in institutions M and B, the system calculated that the probability of identity matching was 92.5%, while the probability of matching the academic papers published by "J.**" in universities A and B was only 14.3%, which indicates that the two groups of "J.**" may be different people.

[0073] Finally, based on the above analysis, the system successfully identified 9 different "J.**" actually involved in 745 academic papers. The database team further verified the author information of these academic papers, communicated with some scholars, and confirmed the accuracy of the disambiguation results. In the embodiment, a "J.**" researcher (Department of Computer Science) stated that his research direction is artificial intelligence. However, his personal page in the database shows that he published an academic paper on "Cancer Biology" in 2023. This is obviously a mismatch. After the method of the present invention is applied, this error is corrected, so that the academic document attribution of the scholar is restored to accuracy. The comparative experimental data are shown in Table 1 below:

[0074] Table 1 Comparative data of author name disambiguation experiment method Disambiguation accuracy (%) Recall rate (%) F1-score (%) Mismatch rate (%) Calculation time (seconds) Rule matching (TF-IDF) 76.8 62.5 68.9 13.2 250 Machine Learning (Random Forest) 83.4 71.2 76.8 9.5 480 The present invention (BERT + MEMRF) 94.2 91.8 93.0 3.2 310

[0075] As can be seen from Table 1, the method of the present invention has high accuracy and can effectively disambiguate author names and correct erroneous assignment information.

[0076] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

Claims

1. A method for author name disambiguation and misassignment detection based on a large language model, characterized in that: The steps include: S1. Collect original data from the database and construct an original academic literature data set, preprocess the original academic literature data set to form a standardized academic literature data set; S2. Use the BERT model to perform deep semantic analysis and vector encoding on the standardized academic literature data set to generate high-dimensional semantic feature vectors; S3. Based on the high-dimensional semantic feature vector, the author signatures and their context information in each academic document are fused to form a unified author identity representation vector; S4. Use the original data set to construct a multivariate relationship graph of academic literature citations, authors and academic literature, authors and institutions, and co-authorship relationships, and extract structural features that reflect the academic cooperation network and author attribution information based on the structural feature extraction method; S5. Combine author identity representation vectors and structural features to perform preliminary matching and clustering of author signatures in academic literature, identify multiple signature variants and duplicate names of the same author, and generate preliminary author identity matching results; S6. Based on the preliminary author identity matching results, high-dimensional semantic features and structural features, combined with the maximum entropy Markov field model, the maximum entropy principle is used to make a probabilistic determination of the author identity and generate the author identity probability distribution prediction results; S7. Perform multiple rounds of reasoning on the matching results based on the author identity probability distribution prediction results, use a dynamic threshold filtering strategy to distinguish high-confidence matching results from low-confidence matching results, and perform secondary verification and optimization on the low-confidence matching results to generate optimized author identity matching data; S8. Use reverse reasoning strategy to automatically correct author identity matching data to form the final author identity attribution result.

2. The method according to claim 1, characterized in that The original data includes academic document titles, academic document abstracts, author information, academic document citation information, author institution information and author co-author information.

3. The method according to claim 1, characterized in that The S1 comprises the following steps: S11. Collect raw data from academic literature database and establish extracted academic literature dataset D raw : , , i=1,2, ,M; Where, d i represents the data of the i-th academic document, T i is the title text of the i-th academic document, A i is the abstract text of the i-th academic document, S i is the author signature text of the i-th academic document, and M is the number of academic documents; S12. i Carry out format standardization and unified coding; S13. Perform noise data filtering on the uniformly coded academic literature dataset: remove academic literature containing invalid characters, garbled characters and abnormal formats, and generate a noise-filtered academic literature dataset; S14. Calculate the text similarity between any two documents based on the academic document title text, identify and remove duplicate documents, and then use the remaining documents to construct a standardized academic document dataset D final .

4. The method according to claim 1, characterized in that: The S2 comprises the following steps: S21. From the standardized academic literature data set D final Extract academic literature titles, abstracts, and author signature data from the dataset and construct a text dataset X input : j=1,2,...,N; Where, d j represents the jth standardized academic literature data, T j is the title text of the jth standardized academic document, A j is the abstract text of the jth standardized academic document, S j is the author signature text of the jth standardized academic document, and N is the number of standardized academic documents; S22. Use BERT language model to input Semantic parsing: Perform word segmentation on the academic document title text, academic document abstract text and author signature text respectively, and convert the word segmented text sequence into the corresponding word vector representation X bert : in, , and Respectively represent The title text, abstract text and author signature text of a standardized academic document are represented by word vectors after being processed by the BERT language model; S23. Use the BERT model to bert Perform deep semantic analysis on each text sequence in the text to generate a high-dimensional text feature matrix H bert : Where t represents the text type; L t is the number of words of text type t; x t,n is the word vector of the nth word in text type t; α t,n is the self-attention weight; f(x t,n ,θ) represents the BERT model for x t,n The encoding function of , θ is the parameter set: Where H is the number of attention heads; ω h is the learnable scaling factor of the h-th attention head; GAttn h (x t,n ,PE(x t,n )) represents the output of the h-th gated multi-head attention module; PE(x t,n ) is a dynamic position encoding function; LN( ) represents a layer normalization function, and FFN( ) is a feed-forward neural network; S24. Fuse the high-dimensional features of each text output by BERT to generate a comprehensive semantic representation V fusion : in, , and are the high-dimensional text feature vectors calculated by BERT for the academic document title, abstract, and author signature text; T , A and λ S are the importance calibration coefficients of title, abstract and author respectively; the symbol concat() represents the vector concatenation operation; W f is the learnable weight matrix for linear fusion; b f is the bias vector; σ( ) is the activation function; the symbol ⊙ represents element-wise multiplication; W int is a learnable weight matrix used to calculate the interaction term; tanh( ) is the hyperbolic tangent activation function; the symbol ⊕ represents the vector addition operation; S25. Comprehensive semantic feature vector V fusion Normalize it to get a high-dimensional semantic feature vector .

5. The method according to claim 1, characterized in that The S3 comprises the following steps: S31. Using high-dimensional semantic feature vector As input, for the standardized academic literature dataset D final Each academic article in i Extract author signature text S i and its corresponding context information, constructing an input feature set containing multi-level author identity representation: in, is the comprehensive semantic feature vector of the i-th academic document, S i The author signature text of the i-th academic document; S32. Calculate the author signature text S i The bidirectional attention weight between the contextual information of the corresponding academic document is used to calculate the attention of the author's signature to the content of the academic document: Calculate the attention of academic literature content to author signatures: Among them, A context is the attention matrix of the author's signature on the content of academic literature, A fusion is the attention matrix of academic literature content to author signatures, Q S is the query vector of the author’s signature, K fusion is the key vector of academic literature title and abstract, Q fusion is the query vector of academic literature content, K S is the key vector of the author's signature, d k is the vector dimension scaling factor; S33. Calculate the weighted feature vector by the bidirectional attention weights of the author's signature and the academic document content: in, Indicates The author signature fusion features of academic papers, Indicates the content integration characteristics of academic literature; S34. Concatenate the features processed by the bidirectional attention to form the final semantic fusion features: Among them, concat() is a vector concatenation operation. As the final fusion feature, characterize the signature text S i its semantic relevance within the context of the scholarly literature to which it belongs; S35. Calculate the signature variants of the same author in different academic documents and construct a signature consistency discrimination matrix: Among them, D variant (S i , S j ) indicates the author's signature S i and S j The semantic similarity measure between them, sim( ) is a vector similarity calculation function, which is used to measure the signature variants that the same author may use in different academic papers; S36. Cluster all calculated similarities of author signature variants and construct an author identity normalization mapping to normalize different signature variants of the same author to a unified identity: Among them, S normalized is the normalized set of author signatures, is the signature variant group after clustering, τ is the signature similarity threshold; S37. The fusion mechanism is used to construct the final author identity representation vector, which represents the signature consistency and academic style characteristics of the same author in different academic documents: Among them, V identity is the identity representation vector; MLP( ) represents a multi-layer perceptron.

6. The method for author name disambiguation and misassignment detection based on a large language model according to claim 1, characterized in that: The S4 comprises the following steps: S41. Extract academic literature citation information, author information, institution information, and co-author information from raw data; S42. Based on the academic literature citation information and the standardized academic literature data set, each academic literature is regarded as a node, and directed edges are constructed based on the citation relationship, thereby constructing the academic literature citation structure; S43. Combined with the relationship between authors and academic documents, combined with the standardized academic document data set, all authors are treated as independent entity nodes, and the connection relationship between authors and academic documents is established according to the signature of academic documents; S44. Construct a mapping network between authors and institutions based on the author's institutional information, bind the author to his or her institution, and establish corresponding edge connection relationships to represent the author's professional affiliation; S45. Construct an academic cooperation network based on the co-author information, with each author as a node in the network. If two authors sign the same academic document, an undirected edge is established in the academic cooperation network, indicating that the two authors have a direct academic cooperation relationship; S46. Based on the constructed academic cooperation network, calculate the belonging weight of each author in different network relationships: for authors whose academic papers published are higher than the preset value and whose citation frequency is higher than the preset value, assign high academic influence weight; for authors who have been affiliated with a certain institution for a long time, adjust the belonging weight in consideration of the stability of the institution; for authors with close cooperation networks, calculate their social influence by co-occurrence frequency and the number of cooperative academic papers; S47. Using the structural feature extraction method, combined with the attribution weights calculated in S46, the global structural features of the academic collaboration network are extracted from the academic literature citation relationship graph, the author and academic literature relationship graph, the author and institution relationship graph, and the co-author relationship graph to generate the structural feature vector V structure (A k ); The structural feature extraction method includes random walk and graph embedding method, and graph convolutional network method.

7. The method according to claim 1, characterized in that S5 includes the following sub-steps: S51. Fuse the author identity representation vector and the structural features into a comprehensive feature representation; S52. Calculate the comprehensive feature similarity of any two authors. When the similarity is greater than a threshold, the two authors can be preliminarily considered to be the same person. S53. Use a clustering algorithm to match the results of S52. Each cluster represents a different signature variant of the same author, and output a preliminary author identity matching set M. initial .

8. The method according to claim 1, characterized in that The S6 comprises the following steps: S61. Matching the preliminary author identity to the set M initial The observation information of each author in is constructed into serialized input, and the observation vector corresponding to each author is defined: k=1,2,...,W; Among them, concat() represents the vector concatenation operation, X k As the author A k The comprehensive observation features of V include both text semantic information and structural features; structure (A k ) represents the structural feature vector; V identity represents the identity representation vector; W is the author identity matching set M initial The number of authors, k represents M initial The author number in it; S62. Construct a maximum entropy Markov field model to determine the author's identity probabilistically. Define the model as a sequence model, where the state sequence s=(s1,s2,...,s k ) represents the identity of each author, and the local conditional probability of the maximum entropy Markov field model is expressed in log-linear form: ; Among them, P(s k |s k-1 ,X k ) is the local conditional probability of the kth author; s k is the author status selected at the kth position; f i (s k-1 ,s k ,X k ) represents the i-th feature function, and the input is the previous state s k-1 , current status k And the corresponding observation vector X k ; i is the learnable weight parameter of the i-th feature; n is the number of feature functions; S63. Construct a global probability distribution based on the local conditional probability and calculate the joint probability for the entire author status sequence s: Where X={X1,X2,...,X K } is the set of observation vectors of all authors; S64. Calculate author A using global probability distribution k The final identity probability of is obtained by marginalizing the state sequence to obtain the author identity probability prediction result: Among them, P(A k ) indicates author A k the final probability of being assigned a particular identity; By summing up and summing up the possible state sequences, the identity attribution with the highest probability is selected as the final identity determination result of the author.

9. The method according to claim 8, characterized in that S7 includes the following sub-steps: S71. According to the current matching result, adjust the author identity representation vector of S3, then repeat step S5, and adjust the learnable weight parameter in S62, and repeat the above operation multiple times until the preset number of repetitions is reached or the prediction result of S6 is stable; S72. According to the final result of S71, the distribution characteristics of the overall confidence are calculated; the matching results with confidence higher than the dynamic threshold are regarded as "high confidence", and those below the threshold are regarded as "low confidence"; the high confidence part is directly retained and used, and the low confidence part enters S73; S73. Add fine-grained features in S1 and S4, recalculate S1 to S6 for the low confidence part, and compare the recalculated result with the original result, and then enter S74; Alternatively, an auxiliary model including a graph neural network is used to recalculate the result of the low confidence part, and the new calculation result is compared with the original result, and then enters S74; S74. If the new calculation result matches the original result, the original result is considered to be calculated correctly; if the new calculation result conflicts with the original result, manual review is introduced.

10. The method according to claim 9, characterized in that In S73, the fine-grained features include signature details, cooperation depth, and field-related knowledge. The signature details include the abbreviation of the organization and the email address. The cooperation depth includes the frequency of joint signatures and the timing of cooperation. The field-related knowledge includes keywords.

Citation Information

Patent Citations

  • Literature author name duplication disambiguation method and literature author name duplication disambiguation construction system

    CN112131872A

Cited By

  • Method and system for automatically and accurately tracking academic activities of scholars based on large language model

    CN121616285A

  • Machine-generated text author affiliation detection method based on multi-source feature fusion

    CN122221038A

  • Machine-generated text authorship attribution detection method based on multi-source feature fusion

    CN122221038B