Method and device for judging authors with same name based on multi-modal spatial-temporal feature fusion

By using a multimodal spatiotemporal feature fusion method, knowledge graphs and unit graphs are constructed by comprehensively utilizing multi-dimensional data, and the judgment threshold is dynamically adjusted. This solves the problems of misjudgment and omission in traditional methods of determining authors with the same name, and achieves accurate judgment and in-depth analysis in different academic scenarios.

CN120950706APending Publication Date: 2025-11-14SHANGHAI TEACHERS EDUCATION COLLEGE (TEACHING RESEARCH OFFICE OF SHANGHAI EDUCATION COMMISSION) +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510766276.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Traditional methods for determining authors with the same name are often based on single-dimensional information, which makes it difficult to fully characterize the author's academic features, leading to misjudgments and omissions, and affecting the fairness of the attribution of academic achievements and the evaluation of scientific research performance.

Method used

A multimodal spatiotemporal feature fusion method is adopted. By acquiring multidimensional data of the target author, a multimodal knowledge graph and a unit knowledge graph are constructed. Combined with the spatiotemporal correlation matrix and the domain knowledge graph, the judgment threshold is dynamically adjusted, and the judgment result and confidence score of the same author are output.

Benefits of technology

It improves the accuracy and reliability of identifying authors with the same name, can adapt to different academic scenarios, comprehensively portrays the academic characteristics of authors, reduces the probability of misjudgment and omission, and provides in-depth academic research analysis and information for discovering scientific research collaborations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120950706A_ABST
    Figure CN120950706A_ABST
Patent Text Reader

Abstract

The invention provides a homonymous author judgment method and equipment based on multi-modal spatial-temporal feature fusion, and relates to the technical field of data processing.The method comprises the steps that all paper metadata under a target author name are obtained, and a multi-modal knowledge graph is constructed, extracting text semantic features through a pre-training language model, extracting structure semantic features through a graph neural network, and extracting research interest evolution features through time sequence analysis; the names of the published units are clustered, a unit knowledge graph is constructed, and the academic association degree between the units is calculated; fusing the multi-modal knowledge graph and the unit knowledge graph, and calculating the final similarity between the papers; constructing a space-time incidence matrix and performing tensor decomposition, and extracting space-time fingerprint features of the target author research mode; according to the space-time fingerprint features, dynamically adjusting a judgment threshold through a domain knowledge graph, and outputting a homonymous author judgment result and a confidence score; the multi-dimensional information is comprehensively utilized, the knowledge graph is effectively fused, the dynamic adjustment capability is achieved, and the accuracy, reliability and adaptability of judgment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method and device for determining authors with the same name based on multimodal spatiotemporal feature fusion. Background Technology

[0002] With the vigorous development of academic research, the number of academic documents has grown exponentially, making the identification of authors with the same name increasingly prominent. In fields such as academic evaluation, research data analysis, and literature retrieval, accurately distinguishing authors with the same name is crucial to ensuring data reliability and the validity of research conclusions. Failure to accurately identify authors with the same name can lead to confusion regarding the attribution of academic achievements, affecting the fairness of research performance evaluation, and interfering with the normal conduct of academic research and the organization of knowledge.

[0003] Traditional methods for identifying authors with the same name often rely on single-dimensional information, such as the author's institution, keywords in the paper, or citation relationships. Because these methods use limited information sources, they fail to comprehensively depict the author's academic characteristics and are prone to misjudgment and omission. For example, relying solely on the institution makes it impossible to accurately distinguish authors with the same name working at the same institution; similarly, relying solely on keywords makes it difficult to draw reliable conclusions when different authors have similar research interests.

[0004] Therefore, there is an urgent need for a method to determine authors with the same name that can integrate multi-dimensional information, be comprehensive, accurate, and dynamically adaptable to different academic scenarios. Summary of the Invention

[0005] To address the aforementioned issues, this invention proposes a method and device for determining authors with the same name based on multimodal spatiotemporal feature fusion. This method comprehensively utilizes multi-dimensional information, accurately extracts features, effectively integrates knowledge graphs, and has dynamic adjustment capabilities, thereby improving the accuracy, reliability, and adaptability of the determination.

[0006] The objective of this invention is achieved through the following technical solution:

[0007] In a first aspect, the present invention provides a method for determining authors with the same name based on multimodal spatiotemporal feature fusion, the method comprising:

[0008] Retrieve metadata for all papers under the target author's name, including publishing institution, abstract, main text, publication date, citation relationships, and co-author information;

[0009] We construct a multimodal knowledge graph and extract textual semantic features through a pre-trained language model, extract structural semantic features through a graph neural network, and extract research interest evolution features through temporal analysis.

[0010] Cluster the names of publishing institutions to construct an institutional knowledge graph and calculate the academic relevance between institutions;

[0011] By integrating multimodal knowledge graphs and unit knowledge graphs, the final similarity between papers is calculated.

[0012] Construct a spatiotemporal correlation matrix and perform tensor decomposition to extract spatiotemporal fingerprint features of the target author's research patterns;

[0013] Based on spatiotemporal fingerprint characteristics, the judgment threshold is dynamically adjusted through the domain knowledge graph, and the judgment result and confidence score of the same author are output.

[0014] Preferably, the construction of the multimodal knowledge graph includes:

[0015] Keywords are mapped to the semantic space of a pre-trained language model to generate context-aware text semantic vectors;

[0016] Based on graph neural networks, a paper citation graph and a co-author social network graph are constructed to extract structural semantic features;

[0017] We conduct time series analysis on the publication time series of papers to construct the research interest evolution feature vector of the target authors.

[0018] Preferably, keywords are mapped to the semantic space of a pre-trained language model to generate context-aware text semantic vectors; including:

[0019] The abstract and main text of the paper are segmented into words, converting the text into multiple independent words;

[0020] The independent words are preprocessed by restoring part-of-speech tags, filtering part-of-speech tags, and removing stop words to obtain words after the first preprocessing.

[0021] The overall weight of a word is determined based on its positional weight in the paper, inverse document frequency, and time decay coefficient.

[0022]

[0023] Where α and β are position weight coefficients; α > β and α + β = 1; IDF(t) is the inverse document frequency; γ is the time decay coefficient; the value of γ ranges from [0.05, 0.2].

[0024] By employing an attention mechanism, text semantics are extracted and generated based on standardized keywords and the calculated comprehensive weights of words.

[0025] Preferably, the step of clustering the names of publishing institutions and calculating academic relevance to construct an institution knowledge graph includes:

[0026] String similarity calculation and hierarchical clustering are performed on the names of publishing institutions to form institution clusters;

[0027] Calculate the academic correlation between institutions based on the co-occurrence relationship of papers and author collaboration relationship within the unit cluster;

[0028] Construct a knowledge graph of units with units as nodes and academic relevance as edge weights.

[0029] Preferably, the calculation of academic correlation between institutions based on paper co-occurrence relationships and author collaboration relationships within the institution cluster includes:

[0030] For each pair of units within a unit cluster, calculate the co-occurrence strength of their papers;

[0031]

[0032] Among them, S p To determine the co-occurrence intensity of papers between institutions, U i and U j Represents any two distinct units within a unit cluster; Unit U i A collection of all published papers; Unit U j A collection of all published papers; ∩ is the intersection symbol; ∪ is the union symbol; Unit U i and U j The number of papers co-authored; Unit U i and U j The union of all published papers;

[0033] Extracting cross-unit co-author pairs C ij Calculate the cooperation strength S c And perform normalization processing;

[0034]

[0035] The academic relevance between institutions can be obtained by analyzing the co-occurrence intensity and collaboration intensity of papers.

[0036] Academic relevance = λ·S p +(1-λ)·S cz

[0037] Among them, S cz λ represents the normalized cooperation strength; λ is a coefficient, 0 < λ < 1.

[0038] Preferably, the method of fusing multimodal knowledge graphs and unit knowledge graphs to calculate the similarity between papers includes:

[0039] Calculate semantic similarity between papers based on text semantic vectors;

[0040] Calculate the structural similarity between papers based on structural semantic features;

[0041] Calculate the temporal evolution similarity between papers based on the evolutionary feature vector of research interests;

[0042] The initial similarity between papers is obtained by weighted fusion of semantic similarity, structural similarity, and temporal evolution similarity; and the weights are dynamically allocated through an attention mechanism.

[0043] The initial similarity is calibrated by the academic relevance between institutions; the final similarity between papers is then obtained.

[0044] S simf =S sim ·(1+μ·log(1+S i,j )

[0045] Among them, S simf S represents the final similarity between the two papers. sim S represents the initial similarity between the two papers; μ is the coefficient, 0 < μ < 1; S i,j This represents the academic relevance between the institutions to which the two papers belong.

[0046] Preferably, the step of constructing a spatiotemporal correlation matrix and performing tensor decomposition to extract spatiotemporal fingerprint features of the target author's research pattern includes:

[0047] Based on the final similarity between papers, hierarchical clustering is performed on all papers under the target author to generate a set of research topic clusters.

[0048] A three-dimensional spatiotemporal correlation tensor is constructed by the number of time slices, the number of geographical regions, and the number of research topic clusters, where each element in the tensor represents the academic output intensity of the target author.

[0049] An improved CP decomposition algorithm is used to approximately decompose the three-dimensional spatiotemporal correlation tensor into the sum of multiple components. Each component is obtained by outer product operation from the diagonal elements of the kernel tensor, the temporal pattern vector that captures the characteristics of the research cycle, the spatial distribution vector that reflects the regional cooperation preference, and the theme evolution vector that represents the change of research direction.

[0050] The core fingerprint components are extracted from the decomposed factor matrix, and the spatiotemporal evolution trajectory matrix of the target author's research pattern is constructed, thereby extracting the spatiotemporal fingerprint features of the target author's research pattern.

[0051] Preferably, the construction of the spatiotemporal evolution trajectory matrix of the target author's research pattern includes:

[0052] The time pattern vector sequence is processed using a long short-term memory network to obtain the first processing result, which means the evolution pattern of the target author's research activities in the time dimension.

[0053] A graph attention network is used to model the spatially distributed vector sequence to obtain a second processing result, which is the spatial correlation.

[0054] The first and second processing results are connected sequentially to form the spatiotemporal evolution trajectory matrix of the target author's research pattern.

[0055] Preferably, the step of dynamically adjusting the judgment threshold based on spatiotemporal fingerprint features using a domain knowledge graph, and outputting the judgment result and confidence score for authors with the same name, includes:

[0056] In the domain knowledge graph, a concept hierarchy tree is built. The semantic distance is obtained by calculating the number of edges traversed from one concept to another in the concept hierarchy tree and dividing it by the number of edges of the longest path in the concept hierarchy tree.

[0057] The judgment threshold is dynamically adjusted by comprehensively considering semantic distance and time factors;

[0058] Based on the similarity score, dynamic threshold, and the magnitude of the weight vector of related concepts in the domain knowledge graph, the confidence score is calculated, and the final output is the result of the same-name author determination and the confidence score.

[0059] In a second aspect, the present invention provides an electronic device, the electronic device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the method of the present invention.

[0060] The beneficial effects of this invention include: by abandoning the single-dimensional judgment method, comprehensively acquiring multi-source metadata such as the publishing institution, abstract, main text, publication time, citation relationship and co-author information of the paper, constructing a multimodal knowledge graph and an institution knowledge graph, and integrating the two to calculate the paper similarity; extracting accurate textual semantic features through a pre-trained language model, considering the positional weight of keywords in the paper, inverse document frequency and time decay coefficient, combining graph neural network to mine structural semantic features, and using time series analysis to capture the evolution characteristics of research interests, comprehensively and meticulously characterizing the academic features of authors, effectively avoiding misjudgment and omission due to one-sided information, and greatly improving the accuracy of identifying authors with the same name. By leveraging domain knowledge graphs and dynamically adjusting judgment thresholds based on semantic distance and temporal factors, the method flexibly adapts to various academic scenarios, delivering reliable results across both emerging interdisciplinary fields and established traditional disciplines. Through constructing a spatiotemporal correlation matrix and performing tensor decomposition, in-depth analysis of spatiotemporal features not only aids in identifying authors with the same name but also reveals the research cycle patterns of authors' academic activities over time, their regional collaboration preferences in space, and the evolution of their research directions over time, providing deeper information for academic research analysis and the discovery of research collaborations. Furthermore, clustering the names of publishing institutions to construct an institutional knowledge graph and calculating the academic correlation between institutions based on paper co-occurrence relationships and author collaboration relationships fully considers the actual academic connections between institutions, more accurately reflecting the closeness of academic research between different institutions, further improving paper similarity calculation, and providing stronger support for identifying authors with the same name. Attached Figure Description

[0061] Figure 1 This is a schematic diagram of the same-name author determination method based on multimodal spatiotemporal feature fusion provided in an embodiment of the present invention. Detailed Implementation

[0062] The present invention will now be further described in conjunction with the accompanying drawings and specific embodiments. It should be noted that, without conflict, the various embodiments or technical features described below can be arbitrarily combined to form new embodiments.

[0063] See Figure 1 A method for determining authors with the same name based on multimodal spatiotemporal feature fusion, the method comprising:

[0064] Retrieve metadata for all papers under the target author's name, including publishing institution, abstract, main text, publication date, citation relationships, and co-author information;

[0065] We construct a multimodal knowledge graph and extract textual semantic features through a pre-trained language model, extract structural semantic features through a graph neural network, and extract research interest evolution features through temporal analysis.

[0066] Cluster the names of publishing institutions to construct an institutional knowledge graph and calculate the academic relevance between institutions;

[0067] By integrating multimodal knowledge graphs and unit knowledge graphs, the final similarity between papers is calculated.

[0068] Construct a spatiotemporal correlation matrix and perform tensor decomposition to extract spatiotemporal fingerprint features of the target author's research patterns;

[0069] Based on spatiotemporal fingerprint characteristics, the judgment threshold is dynamically adjusted through the domain knowledge graph, and the judgment result and confidence score of the same author are output.

[0070] The working principle of the above technical solution is as follows:

[0071] First, we obtain the metadata of all papers under the target author's name, covering multiple dimensions such as publication institution, abstract, main text, publication date, citation relationships, and co-author information. This data forms the basis for subsequent analysis, providing material for characterizing the author's academic features from different perspectives.

[0072] This study utilizes pre-trained language models to process the abstracts and main text of academic papers, extracting semantic features and uncovering the knowledge content and research themes contained within. It also employs graph neural networks to analyze structural data such as citation relationships and co-authorship relationships, extracting structural semantic features and capturing the paper's position and connections within the academic network. Furthermore, through temporal analysis combined with the paper's publication date, it extracts the evolutionary characteristics of research interests, understanding the changing trends of authors' research directions over time, and ultimately constructing a multimodal knowledge graph.

[0073] The names of the publishing institutions of the papers are clustered, and institutions with similar or related names are grouped together. Based on this, an institutional knowledge graph is constructed. By analyzing information such as paper collaborations and personnel among institutions, the academic correlation between institutions is calculated, and the closeness of academic research among different institutions is assessed.

[0074] By fusing multimodal knowledge graphs with unit knowledge graphs, integrating information from multiple aspects such as textual semantics, structural semantics, research interest evolution, and unit academic associations, the final similarity between papers is calculated through a specific algorithm, thus more comprehensively measuring the degree of association between papers.

[0075] By constructing a spatiotemporal correlation matrix and comprehensively considering factors such as the publication time of the paper and the geographical location of the publishing institution, the spatiotemporal fingerprint features of the research patterns of the target authors are extracted through tensor decomposition technology, thus characterizing the academic activity patterns of the authors from the temporal and spatial dimensions.

[0076] Based on the extracted spatiotemporal fingerprint features and combined with the domain knowledge graph, the judgment threshold is dynamically adjusted. According to the adjusted threshold, authors with the same name are identified, and the judgment result and corresponding confidence score are output to clarify the reliability of the judgment.

[0077] The effects of the above technical solution are as follows:

[0078] This approach comprehensively considers various aspects of information, including textual semantics, structural semantics, evolution of research interests, academic connections within institutions, and spatiotemporal characteristics. Compared to single-dimensional judgment methods, it can more comprehensively and accurately characterize the academic features of authors, effectively reducing the probability of misjudgment and omission of authors with the same name, and improving the accuracy of the judgment results. By dynamically adjusting the judgment threshold and combining it with domain knowledge graphs, the method can adapt to the needs of identifying authors with the same name in different disciplines and research scales. Different disciplines have different research characteristics and academic network structures; this method can be flexibly adjusted according to the knowledge and data characteristics of specific domains to ensure the reliability of the judgment results.

[0079] In one possible implementation, the construction of the multimodal knowledge graph includes:

[0080] Keywords are mapped to the semantic space of a pre-trained language model to generate context-aware text semantic vectors;

[0081] Based on graph neural networks, a paper citation graph and a co-author social network graph are constructed to extract structural semantic features;

[0082] We conduct time series analysis on the publication time series of papers to construct the research interest evolution feature vector of the target authors.

[0083] In one possible implementation, the construction of the paper citation graph and co-author social network graph based on graph neural networks includes:

[0084] Construct a heterogeneous academic relationship graph, which includes three types of entities: paper nodes, author nodes, and institution nodes;

[0085] Define two types of edge relationships, including citation relationships between papers (directed edges) and co-authorship relationships between authors (weighted undirected edges, where weight = number of collaborations);

[0086] Node embeddings are generated using the Metapath2Vec algorithm; the metapath is defined as: paper—author—institution—author—paper.

[0087] The working principle of the above technical solution is as follows:

[0088] This invention constructs a heterogeneous academic relationship graph containing paper nodes, author nodes, and institution nodes. Each node represents an entity; for example, a paper node contains information such as the paper's title, abstract, and keywords; an author node contains basic information about the author; and an institution node contains information such as the institution's name and geographical location.

[0089] Citation relationships between papers are represented by directed edges, reflecting the inheritance and development of academic knowledge; co-authorship relationships between authors are represented by weighted undirected edges, with the weight being the number of collaborations, reflecting the closeness of collaboration between authors.

[0090] Node embeddings are generated using the Metapath2Vec algorithm, with the metapath defined as: paper—author—institution—author—paper. This metapath design allows the algorithm to capture semantic relationships between different types of nodes.

[0091] For example, starting with one paper, one can connect to other papers through citations, to co-authors through author relationships, to the institutions to which the co-authors belong through institutional relationships, and finally back to another paper through author relationships. This path can capture indirect relationships between papers, such as the potential connections between papers published by different authors from the same institution.

[0092] Based on the generated node embeddings, the system can extract structural semantic features from the papers. These features not only contain the content information of the papers themselves, but also the position and relationship information of the papers in the academic network. For example, although two papers have different research content, if their authors are from the same institution and frequently collaborate, then this potential connection can be captured through structural semantic features.

[0093] The effects of the above technical solution are as follows:

[0094] Heterogeneous academic relationship graphs can simultaneously represent the complex relationships between papers, authors, and institutions, providing a more comprehensive reflection of the academic network structure than single-type graphs. For example, traditional paper citation networks cannot directly represent collaborations between authors or connections between institutions. This method, through heterogeneous graphs, can capture these relationships simultaneously, providing richer information for identifying authors with the same name. Metapath design allows for the capture of indirect connections between different types of nodes. For example, a path of paper-author-institution-author-paper can reveal potential connections between different authors within the same institution, even if they have not directly collaborated. This indirect connection is highly valuable for identifying authors with the same name, as authors with the same name within the same institution may have similar research areas or collaboration patterns. The Metapath2Vec algorithm generates high-quality node embeddings that consider not only the attributes of the nodes themselves but also their structural position and relationships within the network.

[0095] In a possible implementation, mapping the keywords to the semantic space of the pre-trained language model to generate context-aware text semantic vectors includes:

[0096] Performing word segmentation on the abstract and the body content of the paper to convert the text into multiple independent words;

[0097] Performing preprocessing such as lemmatization,词性 filtering, and removing stop words on the independent words to obtain the first preprocessed words;

[0098] Based on the position weight, inverse document frequency, and time decay coefficient of the word in the paper, determining the comprehensive weight of the word;

[0099]

[0100] Among them, α and β are position weight coefficients; α>β and α + β = 1; IDF(t) is the inverse document frequency; γ is the time decay coefficient.

[0101] The value of α is n times that of β, n≥2; the value range of γ is [0.05, 0.2], and it is dynamically adjusted according to the update speed of domain knowledge.

[0102] Through the attention mechanism, according to the standardized keywords and the calculated comprehensive weight of the words, extracting the fused text semantics to generate text semantic vectors.

[0103] The working principle of the above technical solution is as follows: First, perform word segmentation on the abstract and the body content of the paper to break the continuous text into multiple independent words. For example, break "Research on Image Recognition Based on Deep Learning" into words such as "Based on", "Deep Learning", "Image Recognition", "Research", etc. Then, perform lemmatization on these independent words, restoring "running" to "run"; perform词性 filtering, retaining keywords such as nouns and verbs, and removing auxiliary words and conjunctions; remove stop words such as "of", "already", "in", etc., so as to obtain the first preprocessed words, reducing the interference of invalid information and laying a foundation for subsequent processing.

[0104] Based on the position weight, inverse document frequency, and time decay coefficient of the word in the paper, determining the comprehensive weight of the word. The position weight coefficients α and β correspond to the abstract and the body respectively, and α>β, α + β = 1, and the value of α is n times that of β (n≥2), which means that the words in the abstract are more important in reflecting the core content of the paper. The inverse document frequency

[0105] It should be noted that the "词性" in the original text seems to be an incorrect expression. It might be a misspelling or an incomplete term. I translated it as "词性" as it was in the original, but it may need to be corrected to a more accurate term in the actual context.IDF(t) measures the importance of words; words that appear less frequently in the entire document collection have higher IDF(t) values. A time decay coefficient γ (ranging from [0.05, 0.2] and dynamically adjusted based on the speed of knowledge updates in the domain) is used to adjust word weights based on the year difference between the paper's publication year and the current time. For example, in the computer science field, where knowledge updates rapidly, the γ value can be appropriately increased, causing the weights of words in earlier published papers to decrease more quickly over time; while in the field of historical research, the γ value can be relatively decreased.

[0106] The attention mechanism extracts semantic information from text based on standardized keywords and calculated word weights. It prioritizes words with high weights, extracting semantic information that better reflects the core content of the paper and ultimately generating context-aware semantic vectors. For example, in a paper on "novel solar cell materials," words like "solar cell" and "novel material" have high weights. The attention mechanism focuses on these words, generating semantic vectors that better reflect the paper's research theme of novel solar cell materials.

[0107] The effects of the above technical solution are as follows:

[0108] By considering the differences in word position between the abstract and the main text, the rarity of words in the document set, and the time factor, word weights can be accurately determined, thereby extracting text semantics more accurately. Compared with traditional methods that extract semantics based solely on word frequency, this method avoids the influence of high-frequency but irrelevant words (such as general descriptive words) on semantic understanding. For example, in medical papers, although the word "treatment" is high-frequency, if it appears less frequently in the abstract but is common in the document set, its weight will not be too high after combining inverse document frequency and position weight. Instead, truly key words such as disease names and special treatment methods will be highlighted, making the generated semantic vector more accurately reflect the core of the paper.

[0109] The time decay coefficient can be dynamically adjusted according to the speed of knowledge updates in the domain, enabling the method to adapt to different disciplines and stages of academic development. In the rapidly developing field of artificial intelligence, where knowledge iterates quickly, increasing the value can promptly reduce the weight of words in older papers, focusing more on new research results. In relatively slower-developing fields such as archaeology, decreasing the value preserves the importance of words in older literature, ensuring that the semantic vectors meet the semantic extraction needs of papers in different academic environments.

[0110] The generated context-aware text semantic vectors more accurately represent the core content and research direction of the paper, providing a reliable textual semantic basis for determining authors with the same name. When determining whether multiple papers by the same author belong to the same author, calculating the semantic similarity between papers based on these precise semantic vectors can effectively distinguish authors with the same name but different research directions, reduce misjudgments caused by semantic comprehension biases, and improve the accuracy and reliability of determining authors with the same name.

[0111] In one possible implementation, the step of clustering the names of publishing institutions and calculating academic relevance to construct an institution knowledge graph includes:

[0112] String similarity calculation and hierarchical clustering are performed on the names of publishing institutions to form institution clusters;

[0113] Calculate the academic correlation between institutions based on the co-occurrence relationship of papers and author collaboration relationship within the unit cluster;

[0114] Construct a knowledge graph of units with units as nodes and academic relevance as edge weights.

[0115] In one possible implementation, the calculation of academic correlation between institutions based on paper co-occurrence relationships and author collaboration relationships within an institution cluster includes:

[0116] For each pair of units (U) within a unit cluster i U j ), calculate the co-occurrence intensity of its papers;

[0117]

[0118] Among them, S p To determine the co-occurrence strength of the papers, U i and U j Represents any two distinct units within a unit cluster; Unit U i A collection of all published papers; Unit U j A collection of all published papers; ∩ is the intersection symbol; ∪ is the union symbol; Unit U i and U j Number of co-authored papers Unit U i The union of all papers published by Uj;

[0119] Extracting cross-unit co-author pairs C ij And calculate the cooperation strength S. c And perform normalization processing;

[0120]

[0121] The academic relevance between institutions can be obtained by analyzing the co-occurrence intensity and collaboration intensity of papers.

[0122] Academic relevance = λ·S p +(1-λ)·S cz

[0123] Among them, S cz λ represents the normalized cooperation strength; λ is a coefficient, 0 < λ < 1; λ is set according to the cooperation mode of the field, for example, λ = 0.3 for the experimental science field and λ = 0.6 for the theoretical field.

[0124] The working principle of the above technical solution is as follows:

[0125] First, string similarity is calculated for the names of the publishing institutions to identify institutions with similar names, such as "Computer Science Department of XX University" and "XX University". Then, a hierarchical clustering algorithm is used to organize these institutions into clusters, addressing the issue of institution name diversity, such as grouping different abbreviations or former names of the same institution into the same cluster.

[0126] Extracting cross-unit co-author pairs C ij And calculate the cooperation strength S. c For example, author A from unit U1 and author B from unit U2 co-authored 3 papers, forming a cross-unit co-author pair (A, B). Simultaneously, author C from unit U1 and author D from unit U2 also co-authored 1 paper, forming another cross-unit co-author pair (C, D). Therefore, the set of cross-unit co-author pairs C... 12 This includes the two co-author pairs, namely (C) 12 = {(A,B),(C,D)}; When author A of unit U1 and authors B and C of unit U2 co-author three papers, two co-author pairs will be created, namely (A,B) and (A,C). Both of these co-author pairs will be included in the cross-unit co-author pair set C. 12 In the case of co-author pair (A,B), the number of collaborative papers (A,B) is 3. Calculate the total number of papers for author A and author B respectively. Assuming author A has published a total of 10 papers and author B has published a total of 15 papers, then the total number of papers by the authors is 25.

[0127] Extracting cross-unit co-author pairs C ij The cooperation strength is calculated and normalized; the normalization process involves finding the maximum value S among the original values ​​of the cooperation strength between all unit pairs. cmax and minimum value S cmin Then normalize using the following formula:

[0128]

[0129] Alternatively, the original cooperation strength value can be Z-score normalized first, and then the result can be mapped to the 0 to 1 interval; choose according to the actual situation.

[0130] Each unit is treated as a node, and the academic correlation between units is used as the edge weight to construct a weighted graph. For example, if the academic correlation between unit A and unit B is 0.8, then the edge weight connecting A and B in the graph is 0.8.

[0131] The beneficial effects of the above technical solution are as follows:

[0132] Clustering algorithms effectively identify different expressions such as "University of California Berkeley," "UCBerkele y," and "Berkeley," grouping them into the same institution cluster and avoiding errors in association analysis caused by name differences. Simultaneously, co-occurrence of papers and author collaboration are considered to provide a more comprehensive assessment of academic associations. For example, even if two institutions have no jointly published papers, but have established a connection through co-authorship, this association can be effectively captured by the collaboration strength. Influence control of prolific authors: the logarithmic penalty term in the collaboration strength formula avoids the dominant role of prolific authors in association degree. When authoritative scholars in a certain field frequently collaborate with multiple institutions, this mechanism can more accurately reflect the contribution of ordinary researchers' collaborations to institution associations. In the determination of authors with the same name, the institution knowledge graph provides additional contextual information. For example, if two papers have authors with the same name, but their institutions have extremely low association in the knowledge graph, they can be inferred to be different authors. By adjusting the parameter λ, it can adapt to collaboration patterns in different disciplines. In fields with frequent inter-institutional collaborations (such as high-energy physics), the λ value can be increased to highlight the co-occurrence relationship of papers; while in fields dominated by small team collaborations (such as theoretical mathematics), the λ value can be decreased to emphasize the collaboration relationship between authors.

[0133] In one possible implementation, the fusion of multimodal knowledge graphs and unit knowledge graphs to calculate the similarity between papers includes:

[0134] Calculate semantic similarity between papers based on text semantic vectors;

[0135] Structural similarity between papers is calculated based on structural semantic features; paper nodes are mapped to low-dimensional vectors through graph embedding (e.g., Node2Vec), and Euclidean distance is calculated to obtain structural similarity;

[0136] The temporal evolution similarity between papers is calculated based on the research interest evolution feature vector; the time series interest vectors of two papers are aligned, and the similarity is measured using dynamic time warping (DTW);

[0137] The initial similarity between papers is obtained by weighted fusion of the three similarity measures mentioned above; the weights of the three are dynamically allocated through an attention mechanism.

[0138] The initial similarity is calibrated by the academic relevance between institutions; the final similarity between papers is then obtained.

[0139] S simf =S sim ·(1+μ·log(1+S i,j )

[0140] Among them, S simf S represents the final similarity between the two papers. sim S represents the initial similarity between the two papers; μ is the coefficient, 0 < μ < 1; S i,j This represents the academic relevance between the institutions to which the two papers belong.

[0141] The working principle of the above technical solution is as follows:

[0142] Based on the text semantic vectors generated by the pre-trained language model, cosine similarity is calculated to capture the semantic associations of the paper content; for example, two papers on "the application of deep learning in image recognition" will have a high text similarity.

[0143] By mapping the paper nodes to a low-dimensional space using Node2Vec, the Euclidean distance is calculated. If the network structures of two papers are similar, such as authors and citation relationships (e.g., belonging to the same research team or citing the same batch of literature), the structural similarity is high.

[0144] The Dynamic Time Warping (DTW) algorithm is used to align the research interest evolution feature vectors of two papers, allowing for non-uniform alignment of time series. For example, an earlier paper on "traditional machine learning" and a later paper that shifted to "deep learning" can still be identified as a continuation of the same research direction.

[0145] The attention mechanism dynamically assigns weights to the three similarity metrics based on the specific features of the paper. For example, for theoretical papers, the weight of textual semantic similarity may be higher; while for papers produced through cross-team collaboration, the weight of structural similarity may be increased.

[0146] The initial similarity is calibrated using academic relevance within an institutional knowledge graph. The calibration formula indicates that if the institutions of the authors of two papers have a high degree of relevance (e.g., frequent collaboration), the final similarity will be enhanced. For example, papers by authors with the same name from the same institution are more likely to be identified as belonging to the same person.

[0147] The beneficial effects of the above technical solution are as follows:

[0148] Simultaneously considering three dimensions—text content, network structure, and temporal evolution—avoids the limitations of a single dimension; for example, two papers with similar research topics but published far apart can still be linked through temporal similarity analysis. Adaptive weight allocation: The attention mechanism enables the model to automatically adjust similarity weights based on the characteristics of the papers, improving flexibility.

[0149] Institutional affiliation calibration effectively utilizes information about institutional collaborations. For example, in the biomedical field, collaborations between different hospitals and research institutions are frequent, and institutional affiliation calibration can significantly improve the accuracy of identifying authors with the same name. For processing time-series data: the DTW algorithm allows for asynchronous time-series comparisons, making it suitable for capturing the evolution of research interests. In the field of artificial intelligence, this method can identify shifts in research topics from "neural networks" to "deep learning," avoiding similarity drops caused by keyword changes.

[0150] By adjusting the coefficient μ, collaboration models across different disciplines can be adapted. In large-scale scientific fields such as high-energy physics, the value of μ can be appropriately increased to strengthen the influence of unit correlation; while in fields such as mathematics where individual research is the primary focus, the value of μ can be decreased.

[0151] In one possible implementation, the step of constructing a spatiotemporal correlation matrix and performing tensor decomposition to extract spatiotemporal fingerprint features of the target author's research pattern includes:

[0152] Based on the final similarity between papers, hierarchical clustering is performed on all papers under the target author to generate a set of research topic clusters.

[0153] A three-dimensional spatiotemporal correlation tensor is constructed by the number of time slices, the number of geographical regions, and the number of research topic clusters, where each element in the tensor represents the academic output intensity of the target author.

[0154]

[0155] in, To analyze the research topics belonging to the k-th cluster (Cluster) k For each paper p in the topic cluster, calculate its centroid with respect to that topic cluster. k Final similarity; tensor element T o,q,k This represents the normalized academic output intensity of the target author within the o-th time period, q-th region, and k-th topic cluster.

[0156] An improved Canonical Polyadic Decomposition (CP) algorithm is used to approximately decompose the three-dimensional spatiotemporal correlation tensor into the sum of multiple components. Each component is obtained by outer product operation from the diagonal elements of the kernel tensor, the temporal pattern vector that captures the characteristics of the research cycle, the spatial distribution vector that reflects the regional cooperation preference, and the theme evolution vector that represents the change of research direction.

[0157] The core fingerprint components are extracted from the decomposed factor matrix; and the spatiotemporal evolution trajectory matrix of the target author's research pattern is constructed, thereby extracting the spatiotemporal fingerprint features of the target author's research pattern.

[0158]

[0159] Where max(x) r ) represents the time pattern vector x r The maximum value in argmax(b) reflects the strongest research activity intensity of the target author at a certain time slice under the corresponding component r; r ) is the index of the maximum value in the spatial distribution vector, which can determine the geographical region where the target author is most inclined to cooperate under component r; KL(c r |c bg () is a measure of topic evolution vector c using KL divergence. r Distribution of background topics in the domain knowledge graph c bg The differences reflect the degree to which the target author's research topic deviates from the conventional topic in the field under component r; these information are combined according to component r from 1 to R to form the core fingerprint components.

[0160] The working principle of the above technical solution is as follows:

[0161] Based on the final similarity between papers, a hierarchical clustering algorithm is used to divide the target author's papers into different topic clusters. For example, papers researching "machine learning algorithms" are grouped into one cluster, and papers on "data mining applications" are grouped into another cluster.

[0162] A three-dimensional spatiotemporal correlation tensor is constructed by the number of time slices, geographical regions, and research topic clusters, where each element in the tensor represents the academic output intensity of the target author.

[0163] The three dimensions of the tensor are time slices (e.g., every 5 years is a slice), geographical regions (e.g., divided by country or continent), and research subject clusters; the tensor element T o,q,k This represents the intensity of an author's academic output within a specific time, region, and topic; for example, constructing a three-dimensional spatiotemporal relational tensor T∈R. N×M×KWhere N is the number of time slices divided by publication time (determined by time series analysis), M is the number of geographical regions (obtained through the resolution of the institution's address), and K is the number of research topic clusters (based on S). simf Cluster generation);

[0164] An improved CP decomposition algorithm is employed to approximately decompose the three-dimensional spatiotemporal correlation tensor into the sum of R components. Each component is constructed by outer product operations of the diagonal elements of the kernel tensor, a temporal pattern vector capturing the characteristics of the research cycle, a spatial distribution vector reflecting regional cooperation preferences, and a topic evolution vector representing changes in research direction. A weight matrix composed of the final similarity scores of the papers is introduced during decomposition, with weighted least squares (using the Frobenius norm to measure error) as the optimization objective, making the decomposition result more consistent with the characteristics of the original tensor. The three-dimensional tensor is decomposed into the sum of multiple components, each component consisting of the outer product of the temporal pattern vector, spatial distribution vector, topic evolution vector, and the diagonal elements of the kernel tensor. For example, the temporal pattern vector might capture that the authors focused on theoretical research from 2010 to 2015, and shifted to applied research after 2016.

[0165] Core components were extracted from the decomposed factor matrix to construct a spatiotemporal evolution trajectory matrix. This matrix contains the changing patterns of the authors' research interests over time and space, forming a unique "spatiotemporal fingerprint." For example, the matrix might show that the authors primarily researched artificial intelligence in Asia, while focusing on robotics in Europe.

[0166] The core fingerprint components are extracted from the decomposed factor matrix; and the spatiotemporal evolution trajectory matrix of the target author's research pattern is constructed, thereby extracting the spatiotemporal fingerprint features of the target author's research pattern.

[0167] For each decomposed component r, the maximum value of the time pattern vector is calculated. This value represents the intensity of the author's strongest research activity under that component, corresponding to a specific time slice. For example, if x r The maximum value was achieved on the 2015-2020 time slice, indicating that the authors invested the most in research on topics related to this component during this period; (The last part, "argmax(b...", appears to be incomplete and requires further context.) r Determine the spatial distribution vector b rThe index of the maximum value corresponds to the geographical region where the author is most inclined to collaborate under component r. For example, if the maximum value index corresponds to "North America," it indicates that the author collaborates most frequently with North American institutions in this research direction. Topic deviation feature extraction: KL divergence is used to measure the difference between the topic evolution vector and the distribution of topics in the domain context. The larger the KL divergence, the more the author's research topic deviates from the domain norm under this component. For example, in the field of computer vision, if an author has a high KL value, it may indicate that they have explored new research directions (such as cross-modal learning). Core fingerprint construction: The temporal intensity, spatial preference, and topic deviation of each component r are combined into triples to form core fingerprint components; the set of these components constitutes the author's unique "spatiotemporal fingerprint," reflecting the uniqueness of their research pattern.

[0168] The beneficial effects of the above technical solution are as follows:

[0169] Simultaneously capturing features across three dimensions—time, space, and theme—to comprehensively characterize an author's research patterns. For example, it can identify changes in an author's research focus across different time periods and regions; spatiotemporal fingerprint features are unique and can be used to detect anomalous papers; by analyzing time pattern vectors, the future research direction of an author can be predicted; for example, if a time pattern shows a continuous increase in research intensity in a certain field, it can be predicted that the author will continue to delve deeper into that field. The granularity of time slices and the method of geographical region division can be adjusted to adapt to the research characteristics of different disciplines. For example, in the field of high-energy physics, geography can be divided according to the location of large-scale experimental facilities; the extracted spatiotemporal features can be used to enhance domain knowledge graphs, providing richer information for academic navigation and recommendation systems. For example, recommending researchers or papers similar to the author's current spatiotemporal patterns; transforming complex tensor decomposition results into concise feature vectors, retaining key information while reducing dimensionality.

[0170] In one possible implementation, constructing the spatiotemporal evolution trajectory matrix of the target author's research pattern includes:

[0171] The time pattern vector sequence is processed using a long short-term memory network to obtain the first processing result, which means the evolution pattern of the target author's research activities in the time dimension.

[0172] A graph attention network is used to model the spatially distributed vector sequence to obtain a second processing result, which is the spatial correlation.

[0173] The first and second processing results are connected sequentially to form the spatiotemporal evolution trajectory matrix of the target author's research pattern.

[0174] The working principle of the above technical solution is as follows:

[0175] The temporal pattern vector sequence is processed using a Long Short-Term Memory (LSTM) network, which excels at capturing long-term dependencies in time series, thereby capturing the evolutionary pattern of the target author's research activities in the temporal dimension. The spatially distributed vector sequence is modeled using a Graph Attention Network (GAT), which can effectively process graph-structured data and mine spatial relationships. Finally, the results of these two processing steps are connected sequentially to form a spatiotemporal evolution trajectory matrix, which comprehensively reflects the research dynamics of the target author in the spatiotemporal dimension.

[0176] The effects of the above technical solution are as follows:

[0177] LSTM networks effectively handle the nonlinear evolution of research interests, identifying long-term trends and short-term fluctuations. For example, in the field of artificial intelligence, they can capture the phased transitions from "neural networks" to "deep learning" and then to "large models," avoiding the pattern breaks caused by time intervals in traditional methods.

[0178] The GAT network uses an attention mechanism to uncover potential connections between geographical regions and discover hidden collaboration preferences. For example, it can identify cyclical collaboration patterns among authors across Asia, Europe, and North America, or regional research clusters within specific fields. The spatiotemporal trajectory matrix integrates temporal evolution and spatial correlations to form a multidimensional feature representation.

[0179] In one possible implementation, the step of dynamically adjusting the judgment threshold based on spatiotemporal fingerprint features and using a domain knowledge graph to output the same-name author judgment result and confidence score includes:

[0180] In the domain knowledge graph, a concept hierarchy tree is built. The semantic distance is obtained by calculating the number of edges (path length) that a concept pair passes through from one concept to another in the concept hierarchy tree and dividing it by the number of edges of the longest path in the concept hierarchy tree (i.e., the maximum path length).

[0181] The threshold for determining authors with the same name is dynamically adjusted by taking into account both semantic distance and time factors.

[0182] A probabilistic calibration method is used to calculate the confidence score based on the similarity score, dynamic threshold, and the magnitude of the weight vector of related concepts in the domain knowledge graph. Finally, the result of the same-name author determination and the confidence score are output.

[0183] In some embodiments, the threshold for determining authors with the same name is dynamically adjusted by combining semantic distance and time factors; including:

[0184] Combining semantic distance and time factors, a normalization adjustment term is constructed.

[0185]

[0186] Where, dc (ge, gf) represents the semantic distance, reflecting the degree of semantic association between concepts ge and gf; te and tf represent the time of events related to concepts ge and gf (such as events related to the author's research activities, such as paper publication); n is the total number of concept pairs; σ is the time decay window.

[0187] The threshold for determining authors with the same name is dynamically adjusted based on the baseline threshold, sensitivity coefficient, and normalization adjustment term.

[0188]

[0189] Where Th0 is the baseline threshold; η is the sensitivity coefficient, 0 < η < 1. This is a normalization adjustment term.

[0190] The working principle of the above technical solution is as follows:

[0191] In the concept hierarchy tree of a domain knowledge graph, the path length (number of edges) between two concepts is calculated and normalized to a proportion relative to the longest path. For example, in a computer science knowledge graph, the path length between "machine learning" and "deep learning" is 2. If the longest path is 10, the semantic distance is 0.2.

[0192] A time decay factor is applied to the semantic distance of each concept pair, allowing the semantic association of recent events to have a greater impact on threshold adjustment. For example, two papers published 5 years apart will have a lower weight for their concept association than a paper published 1 year apart. Normalization adjustment term construction: The normalization adjustment term is obtained by averaging the time decay semantic distance of all concept pairs. This value reflects the combined semantic and temporal association strength of the currently compared set of papers. The judgment threshold is adjusted using a formula; when the semantic association between papers is strong and their publication dates are close, a threshold is set. As the threshold increases, a higher similarity is required to determine if the authors are the same person.

[0193] Beneficial Adaptive Capabilities: Leveraging the hierarchical structure of domain knowledge graphs, the decision threshold automatically adapts to the semantic characteristics of different disciplines. In the biomedical field, due to the more complex concept hierarchy and more refined semantic distance calculation, threshold adjustment is more sensitive; while in the mathematical field, threshold adjustment focuses more on time factors. Time-Sensitive Decision Making: The time decay mechanism ensures that recent research activities have a greater impact on the decision results. For example, when an author changes their research direction, the similarity threshold between old and new papers dynamically decreases, avoiding misjudgments caused by changes in research direction. Enhanced Confidence Measurement: By integrating semantic distance, time factors, and concept weights, the confidence score more accurately reflects the reliability of the decision; the dynamic threshold mechanism effectively addresses abrupt changes in research patterns. For example, when an author suddenly shifts to a new field, the threshold automatically decreases, reducing missed judgments; while when an author continues to delve deeply into a related field, the threshold increases, reducing misjudgments. Enhanced Interpretability: Each parameter in the decision process (such as semantic distance and time decay factor) has a clear physical meaning, making the results easier to interpret. For example, when the decision threshold increases, it can be traced back to specific concept associations and time factors. Parameter adaptability: By adjusting the sensitivity coefficient, it can adapt to different application scenarios.

[0194] This invention also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to perform the steps of the method described in this invention.

[0195] The working principle and effect of the above technical solution are the same as those in the embodiments of the present invention, and will not be repeated here.

[0196] This invention is described from the perspectives of purpose, effectiveness, progress, and novelty, and meets the functional enhancement and use requirements emphasized by the Patent Law. The above description and accompanying drawings are only preferred embodiments of this invention and are not intended to limit the invention. Therefore, all structures, devices, features, etc., that are similar to or identical to those of this invention, i.e., all equivalent substitutions or modifications made in accordance with the scope of this patent application, should fall within the scope of protection of this patent application.

Claims

1. A method for determining authors with the same name based on multimodal spatiotemporal feature fusion, characterized in that, The method includes: Retrieve metadata for all papers under the target author's name, including publishing institution, abstract, main text, publication date, citation relationships, and co-author information; We construct a multimodal knowledge graph and extract textual semantic features through a pre-trained language model, extract structural semantic features through a graph neural network, and extract research interest evolution features through temporal analysis. Cluster the names of publishing institutions to construct an institutional knowledge graph and calculate the academic relevance between institutions; By integrating multimodal knowledge graphs and unit knowledge graphs, the final similarity between papers is calculated. Construct a spatiotemporal correlation matrix and perform tensor decomposition to extract spatiotemporal fingerprint features of the target author's research patterns; Based on spatiotemporal fingerprint characteristics, the judgment threshold is dynamically adjusted through the domain knowledge graph, and the judgment result and confidence score of the same author are output.

2. The method for determining authors with the same name according to claim 1, characterized in that, The construction of the multimodal knowledge graph includes: Keywords are mapped to the semantic space of a pre-trained language model to generate context-aware text semantic vectors; Based on graph neural networks, a paper citation graph and a co-author social network graph are constructed to extract structural semantic features; We conduct time series analysis on the publication time series of papers to construct the research interest evolution feature vector of the target authors.

3. The method for determining authors with the same name according to claim 2, characterized in that, Keywords are mapped to the semantic space of a pre-trained language model to generate context-aware text semantic vectors; including: The abstract and main text of the paper are segmented into words, converting the text into multiple independent words; The independent words are preprocessed by restoring part-of-speech tags, filtering part-of-speech tags, and removing stop words to obtain words after the first preprocessing. The overall weight of a word is determined based on its positional weight in the paper, inverse document frequency, and time decay coefficient. Where α and β are position weight coefficients; α > β and α + β = 1; IDF(t) is the inverse document frequency; γ is the time decay coefficient; the value of γ ranges from [0.05, 0.2]. By employing an attention mechanism, text semantics are extracted and generated based on standardized keywords and the calculated comprehensive weights of words.

4. The method for determining authors with the same name according to claim 1, characterized in that, The process of clustering the names of publishing institutions and calculating their academic relevance to construct an institutional knowledge graph includes: String similarity calculation and hierarchical clustering are performed on the names of publishing institutions to form institution clusters; Calculate the academic correlation between institutions based on the co-occurrence relationship of papers and author collaboration relationship within the unit cluster; Construct a knowledge graph of units with units as nodes and academic relevance as edge weights.

5. The method for determining authors with the same name according to claim 4, characterized in that, The calculation of academic correlation between institutions based on paper co-occurrence relationships and author collaboration relationships within unit clusters includes: For each pair of units within a unit cluster, calculate the co-occurrence strength of their papers; Among them, S p To determine the co-occurrence intensity of papers between institutions, U i and U j Represents any two distinct units within a unit cluster; Unit U i A collection of all published papers; Unit U j A collection of all published papers; ∩ is the intersection symbol; ∪ is the union symbol; Unit U i and U j The number of papers co-authored; Unit U i and U j The union of all published papers; Extracting cross-unit co-author pairs C ij Calculate the cooperation strength S c And perform normalization processing; The academic relevance between institutions can be obtained by analyzing the co-occurrence intensity and collaboration intensity of papers. Academic relevance = λ·S p +(1-λ)·S cz Among them, S cz λ represents the normalized cooperation strength; λ is a coefficient, 0 < λ < 1.

6. The method for determining authors with the same name according to claim 1, characterized in that, The method of integrating multimodal knowledge graphs and unit knowledge graphs to calculate the similarity between papers includes: Calculate semantic similarity between papers based on text semantic vectors; Calculate the structural similarity between papers based on structural semantic features; Calculate the temporal evolution similarity between papers based on the evolutionary feature vector of research interests; The initial similarity between papers is obtained by weighted fusion of semantic similarity, structural similarity, and temporal evolution similarity; and the weights are dynamically allocated through an attention mechanism. The initial similarity is calibrated by the academic relevance between institutions; the final similarity between papers is then obtained. S simf =S sim ·(1+μ·log(1+S i,j ) Among them, S simf S represents the final similarity between the two papers. sim S represents the initial similarity between the two papers; μ is the coefficient, 0 < μ < 1; S i,j This represents the academic relevance between the institutions to which the two papers belong.

7. The method for determining authors with the same name according to claim 1, characterized in that, The process of constructing a spatiotemporal correlation matrix and performing tensor decomposition to extract spatiotemporal fingerprint features of the target author's research patterns includes: Based on the final similarity between papers, hierarchical clustering is performed on all papers under the target author to generate a set of research topic clusters. A three-dimensional spatiotemporal correlation tensor is constructed by the number of time slices, the number of geographical regions, and the number of research topic clusters, where each element in the tensor represents the academic output intensity of the target author. An improved CP decomposition algorithm is used to approximately decompose the three-dimensional spatiotemporal correlation tensor into the sum of multiple components. Each component is obtained by outer product operation from the diagonal elements of the kernel tensor, the temporal pattern vector that captures the characteristics of the research cycle, the spatial distribution vector that reflects the regional cooperation preference, and the theme evolution vector that represents the change of research direction. The core fingerprint components are extracted from the decomposed factor matrix, and the spatiotemporal evolution trajectory matrix of the target author's research pattern is constructed, thereby extracting the spatiotemporal fingerprint features of the target author's research pattern.

8. The method for determining authors with the same name according to claim 7, characterized in that, The spatiotemporal evolution trajectory matrix for constructing the target author's research pattern includes: The time pattern vector sequence is processed using a long short-term memory network to obtain the first processing result, which means the evolution pattern of the target author's research activities in the time dimension. A graph attention network is used to model the spatially distributed vector sequence to obtain a second processing result, which is the spatial correlation. The first and second processing results are connected sequentially to form the spatiotemporal evolution trajectory matrix of the target author's research pattern.

9. The method for determining authors with the same name according to claim 7, characterized in that, The process of dynamically adjusting the judgment threshold based on spatiotemporal fingerprint features and using a domain knowledge graph to output the judgment result and confidence score for authors with the same name includes: In the domain knowledge graph, a concept hierarchy tree is built. The semantic distance is obtained by calculating the number of edges traversed from one concept to another in the concept hierarchy tree and dividing it by the number of edges of the longest path in the concept hierarchy tree. The judgment threshold is dynamically adjusted by comprehensively considering semantic distance and time factors; Based on the similarity score, dynamic threshold, and the magnitude of the weight vector of related concepts in the domain knowledge graph, the confidence score is calculated, and the final output is the result of the same-name author determination and the confidence score.

10. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the method according to any one of claims 1-9.