Author name disambiguation method and related device

Through multi-dimensional similarity mining and fusion processing, especially in-depth analysis of collaboration networks, the problem of insufficient accuracy in author name disambiguation in existing technologies is solved, and higher matching accuracy is achieved.

CN120725014AActive Publication Date: 2025-09-30INST OF MEDICAL INFORMATION CHINESE ACAD OF MEDICAL SCI
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202511250929.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-03
Publication Date
2025-09-30
Estimated Expiration
2045-09-03

AI Technical Summary

Technical Problem

Existing technologies lack accuracy in author name disambiguation, especially in collaboration network analysis, where they fail to deeply explore network topology characteristics and temporal evolution patterns, resulting in limited ability to identify complex academic relationships.

Method used

A multi-dimensional similarity mining method is adopted, including the similarity mining of direct cooperation relationships, indirect network relationships, network structure topology characteristics and temporal network evolution of the cooperation network. Combined with dimensions such as identity identification code, name, organization and time, a weighted graph is constructed through the cooperation network for similarity mining, and multi-dimensional similarity fusion processing is performed.

Benefits of technology

The accuracy of author name disambiguation has been significantly improved, and higher matching accuracy has been achieved through a deep understanding of the topological characteristics and dynamic evolution patterns of the collaboration network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120725014A_ABST
    Figure CN120725014A_ABST
Patent Text Reader

Abstract

The invention discloses an author name disambiguation method and a related device. The method comprises the following steps: acquiring author information to be analyzed and a candidate information set; performing multi-dimensional similarity mining on the plurality of pieces of author information in the candidate information set and the author information to be analyzed; the cooperative network is one of a plurality of dimensions. And performing fusion processing on a multi-dimensional similarity mining result between the same author information in the candidate information set and the author information to be analyzed to obtain a multi-dimensional similarity fusion result, and obtaining a matching result of the author to be analyzed and the authors in the candidate information set according to the multi-dimensional similarity fusion result. According to the scheme, the effectiveness of the cooperation network in the author name disambiguation aspect is improved, deep understanding of structural topological characteristics and dynamic mastering of a network evolution mode are achieved, compared with the prior art, the author name disambiguation method has the advantage of higher accuracy, and author name disambiguation accuracy can be further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a method for disambiguating author names and related devices. Background Art

[0002] With the continuous progress of society and the rapid development of science and technology, the volume of literature across the internet is exploding. This vast amount of literature comes in a variety of forms, including papers, patents, funding projects, software copyrights, and monographs. A number of factors can lead to different author information in published works by the same author. Therefore, author name disambiguation is necessary to accurately identify and match name variants of the same author appearing in different documents. There are many factors that can lead to different author information, such as changes in the author's institution, the presence of the same name in different authors, and name spelling variations. Currently, research on author name disambiguation primarily relies on rule-based string matching and simple machine learning methods. Existing methods typically only consider basic features such as name and institution, resulting in accuracy rates generally ranging from 70-85%, and processing time is relatively long. In particular, existing technologies for collaborative network analysis remain limited to simply counting co-authors, failing to delve deeper into network topology and temporal evolution patterns. This results in limited ability to identify complex academic relationships, and the accuracy of author name disambiguation remains to be improved. Summary of the Invention

[0003] Based on the above problems, this application provides an author name disambiguation method and related devices, the purpose of which is to improve the accuracy of author name disambiguation.

[0004] The embodiments of this application disclose the following technical solutions: In a first aspect, the present application provides a method for disambiguating author names, the method comprising: Obtaining a piece of author information to be analyzed and a candidate information set; the candidate information set includes multiple pieces of author information; the author information to be analyzed includes a list of collaborators; Performing multiple-dimensional similarity mining on the multiple pieces of author information and the author information to be analyzed, respectively; wherein the collaboration network is one of the multiple dimensions, wherein the authors are represented by nodes in the collaboration network, and the collaboration relationships are represented by lines between the nodes; the similarity mining of the collaboration network includes at least one of the following: direct collaboration relationship similarity mining based on the collaboration network, indirect network relationship similarity mining based on the collaboration network, network structure topology feature similarity mining based on the collaboration network, or temporal network evolution similarity mining based on the collaboration network; fusing the results of multi-dimensional similarity mining between the same author information in the candidate information set and the author information to be analyzed to obtain a multi-dimensional similarity fusion result; According to the fusion results of the multiple dimensions of similarity between the multiple pieces of author information and the author information to be analyzed, a matching result between the author to be analyzed and the authors in the candidate information set is obtained.

[0005] A second aspect of the present application provides an author name disambiguation device, the device comprising: An information acquisition module is used to acquire a piece of author information to be analyzed and a candidate information set; the candidate information set includes multiple pieces of author information; the author information to be analyzed includes a list of collaborators; a multi-dimensional similarity mining module for performing multi-dimensional similarity mining on the multiple pieces of author information and the author information to be analyzed; wherein the collaborative network is one of the multiple dimensions, wherein the authors are represented by nodes in the collaborative network and the collaborative relationships are represented by lines between the nodes; and similarity mining of the collaborative network includes at least one of the following: direct collaborative relationship similarity mining based on the collaborative network, indirect network relationship similarity mining based on the collaborative network, network structure topology feature similarity mining based on the collaborative network, or temporal network evolution similarity mining based on the collaborative network; A multi-dimensional similarity fusion module is used to fuse the results of multi-dimensional similarity mining between the same author information in the candidate information set and the author information to be analyzed to obtain a multi-dimensional similarity fusion result; The matching module is used to obtain a matching result between the author to be analyzed and the authors in the candidate information set based on the fusion results of the multiple dimensions of similarity between the multiple author information and the author information to be analyzed.

[0006] A third aspect of the present application provides an author name disambiguation device, the device comprising: a processor and a memory communicatively connected to each other; The memory stores a computer program; The processor is configured to run the computer program to implement the author name disambiguation method as described in the first aspect.

[0007] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when processed, implements the steps of the author name disambiguation method as described in the first aspect.

[0008] Compared with the existing technology, this application has the following beneficial effects: In the technical solution of the present application, in order to achieve accurate disambiguation of author names, first obtain an author information to be analyzed and a candidate information set; the candidate information set includes multiple author information; the author information to be analyzed includes a list of collaborators. Then, multiple dimensions of similarity mining are performed between the multiple author information and the author information to be analyzed respectively; wherein, the cooperation network is one of the multiple dimensions, and the authors are represented by nodes in the cooperation network, and the cooperation relationship is represented by the lines between the nodes. Thereafter, the results of the multiple dimensions of similarity mining between the same author information in the candidate information set and the author information to be analyzed are fused to obtain the multiple dimensions of similarity fusion results. Finally, based on the multiple dimensions of similarity fusion results between the multiple author information and the author information to be analyzed, the matching result of the author to be analyzed and the author in the candidate information set is obtained.

[0009] In the technical solution of the present application, consideration is given to similarity mining of author information in multiple dimensions, which involves mining of collaborative networks. And in the process of mining collaborative networks, not only can the similarity of direct collaborative relationships and the similarity of indirect network relationships be mined, but it is also specifically proposed to mine the similarity of network structure topology features in the collaborative network and the similarity of temporal network evolution, which greatly enhances the effectiveness of the collaborative network in author name disambiguation. The mining of the similarity of network structure topology features and the similarity of temporal network evolution can achieve a deep understanding of the structural topology features and a dynamic grasp of the network evolution pattern. Therefore, in the present application, the fusion results of the similarity mining results of the collaborative network dimension have a higher accuracy advantage than the existing technology, and can achieve a further improvement in the accuracy of author name disambiguation. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0011] Figure 1 A flowchart of a method for disambiguating author names provided in an embodiment of the present application; Figure 2 This is the implementation architecture diagram of the author name disambiguation method; Figure 3 Schematic diagram of four aspects of collaborative network similarity mining proposed in the embodiment of this application; Figure 4 A schematic diagram of mining similarity in six dimensions for author information and obtaining the corresponding similarity mining results; Figure 5This is a diagram illustrating an implementation architecture of an author name disambiguation method provided in an embodiment of the present application; Figure 6 A schematic diagram of the structure of an author name disambiguation device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0012] While research on author name disambiguation has focused on collaboration networks, it generally remains limited to simply counting co-authors, leaving the study of collaboration networks inadequate. Consequently, existing techniques fail to fully capture the knowledge directly or implicitly reflected in collaboration networks, resulting in limited ability to identify complex academic relationships and, in turn, hindering the effectiveness of collaboration networks for author name disambiguation.

[0013] After research, the inventors proposed a method for disambiguating author names and related devices. In the technical solution of this application, attention is paid to similarity mining of multiple dimensions of author information, which also includes similarity mining of cooperation networks. Similarity mining of cooperation networks can be started from multiple aspects, for example, similarity mining of direct cooperation relationships, similarity mining of indirect network relationships, similarity mining of network structure topological features, and similarity mining of temporal network evolution. By mining the similarity of one or more aspects of the cooperation network and integrating the results of similarity mining in other dimensions, it is possible to obtain the similarity fusion results of multiple dimensions between the same author information in the candidate information set and the author information to be analyzed. Finally, based on the similarity fusion results of multiple dimensions of the multiple author information and the author information to be analyzed, the matching results of the author to be analyzed and the authors in the candidate information set are compared and determined. The matching result reflects the disambiguation of the author name of the author to be analyzed.

[0014] In order to help those skilled in the art better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.

[0015] It should be noted that before the author name disambiguation method is formally implemented, the embodiment of the present application can use an intelligent two-level adaptive cache to initialize the system. For example, create a memory LRU cache (3000 items by default), set an adaptive TTL mechanism (dynamically adjust 7-30 days according to the frequency of data access); establish a SQLite persistent cache, use the WAL mode to optimize concurrent access, and set a 256MB memory mapping; implement cache hit rate statistics and performance monitoring, and automatically expand the capacity when the hit rate is lower than 70%. The two-level adaptive cache is an intelligent cache system that combines the memory LRU cache and the SQLite persistent cache. It can automatically adjust the cache strategy according to the access pattern and data characteristics, and supports the separate management of the query result cache and the calculation result cache.

[0016] See also Figure 1 , which is a flow chart of an author name disambiguation method provided in an embodiment of the present application. Figure 1 As shown, the method includes: S101. Obtain a piece of author information to be analyzed and a candidate information set.

[0017] In this application, in order to achieve disambiguation of author names, firstly, a piece of author information to be analyzed and a candidate information set are obtained, wherein the candidate information set includes multiple pieces of author information. Figure 2 This is the implementation architecture diagram of the author name disambiguation method, which can be found in Figure 2 For example, the candidate information set includes N pieces of author information, which can be referred to as "author information entry 1", "author information entry 2", ..., "author information entry N" for easy distinction.

[0018] This application proposes an example implementation method for obtaining a candidate information set. Before obtaining the candidate information set, first identify a piece of author information to be analyzed, and assume that the author pointed to by the author information to be analyzed is the author for whom author name disambiguation is required. Based on the obtained author information to be analyzed, the method for obtaining the candidate information set may include: In the author information database, author information is matched based on the surname and initials of the author information to be analyzed, obtaining a first matching information set that meets the matching criteria; within the first matching information set, author information is matched based on the institution keywords and year range of the author information to be analyzed, obtaining a second matching information set that meets the matching criteria; within the second matching information set, author information is matched based on the identity identification code of the author information to be analyzed, obtaining a third matching information set that meets the matching criteria as a candidate information set. It can be seen that in this embodiment of the present application, the candidate information set is obtained through a three-level screening strategy based on surname and initials, institution keywords and year range, and identity identification code.

[0019] From the above description, it can be known that in order to screen out candidate information sets for similarity mining and matching with the author information to be analyzed, this application can adopt a three-level screening method. The first level starts from the surname and initials and conducts a round of screening; the second level starts from the institution keywords and year range and conducts another round of screening; the third level starts from the identity identification code and conducts another round of screening. The candidate information set finally constructed through the three-level screening has a high degree of matching with the author information to be analyzed in terms of name, initials, institution, year range and identity identification code, which is conducive to the rapid positioning of author information with a high degree of correlation and reducing the candidate information set to a smaller amount of data. In an optional example, the first level of screening controls the number of information entries in the first matching information set to within 500 entries, and the second level of screening controls the number of information entries in the second matching information set to within 200 entries.

[0020] In actual applications, before filtering the author information that meets the matching conditions from the author information database for the author information to be analyzed and adding it to the candidate information set, the technical solution of the present application can also perform a data quality assessment on the author information to be analyzed. For example, the author information to be analyzed can be subjected to integrity check, format verification and outlier detection. The received author information to be analyzed includes at least a list of collaborators. In addition, it can also include: standardized names, a list of institutions, a year of publication, and an identification code. For information of different dimensions in the above information, the method of data quality assessment may be different, and can be specifically performed according to the designed integrity check method, format verification method, and outlier detection method, which are not limited here.

[0021] After performing the aforementioned data quality assessment, a query fingerprint can be created for each author information entry for cache indexing. The query fingerprint can be in the form of MD5 (name + institution keyword + year range). MD5 stands for the MD5 Message-Digest Algorithm, a widely used cryptographic hash function that produces a 128-bit (16-byte) hash value. Creating and caching a query fingerprint for each author information entry facilitates efficient indexing of matching author information when selecting candidate information sets, enabling efficient candidate information set creation.

[0022] To intelligently construct candidate information sets and optimize author information queries, a query plan generator can be constructed to automatically select the optimal query path based on data distribution. Alternatively, a batch query strategy can be adopted to avoid the significant impact of large query results on memory. Furthermore, a query performance monitoring mechanism can be established. When query time exceeds a threshold, an alternative query strategy is automatically switched to, thus avoiding long wait times before disambiguation and improving the efficiency of author name disambiguation.

[0023] In the description of the query plan generator above, the "data distribution" mentioned refers to statistical information such as name frequency distribution (e.g., the number of occurrences of high-frequency surnames like Smith and Wang), institutional distribution characteristics (number of authors per institution, country distribution ratio), ID coverage distribution (ORCID and Scopus ID coverage ratio), and time distribution characteristics (author activity in different years). The "query path" refers to different query strategies designed based on these distribution characteristics, such as: When the input author has a reliable ID and the ID coverage is high, select "ID priority path" (use ID for exact matching first and then verify name and institution); When encountering a high-frequency surname, choose the "name-institution joint path" (using both name and institution keywords to filter to narrow the candidate set); When the information about collaborators is rich, the “collaboration network priority path” is selected (first build the candidate set through the collaboration network and then verify the name organization).

[0024] The query plan generator analyzes the input author information (whether there is an ID, whether the surname is common, whether there are many collaborators, etc.), combines data distribution statistics and historical query performance, and automatically calculates the expected efficiency score of each path. It then selects the optimal path to execute the query and dynamically adjusts the strategy based on the actual query results, thereby achieving intelligent query optimization and load balancing.

[0025] S102: Perform multi-dimensional similarity mining on the multiple pieces of author information and the author information to be analyzed.

[0026] In this application, each piece of author information in the candidate information set needs to be mined for similarity in multiple dimensions with the aforementioned author information to be analyzed.

[0027] The multiple dimensions include the dimension of the collaboration network. The collaboration network is constructed as an undirected weighted graph G = (V, E, W), where V represents the set of author nodes, E represents the set of collaboration edges, and W represents the set of edge weights. Edge weights are calculated based on the frequency of collaboration and a time decay factor: w(i, j) = count(i, j) × exp(-λ × (current year – collaboration year)), where i and j represent two author nodes in the graph, and λ = 0.1 / year. The collaboration network is actually a complete graph network structure, and can be understood as a global collaboration network, rather than a radial network centered around a single author. Specifically, a global collaboration network is a global graph encompassing all author nodes, where any two collaborating authors are connected by an edge, and the edge weights represent the strength of the collaboration. When analyzing a specific author, one can extract that author's ego network, which can be understood as a local network or subnetwork. This ego network is centered in a radial pattern, encompassing the author's first-degree neighbors (direct collaborators), second-degree neighbors, and third-degree neighbors.

[0028] When mining the similarity between two authors in the dimension of collaborative networks, the network topology characteristics (such as local clustering coefficient, structural hole location, network centrality) and temporal network evolution of the two authors' respective self-networks are compared, rather than a simple radial matching.

[0029] That is, in the embodiment of the present application, a global collaboration network is constructed based on the collaborator list, and when analyzing a specific author, an ego network centered on the author is extracted to perform network feature analysis.

[0030] In the example implementation, the collaborative network is constructed with the following key points: Collaborative relationships are extracted and a weighted directed graph is constructed. Weights in the weighted directed graph are calculated based on the product of collaboration frequency and a time decay factor. Network topology features such as node degree centrality, clustering coefficient, betweenness centrality, and shortest path length can also be calculated. Furthermore, an author influence scoring model can be constructed: Score = α × degree centrality + β × clustering coefficient + γ × publication quality.

[0031] Figure 3 This is a schematic diagram of four aspects of collaborative network similarity mining proposed in this application embodiment. Figure 3 As shown, with respect to the dimension of cooperative network, the embodiment of the present application proposes that similarity mining can be performed in at least one of the following four aspects: direct cooperative relationship similarity mining based on cooperative network, indirect network relationship similarity mining based on cooperative network, network structure topology feature similarity mining based on cooperative network, or temporal network evolution similarity mining based on cooperative network.

[0032] Among them, direct collaboration similarity mining focuses on: the similarity of the collaboration relationship directly reflected by the collaboration networks of the two authors being compared; Indirect network relationship similarity mining focuses on: the potential cooperation relationship or academic circle connection reflected by the cooperation networks of two authors compared with each other; The similarity mining of network structure topology features focuses on the similarity of the structural topology features of the collaboration networks of two authors being compared with each other, and deeply analyzing the similarity of authors through network structure; Temporal network evolution similarity mining focuses on the similarity of dynamic information developed over time as reflected by the collaboration networks of two authors being compared.

[0033] The similarity mining in the dimension of collaborative network has been briefly described above. The specific mining methods for the above four aspects will be introduced in detail later, so we will not go into details here.

[0034] In addition to the collaboration network, step S102 performs similarity mining in multiple dimensions, where the multiple dimensions may also include: identity identification code, name, organization, time, and alias. Figure 4 A schematic diagram of similarity mining of author information in six dimensions and obtaining corresponding similarity mining results. Regarding similarity mining in these dimensions, it can be performed based on the specific information types contained in multiple author information in the candidate information set and the information type of the author to be analyzed contained in the author information to be analyzed, or further processing can be performed on the basis of the specific information type contained in the author information to complete the similarity mining work. Through S102, the results of identity code similarity mining, name similarity mining, organization similarity mining, cooperation network similarity mining, time similarity mining and alias similarity mining between each author information in the candidate information set and the author information to be analyzed can be obtained. These similarity mining results of different dimensions will be fused in step S103 described below according to their correspondence with the author information in the candidate information set.

[0035] S103: Fusing the results of multi-dimensional similarity mining between the same author information in the candidate information set and the author information to be analyzed to obtain a multi-dimensional similarity fusion result.

[0036] In an optional implementation, in combination with the example of similarity mining in the above six dimensions, this step can be to fuse the results of identity code similarity mining, name similarity mining, organization similarity mining, cooperation network similarity mining, time similarity mining and alias similarity mining between the same author information in the candidate information set and the author information to be analyzed according to the weights of the corresponding dimensions to obtain the similarity fusion results of multiple dimensions between the same author information and the author information to be analyzed.

[0037] The results of identity code similarity mining, name similarity mining, organization similarity mining, cooperation network similarity mining, time similarity mining and alias similarity mining can each be represented by a numerical value, so they can be fused in a weighted manner to finally obtain the similarity fusion results of multiple dimensions.

[0038] For example, the weight of the identity code dimension is 0.25, the weight of the name dimension is 0.4, the weight of the organization dimension is 0.15, the weight of the collaboration network dimension is 0.15, the weight of the time dimension is 0.03, and the weight of the alias dimension is 0.02. It should be noted that the weight of each dimension can be set or adjusted according to different fusion requirements, or the focus or secondary consideration of the similarity of a specific dimension. The above is only an example of the weight setting and is not a limitation on the specific weight value.

[0039] S104: Obtain matching results between the author to be analyzed and the authors in the candidate information set based on the fusion results of the multiple dimensions of similarity between the multiple pieces of author information and the author information to be analyzed.

[0040] Each author information in the candidate information set can be used to mine the similarity of multiple dimensions with the author information to be analyzed, and the similarity fusion results of multiple dimensions can be obtained, such as Figure 2 The "Multi-dimensional Similarity Fusion Result 1," "Multi-dimensional Similarity Fusion Result 2," ..., and "Multi-dimensional Similarity Fusion Result N" are shown in the figure. After obtaining these multi-dimensional similarity fusion results, their numerical values ​​can be compared to determine the matching result between the author to be analyzed and the authors in the candidate information set. For example, the author information with the highest multi-dimensional similarity fusion result value with the author to be analyzed is determined to be the matching author information. This author information is considered to correspond to the same author as the author to be analyzed, thus completing the author name disambiguation.

[0041] In the technical solution of the present application, consideration is given to similarity mining of author information in multiple dimensions, which involves mining of collaborative networks. And in the process of mining collaborative networks, not only can the similarity of direct collaborative relationships and the similarity of indirect network relationships be mined, but it is also specifically proposed to mine the similarity of network structure topology features in the collaborative network and the similarity of temporal network evolution, which greatly enhances the effectiveness of the collaborative network in author name disambiguation. The mining of the similarity of network structure topology features and the similarity of temporal network evolution can achieve a deep understanding of the structural topology features and a dynamic grasp of the network evolution pattern. Therefore, in the present application, the fusion results of the similarity mining results of the collaborative network dimension have a higher accuracy advantage than the existing technology, and can achieve a further improvement in the accuracy of author name disambiguation.

[0042] As mentioned earlier, within the collaborative network dimension, similarity mining can be performed in at least one of the following four areas: direct collaborative relationship similarity mining based on the collaborative network, indirect network relationship similarity mining based on the collaborative network, network structure topology feature similarity mining based on the collaborative network, or temporal network evolution similarity mining based on the collaborative network. The following describes the similarity mining process for each of these four areas. For ease of introduction, the author to whom a piece of author information in the candidate information set belongs is referred to as the target author, the local network extracted from the global collaborative network for the author to be analyzed is referred to as the first collaborative network, and the local network extracted from the global collaborative network for the target author is referred to as the second collaborative network. This allows for a more concise and accurate description of the similarity mining process for collaborative networks in various aspects.

[0043] (1) Mining the similarity of network structure topological features.

[0044] In the embodiment of the present application, the target author information in the candidate information set and the author information to be analyzed are subjected to similarity mining of network structure topology features based on the collaborative network, including: Analyze the network structure topology characteristics of the first cooperation network to obtain the local clustering coefficient, structural hole location and network centrality of the first cooperation network; analyze the network structure topology characteristics of the second cooperation network to obtain the local clustering coefficient, structural hole location and network centrality of the second cooperation network; Based on the local clustering coefficient of the first collaboration network and the local clustering coefficient of the second collaboration network, the local clustering similarity between the target author and the author to be analyzed is calculated; based on the structural hole position of the first collaboration network and the structural hole position of the second collaboration network, the structural hole position similarity between the target author and the author to be analyzed is calculated; based on the network centrality of the first collaboration network and the network centrality of the second collaboration network, the network centrality similarity between the target author and the author to be analyzed is calculated; Finally, based on the local clustering similarity, structural hole position similarity and network centrality similarity, the network structure topology feature similarity mining results between the target author and the author to be analyzed are calculated.

[0045] As an example, the local clustering similarity, structural hole position similarity and network centrality similarity may each have a fusion weight, and the local clustering similarity, structural hole position similarity and network centrality similarity are weightedly summed based on the fusion weights of the three. The final calculated result is used as the network structure topology feature similarity mining result between the target author and the author to be analyzed.

[0046] Among them, the local clustering coefficient is used to measure the closeness of the author's network; the location of structural holes is used to identify the author's bridge role in the network; and the network centrality is used to achieve a comprehensive evaluation of degree centrality and betweenness centrality.

[0047] The local clustering coefficient is calculated as: C(i) = 2×ei / (ki×(ki-1)), where ei is the actual number of edges between neighbors of node i, and ki is the degree of node i.

[0048] The location of the structural holes is calculated using the effective scale:

[0049] Where p iq represents the normalized connection strength between node i and node q, p qj Represents the normalized connection strength between node q and node j, and the sum The sum of all nodes q except nodes i and j. The larger the effective scale value ES(i), the more important the structural hole position of node i in the network is, that is, there is a lack of direct connection between the neighborhoods of node i, and node i plays an important bridge role. In the formula, is the sum of all neighbor nodes j of node i, and is the “uniqueness” index corresponding to each neighbor node j, that is, the degree to which node j does not depend on i through other nodes. The effective scale value of node i is obtained by adding the “uniqueness” index corresponding to each neighbor node j.

[0050] Network centrality adopts a weighted combination of degree centrality and betweenness centrality: NC(i) = α×DC(i) + β×BC(i), where α+β=1, DC represents degree centrality, and BC represents betweenness centrality.

[0051] In the embodiments of the present application, the network structure topology feature analysis driven by graph neural network, including advanced network features such as degree centrality, clustering coefficient, and structural hole analysis, fills the technical gap in the in-depth analysis of collaborative networks.

[0052] (2) Mining the similarity of temporal network evolution.

[0053] In the embodiment of the present application, the target author information in the candidate information set and the author information to be analyzed are subjected to temporal network evolution similarity mining based on the collaborative network, including: Performing a time window analysis based on the first cooperative network to obtain a first time window analysis result and identifying a dynamic change pattern of the first cooperative network; performing a time window analysis based on the second cooperative network to obtain a second time window analysis result and identifying a dynamic change pattern of the second cooperative network; Based on the analysis results of the first time window and the analysis results of the second time window, obtaining the network time window similarity between the target author and the author to be analyzed; Based on the dynamic change pattern of the first collaboration network and the dynamic change pattern of the second collaboration network, obtaining the similarity of the network dynamic change pattern between the target author and the author to be analyzed; Based on the network time window similarity and the network dynamic change pattern similarity, the temporal network evolution similarity mining results between the target author and the author to be analyzed are calculated.

[0054] As an example, the network time window similarity and the network dynamic change pattern similarity may each have a fusion weight. Based on the fusion weights of the two, the network time window similarity and the network dynamic change pattern similarity are weighted and summed, and the final calculated result is used as the temporal network evolution similarity mining result between the target author and the author to be analyzed.

[0055] It is understandable that the list of collaborators in the author information can not only reflect the specific objects of the collaboration, but also that the collaborative relationship has a time attribute, for example, a work was co-published in July 2021, and another work was co-published in November 2022. Therefore, by analyzing the time window of the collaborative network, it is possible to identify the dynamic characteristics of the collaborative relationship, and then to mine the similarity of the network time window and the similarity of the network dynamic change pattern. It is understandable that if it is the same author, its network time window and network dynamic change pattern should tend to be similar. In this application, the temporal network evolution similarity mining introduced above can more accurately capture the potential dynamic change attributes of the network than the existing technology for mining the collaborative network. Taking this as an aspect of the collaborative network similarity mining can improve the accuracy and utility of the collaborative network in author name disambiguation.

[0056] (3) Mining the similarity of direct cooperative relationships.

[0057] In the embodiment of the present application, the target author information in the candidate information set and the author information to be analyzed are subjected to direct cooperation relationship similarity mining based on the cooperation network, including: Based on the first collaboration network, the first collaboration intensity and first collaboration density of the author to be analyzed and each author in the collaborator list are calculated; based on the second collaboration network, the second collaboration intensity and second collaboration density of the target author and each author in the collaborator list are calculated; Obtain the similarity of the cooperation intensity between the target author and the author to be analyzed based on the first cooperation intensity and the second cooperation intensity of the author to be analyzed and the target author with the same author in the collaborator list; Obtain the similarity of the collaboration density between the target author and the author to be analyzed based on the first collaboration density and the second collaboration density between the author to be analyzed and the target author and the same author in the collaborator list; Based on the cooperation intensity similarity and cooperation density similarity, the direct cooperation relationship similarity mining results between the target author and the author to be analyzed are calculated.

[0058] As an example, the cooperation intensity similarity and the cooperation density similarity may each have a fusion weight, and the cooperation intensity similarity and the cooperation density similarity are weighted and summed based on the fusion weights of the two. The final calculated result is used as the direct cooperation relationship similarity mining result between the target author and the author to be analyzed.

[0059] The following is an example calculation method for cooperation intensity and cooperation density: Collaboration intensity = number of collaborations × time decay factor × document quality weight; Collaboration density = number of common documents / max[number of documents by author A, number of documents by author B].

[0060] Among them, author A and author B in the formula of collaboration density refer to the two authors used to calculate collaboration density.

[0061] From the above formula, we can see that the intensity of cooperation is positively correlated with the number of collaborations, the time decay factor and the weight of the literature quality.

[0062] The longer the collaboration, the smaller the time decay factor; the more recent the collaboration, the larger the time decay factor. This design is because recent collaborations better reflect the author's current academic connections and activity, and have a higher reference value in author disambiguation, while older collaborations may no longer be active or have lower relevance. A typical time decay function is implemented as: decay_factor = exp(-λ × (current year - cooperation year)); Where decay_factor represents the time decay factor, and λ represents the decay parameter. Taking λ = 0.1 / year, the time decay factor for a partnership five years ago becomes exp(-0.5) ≈ 0.61, retaining approximately 60% of the weight. The time decay factor for a partnership ten years ago decreases to approximately 0.37. When the partnership occurred in the current year, the time decay factor is 1.0 (the maximum value). As the time difference increases, the time decay factor gradually decreases, approaching 0. This gives more weight to recent partnerships and less weight to older partnerships in the partnership intensity calculation.

[0063] Collaboration density is positively correlated with the number of common documents and negatively correlated with the maximum number of documents by two authors.

[0064] In this application, taking into account the direct cooperation relationship reflected by the cooperation intensity and cooperation density, by comparing the similarity of such direct cooperation relationships, we can accurately grasp the commonalities of the cooperation relationship between the target author and the author to be analyzed, and then assist in accurately identifying the same author and realizing author name disambiguation.

[0065] (4) Indirect network relationship similarity mining.

[0066] In the embodiment of the present application, the target author information in the candidate information set and the author information to be analyzed are subjected to indirect network relationship similarity mining based on the collaborative network, including: Identify the second-degree connection nodes and the third-degree connection nodes of the author to be analyzed from the first collaboration network; and identify the second-degree connection nodes and the third-degree connection nodes of the target author from the second collaboration network.

[0067] Based on the second-degree connection nodes of the author to be analyzed and the second-degree connection nodes of the target author, the second-degree connection similarity between the target author and the author to be analyzed is obtained; based on the third-degree connection nodes of the author to be analyzed and the third-degree connection nodes of the target author, the third-degree connection similarity between the target author and the author to be analyzed is obtained; Based on the second-degree connection similarity and the third-degree connection similarity, the indirect network relationship similarity mining results between the target author and the author to be analyzed are calculated.

[0068] The second-degree connection nodes of the author to be analyzed in the first collaborative network can be constructed based on the collaborator list of the first-degree connection nodes of the author to be analyzed, and the third-degree connection nodes of the author to be analyzed can be constructed based on the collaborator list of the second-degree connection nodes of the author to be analyzed. Similarly, the second-degree connection nodes of the target author in the second collaborative network can be constructed based on the collaborator list of the first-degree connection nodes of the target author, and the third-degree connection nodes of the target author can be constructed based on the collaborator list of the second-degree connection nodes of the target author.

[0069] As an example, the second-degree connection similarity and the third-degree connection similarity may each have a fusion weight, and the second-degree connection similarity and the third-degree connection similarity are weightedly summed based on the fusion weights of the two. The final calculated result is used as the indirect network relationship similarity mining result between the target author and the author to be analyzed.

[0070] Second-degree connection similarity reflects the strength of connections between co-authors, while third-degree connection similarity can be used to identify potential academic relationships. Therefore, the indirect network relationship similarity mining results calculated based on second-degree and third-degree connection similarity effectively analyze the commonalities between the indirect network relationships of the authors being analyzed and the target authors, deepening and expanding the superficial understanding of the collaborative network, thereby facilitating the accurate identification of the same author.

[0071] After completing the similarity mining based on the collaborative network in the above four aspects, the similarity mining results of the network structure topology features, the similarity mining results of the temporal network evolution, the similarity mining results of the direct collaborative relationship, and the similarity mining results of the indirect network relationship can be fused to obtain the results of the collaborative network similarity mining. For example, the similarity mining results of the above four aspects each have a fusion weight, and the similarity mining results of the four aspects are weighted and summed using the corresponding fusion weights, and the final result is used as the result of the collaborative network similarity mining. The similarity mining results of the collaborative network and the similarity mining results of other dimensions can be fused with the results of similarity mining of several other dimensions using step S103 of the above embodiment method to obtain a multi-dimensional similarity fusion result.

[0072] The previous section introduces the implementation of similarity mining in the collaborative network dimension. Figure 4 The similarity mining methods of the other five dimensions shown are explained.

[0073] (1) Similarity mining based on institutions.

[0074] In this application, an institutional hierarchy identification system and a standardized mapping of country codes are pre-established, so that the diverse expressions and hierarchical relationships of institutional names can be effectively handled.

[0075] In the embodiment of the present application, the target author information in the candidate information set and the author information to be analyzed are subjected to similarity mining based on the institution, including: The institution information in the target author information and the institution information in the author information to be analyzed are standardized and preprocessed separately. This standardization preprocessing includes removing institution type prefixes, extracting core institution keywords, and applying country codes for standardized mapping. Removing institution type prefixes can include removing (e.g., Department of / School of). Extracting core institution keywords can include removing common terms such as "university" and "college." Applying country codes for standardized mapping ensures accurate matching of corresponding countries and geographic locations.

[0076] The Jaccard similarity and fuzzy matching similarity are calculated for the keywords of the two sets of institutional information after standardized preprocessing; the preliminary institutional similarity is calculated based on the Jaccard similarity and fuzzy matching similarity; the preliminary institutional similarity is optimized based on the matching of the geographical locations and institutional hierarchies of the two sets of institutional information, and the optimized similarity result is used as the result of institutional similarity mining between the target author and the author to be analyzed.

[0077] As described above, the institution-based similarity mining process is mainly divided into two stages, which can also be understood as adopting a hierarchical matching strategy: first, the preliminary institution similarity is obtained by weighted averaging the Jaccard similarity (the ratio of keyword intersection to union, |A∩B| / |A∪B|, A∩B represents the keyword intersection of two sets of institution information, and A∪B represents the keyword union of two sets of institution information) and fuzzy matching similarity (string token matching based on edit distance); then geographical and institutional hierarchical optimization is performed.

[0078] In an example implementation, optimizing based on geographic location and organizational level may include: Geographic location optimization adds a geographic adjustment factor based on country matching. For example, +0.1-0.2 for the same country, +0.05 for different countries with a history of international cooperation, and -0.1 for completely different countries); Institutional level optimization adds an institutional level adjustment factor based on level matching. For example, within the same level (e.g., university vs. department), the adjustment factor is +0.1, while across levels (e.g., university vs. department) the adjustment factor is -0.05 to 0.15.

[0079] The calculation formula of the optimized similarity result can be expressed as: Optimized similarity = min[1.0, initial similarity × (1 + geographic adjustment factor + institutional level adjustment factor)]; For the similarity mining of institutions, the above-mentioned hierarchical matching strategy not only ensures the textual similarity basis, but also incorporates the realistic constraints of the geographical distribution and organizational structure of the institutions, so that the obtained institutional similarity mining results have higher numerical rationality, thereby effectively achieving accurate author name disambiguation.

[0080] The embodiment of the present application can also implement organization alias clustering through an algorithm, thereby automatically identifying different expressions of the same organization.

[0081] We previously discussed the concepts of institutional hierarchy and standardized country code mapping. To more accurately mine similarity within the institutional dimension, this application proposes constructing an institutional hierarchy identification system, a standardized country code mapping table, and a database of institutional geographic locations and discipline distribution features. These are described below.

[0082] Constructing an institutional hierarchy identification system: In the institutional hierarchy identification system, the university level is higher than the institute level, and the institute level is higher than the department level. For example, University (3) > Institute (2) > Department (1). Where University represents university, Institute represents institute, and Department represents department. The numbers in brackets represent their specific levels in the institutional hierarchy identification system. The larger the number, the higher the level represented.

[0083] Build a standardized country code mapping table: This table supports multiple representations of over 200 countries and regions. This table accurately and unambiguously identifies country and region representations, mapping them to their corresponding representations and accurately determining geographic information.

[0084] Constructing a database of institutional location and disciplinary distribution features: As the name suggests, this database can assist in determining an institution's location. While the disciplinary distribution features in this database are not used in institutional similarity mining, they can be used as an extension or in other unspecified scenarios related to academic or disciplinary relationships. By leveraging disciplinary distribution features, accurate author name disambiguation can be achieved.

[0085] Applying country codes for standardized mapping, specifically: performing standardized mapping on the country codes in the institution information based on the country code standardized mapping table; The aforementioned matching of the geographical locations of the two sets of institutional information is specifically determined based on the analysis of the institutional geographical locations and discipline distribution feature database; the aforementioned matching of the institutional levels of the two sets of institutional information is specifically determined based on the analysis of the institutional level identification system.

[0086] (2) Similarity mining based on name.

[0087] In the embodiment of the present application, an algorithm process for deep name similarity calculation is innovatively adopted. Specifically, similarity mining based on name is performed on the target author information in the candidate information set and the author information to be analyzed, including: Perform Unicode normalization on the name information in the target author information and the name information in the author information to be analyzed respectively; Recognize compound surnames based on two sets of name information after Unicode standardization; Based on the surname's exact match weight, surname's fuzzy match weight, full name string (Token) sorting match weight, and initial letter sequence match weight, a multi-level matching result of the two sets of name information is calculated; Cultural background matching is used for the two sets of name information; cultural background matching includes: pinyin conversion and initial consonant matching for Chinese names; Based on the multi-level matching results and the cultural background matching results, the results of name similarity mining between the target author and the author to be analyzed are generated.

[0088] Unicode normalization processing can include NFD decomposition, ASCII conversion, and special string cleaning. NFD (Normalization Form Decomposition) is a character normalization specification in the Unicode standard that decomposes characters into basic graphemes or combining characters. This decomposition is based on Unicode's standard decomposition rules and ensures compatibility between character encodings on different platforms.

[0089] The recognition of compound surnames can detect more than 100 preset compound surname prefixes such as van / von / de / della.

[0090] The embodiment of the present application proposes to perform multi-level matching calculations for name similarity mining, which involves exact surname matching, fuzzy surname matching, full name string (Token) sorting matching, and initial sequence matching. As an example, the weight of the exact surname matching is 0.4, the weight of the fuzzy surname matching is 0.3, the weight of the full name string (Token) sorting matching is 0.2, and the weight of the initial sequence matching is 0.1. In other words, the weights of the above four levels of matching calculations decrease in sequence. Through multi-level matching, the problem of poor accuracy caused by single-level name similarity mining is avoided.

[0091] In addition, cultural background adaptation was performed for the two sets of name information. Combined with relevant knowledge of the cultural background, the difficulty of name similarity mining was reduced, and the accuracy of name matching between the target author and the author to be analyzed was assisted.

[0092] (3) Time-based similarity mining.

[0093] In the embodiment of the present application, when similarity mining is performed on the time dimension, the similarity of the years in which the authors published their works is taken into account, the overlap between their career trajectories and active periods is analyzed, and the similarity of the frequencies of their published works is analyzed. Specifically, in the embodiment of the present application, time-based similarity mining is performed on the target author information in the candidate information set and the author information to be analyzed, including: Based on the publication year in the target author information and the publication year in the author information to be analyzed, the publication year similarity is calculated using the Gaussian decay function; The target author's academic career stage is inferred based on the first publication year in the target author's information, and the academic career stage of the author to be analyzed is inferred based on the first publication year in the author's information. Career trajectory analysis is performed based on the two groups of academic career stages to obtain the career trajectory similarity analysis results; Based on the publication year in the target author information and the publication year in the author information to be analyzed, the overlap of the active publication period of the work is calculated; Calculate the publication frequency similarity between the target author and the author to be analyzed based on the annual publication number in the target author's information and the annual publication number in the author to be analyzed's information; Based on the similarity of publication years, the similarity analysis results of career trajectories, the overlap of active publication periods of works, and the similarity of publication frequencies, the results of temporal similarity mining between the target author and the author to be analyzed are generated.

[0094] In practical applications, the above-mentioned similarity analysis results of publication year, career trajectory similarity, overlap of active publication periods, and publication frequency similarity each have corresponding fusion weights. The similarity calculation results of the above four aspects can be weighted and summed according to the corresponding fusion weights, and the final result obtained is used as the similarity mining result of the target author and the author to be analyzed in the time dimension. In the technical solution of this application, by mining the similarity of multiple aspects in the time dimension, not only the similarity of publication-related explicit features (such as publication year and publication frequency) is analyzed, but also the similarity mining of features with a certain time span and pattern expression, such as the overlap of career trajectory and active publication period, is deeply extended to achieve a deep understanding and analysis of the commonalities of the time dimension, which in turn helps to improve the effect of author name disambiguation.

[0095] (4) Alias-based similarity mining.

[0096] In the embodiment of the present application, similarity mining based on aliases is performed on the target author information in the candidate information set and the author information to be analyzed, including: generating an alias for the name information in the target author information based on the author alias knowledge base, and generating an alias for the name information in the author information to be analyzed based on the author alias knowledge base; Perform extended matching based on the two sets of generated aliases to obtain alias matching results; The confidence of the alias matching results is evaluated based on the usage frequency, time distribution and source credibility, and the alias similarity mining results between the target author and the author to be analyzed are generated based on the evaluated confidence and the alias matching results.

[0097] The established author alias knowledge base can include historical aliases, common variants, misspellings, and more. Aliases can be generated for the name information in the author information being analyzed using an alias generation algorithm. This algorithm can generate possible aliases based on phonetic similarity and spelling conventions. In practical applications, frequency, temporal distribution, and source credibility can be used to assess the confidence of alias matching results.

[0098] Alternatively, after generating an alias and before performing alias matching, a confidence assessment can be performed on the alias based on usage frequency, time distribution, and source credibility. After the alias confidence assessment, aliases with higher confidence scores can be advanced to the alias matching process. The alias matching results can be directly used as the alias similarity mining results.

[0099] Aliases are used to perform extended matching of author information that is compared with each other, which expands the utilization rate of name information in author name disambiguation scenarios, increases the matching dimension, and improves the feasibility of author name disambiguation.

[0100] (5) Similarity mining based on identity identification codes.

[0101] In the embodiment of the present application, similarity mining based on identity identification codes is performed on the target author information in the candidate information set and the author information to be analyzed, including: Perform a full match between the ORCID in the target author information and the ORCID in the author information to be analyzed to obtain a first match score; The Scopus ID in the target author information and the Scopus ID in the author information to be analyzed are fuzzy matched using the edit distance algorithm to obtain a second matching score; Based on the documents jointly published by the target author and the author to be analyzed, an indirect ID association between the target author and the author to be analyzed is established to obtain a third matching score; Based on the source, verification status, and usage frequency of the identity identification code, credibility weights are configured for the first matching score, the second matching score, and the third matching score, respectively. Based on the configured credibility weights and the three matching scores, the results of identity identification code similarity mining between the target author and the author to be analyzed are generated.

[0102] As described above, the embodiment of the present application adopts enhanced ID matching and promotion, which takes into account three levels of ID matching, namely: ORCID matching, Scopus ID matching, and indirect ID association matching.

[0103] When performing an ORCID exact match, if an ORCID exact match is achieved, a first match score of 1.0 can be obtained. Fuzzy matching of Scopus IDs can tolerate 1-2 character differences. Each published work or document has a corresponding document ID. In the embodiment of the present application, an indirect ID association between the target author and the author to be analyzed is established based on the documents jointly published by the target author and the author to be analyzed, and a third match score is obtained. Specifically, the ID of the jointly published document is used to establish an indirect ID association between the target author and the author to be analyzed.

[0104] When mining IDs, we consider not only the ID matching itself but also the credibility of the IDs. As described above, each match score can be assigned a credibility weight based on the ID's source, verification status, and frequency of use. This ensures a higher level of credibility in the final ID similarity mining results, improving the accuracy of author name disambiguation.

[0105] The above is a relatively detailed introduction to the similarity mining process of each of the six dimensions of identity identification code, name, organization, cooperation network, time and alias. In the previous embodiment, it was introduced that the results of the similarity mining of multiple dimensions between the same author information in the candidate information set and the author information to be analyzed are fused to obtain the similarity fusion results of multiple dimensions. The following is an introduction to the example adjustment scheme of the weights of the fusion process in combination with the mining of the above six dimensions. It is mentioned in the embodiment of the present application that a method can be used to adaptively adjust the weights used in the fusion of each dimension based on data quality.

[0106] First, obtain the basic weights corresponding to the six dimensions of identification code, name, organization, collaborative network, time, and alias; the sum of the basic weights corresponding to the six dimensions is 1. As an example, the basic weights corresponding to the six dimensions of identification code, name, organization, collaborative network, time, and alias are: 0.25 (identity code), 0.4 (name), 0.15 (organization), 0.15 (collaborative network), 0.03 (time), and 0.02 (alias).

[0107] If the author information meets the weight adaptive adjustment conditions based on data quality, or the author information meets the weight adaptive adjustment conditions based on scenarios, the basic weights corresponding to the six dimensions will be dynamically adjusted according to the weight dynamic adjustment plan corresponding to the conditions met.

[0108] The aforementioned weight adaptive adjustment conditions based on data quality include: one or more of the first condition, the second condition and the third condition; The first condition is: the credibility of the identity identification code information in the author information is greater than the preset credibility threshold; The second condition is: the name information in the author information is incomplete; The third condition is: the institutional information in the author information is missing.

[0109] The above-mentioned scenario-based weight adaptive adjustment conditions include: one or more of the fourth condition, the fifth condition and the sixth condition; The fourth condition is that the author's information is consistent with the international cooperation scenario; The fifth condition is that the author information is consistent with the interdisciplinary research scenario; The sixth condition is that the author’s information is identified as an emerging scholar.

[0110] In the embodiment of the present application, if any one of the first to sixth conditions is satisfied, it is necessary to dynamically adjust the basic weights corresponding to each of the six dimensions according to the dynamic adjustment scheme of the satisfied condition. The reason for adjusting the basic weights of the six dimensions is that the sum of the basic weights of the six dimensions is 1. If one weight is adjusted up or down, the basic weights of the remaining dimensions must also change to meet the requirement that the sum of the basic weights is 1.

[0111] The weight adjustment plan for the first condition is to increase the weight of the ID card and proportionally reduce the weights of the other dimensions. For example, if the ID card dimension has high credibility, its weight will be increased from the base weight of 0.25 to 0.4, and the base weights of the other dimensions will be proportionally reduced.

[0112] The weight adjustment plan for the second condition is to increase the weight of the collaborative network and proportionally reduce the weights of other dimensions. For example, if the name dimension data is incomplete, the weight of the collaborative network dimension is increased to 0.25, and the weights of other dimensions are proportionally reduced.

[0113] The weight adjustment scheme for the third condition is to increase the sum of the weights for the time and alias dimensions, while proportionally reducing the weights for the other dimensions. When institutional information is missing, the sum of the weights for the time and alias dimensions is increased to 0.1, while the weights for the other dimensions are proportionally reduced. This increases the weight of the similarity mining results for the time and alias dimensions in the fused mining results.

[0114] The weight adjustment plan for the fourth condition is to increase the weight of the institution and proportionally reduce the weights of other dimensions. For example, the weight of the institution can be further increased by 0.05, while the weights of other dimensions can be proportionally reduced.

[0115] The weight adjustment plan for the fifth condition is to increase the weight of the collaborative network and proportionally reduce the weights of other dimensions. For example, the weight of the collaborative network can be further increased by 0.1, while the weights of other dimensions can be proportionally reduced.

[0116] The weight adjustment scheme corresponding to the sixth condition is: increase the weight of time and proportionally reduce the weights of other dimensions. For example, increase the weight of time by 0.05 and proportionally reduce the weights of other dimensions.

[0117] Regarding the adaptive adjustment of the above-mentioned weights, the implementation process of step S103 may be specifically as follows: the results of identity code similarity mining, name similarity mining, organization similarity mining, cooperation network similarity mining, time similarity mining and alias similarity mining between the same author information in the candidate information set and the author information to be analyzed are integrated according to the adjusted weights of the corresponding dimensions.

[0118] This application utilizes an innovative dynamic weight adjustment mechanism, taking into account data quality across all dimensions (e.g., integrity, credibility, consistency, and other factors that reflect data quality) while also fully considering contextual factors such as international collaboration, interdisciplinary research, and emerging scholars. This allows the fusion of similarity mining results across multiple dimensions to be tailored to the context and fully consider data quality, achieving a dynamic and flexible fusion effect.

[0119] For data quality assessments, completeness, credibility, and consistency can all be scored. For example, completeness is scored based on the ratio of the number of available data dimensions to the total number of data dimensions. Credibility can be assessed based on data source, verification status, and historical accuracy. Consistency scores verify the consistency of information across multiple dimensions.

[0120] In addition to the previously described dynamic adjustment and optimization of the weights of each dimension, this embodiment of the present application also proposes the application of a multi-level rule engine during the decision-making stage of integrating the similarity mining results of each dimension. This multi-level rule engine helps clarify whether the fusion results require correction, thereby effectively improving the accuracy of author name disambiguation. The following describes the application of the multi-level rule engine in detail.

[0121] In the embodiment of the present application, the results of identity code similarity mining, name similarity mining, organization similarity mining, collaboration network similarity mining, time similarity mining, and alias similarity mining between the same piece of author information in the candidate information set and the author information to be analyzed are fused according to the weights of the corresponding dimensions to obtain a fusion result of similarities in multiple dimensions between the same piece of author information and the author information to be analyzed, including: The results of identity code similarity mining, name similarity mining, organization similarity mining, collaboration network similarity mining, time similarity mining, and alias similarity mining between the same author information in the candidate information set and the author information to be analyzed are fused according to the weights of the corresponding dimensions to obtain a preliminary fusion result. Determine whether the preliminary fusion result needs to be corrected based on the multi-level rule engine. If it is determined that correction is required, obtain a correction result of the preliminary fusion result based on the multi-level rule engine, and use the correction result as the multi-dimensional similarity fusion result; if it is determined that correction is not required, use the preliminary fusion result as the multi-dimensional similarity fusion result; The multi-level rule engine includes at least two categories of rules: Enforce matching rules, suppress rules, enhance rules, or risk identification rules.

[0122] The mandatory matching rule includes: at least one of a first rule and a second rule; The first rule is: if the IDs match exactly, the lower limit of the similarity fusion results of multiple dimensions is controlled. For example, if the ORCIDs match exactly, the overall similarity (i.e., the similarity fusion results of multiple dimensions) is guaranteed to be at least 0.95.

[0123] The second rule is: if there are three or more collaborators with the same name and organization, the lower limit of the similarity fusion results of multiple dimensions is controlled. For example, under the second rule, the minimum guaranteed overall similarity is 0.90.

[0124] The suppression rule includes at least one of a third rule, a fourth rule, and a fifth rule; The third rule is: if the result of name similarity mining is less than a first preset value, then the upper limit of the similarity fusion results of multiple dimensions is controlled. For example, the first preset value is that if the similarity mining result of the name dimension is less than 0.2, then the overall similarity must not exceed 0.3.

[0125] The fourth rule is: if the institution's country of origin does not match and there is no history of international cooperation, the upper limit of the similarity fusion results of multiple dimensions is controlled. For example, under the fourth rule, the overall similarity must not exceed 0.4.

[0126] The fifth rule is: if the difference in publication years exceeds 15 years and there is no ongoing collaboration, the upper limit of the similarity fusion results in multiple dimensions is controlled. For example, under the fifth rule, the overall similarity must not exceed 0.5.

[0127] The enhanced rule includes at least one of the sixth rule, the seventh rule, and the eighth rule; The sixth rule states: If the similarity mining results for three or more dimensions all exceed the second preset value, the value is adjusted upward based on the preliminary fusion result. For example, if the second preset value is 0.8, under the sixth rule, the overall similarity is adjusted upward by 0.15 based on the preliminary fusion result.

[0128] The seventh rule states: If the network centrality similarity of the collaborative network exceeds the third preset value and the institutions match, the value is adjusted upward based on the initial fusion result. For example, under the seventh rule, the overall similarity is adjusted upward by 0.1 based on the initial fusion result.

[0129] The eighth rule states: If the similarity of publication patterns exceeds the fourth preset value, the value is adjusted upward based on the preliminary fusion results. Publication patterns are behavioral patterns formed by integrating characteristics from multiple time dimensions, including publication frequency, distribution of active publication periods, career trajectory evolution, and annual output rhythm. If the similarity of publication patterns exceeds the fourth preset value, it indicates that the publication patterns are highly similar. As an example, under the eighth rule, the overall similarity is adjusted upward by 0.08 based on the preliminary fusion results.

[0130] Risk identification rules include at least one of the following: If an abnormally high score is detected in the initial fusion result, a secondary verification mechanism will be automatically triggered to check whether there are any data quality issues and identify whether there is any malicious matching behavior; Based on detected data quality issues or identified malicious matching behavior, perform at least one of the following actions: Lower the confidence of the match, mark it for manual review, or reject the match.

[0131] As an example, in the risk identification rules, anomalies are detected by setting an upper limit on the similarity threshold (such as above 0.95): when the initial fusion result shows an abnormally high score, the system will automatically trigger a secondary verification mechanism to check whether there are data entry errors (such as duplicate records, format abnormalities), identity code conflicts (the same ORCID is used by multiple authors), institutional information tampering and other data quality issues, and identify malicious matching behaviors (such as intentional falsification of cooperative relationships, batch false identity associations). Once such risks are detected, the system will reduce the confidence of the match, mark it for manual review, or directly reject the match to ensure the reliability and security of the disambiguation results.

[0132] As can be seen from the above description, this embodiment of the application utilizes risk identification rules to detect high-similarity anomalies and, in conjunction with manual review, confirm them, thereby preventing abnormally high scores from interfering with the accuracy of author name disambiguation. Furthermore, by identifying data quality issues and possible malicious matching attempts, the security and reliability of the entire author name disambiguation process are further ensured.

[0133] When integrating similarity mining results from multiple dimensions, we also propose the use of parallel best match search. For example, we implement batch parallel processing: candidate information sets are computed in parallel in batches of 15. We also employ a thread pool optimization strategy: dynamically adjusting the number of threads (2-6) to avoid resource contention. We also employ memory-friendly batch processing: controlling memory usage to prevent out-of-memory (OOM) issues caused by large datasets.

[0134] Furthermore, we also perform intelligent result mapping and verification for similarity fusion results across multiple dimensions. Specifically, this includes: unified entity ID mapping: mapping the consistency of the alias table author_id to the entity table data_id; result confidence assessment: based on similarity distribution, historical success rate, and a comprehensive data quality score; and abnormal result detection: identifying obviously unreasonable matching results and triggering manual review.

[0135] The embodiment of the present application can use the intelligent two-level adaptive cache introduced above to initialize the system. In addition, the embodiment of the present application also provides a three-level cache optimization system. The cache architecture design includes L1 cache (memory LRU), L2 cache (SQLite persistence) and L3 cache (distributed cache). Among them, L1 cache capacity: 3,000 items, adaptively expanded to 10,000 items; TTL: dynamically adjusted based on access frequency (1-24 hours); hit rate monitoring: target above 95%. L2 cache capacity: 1 million items, regularly clearing expired data; index optimization: composite index, query time <5ms; concurrency control: WAL mode, supporting multiple reads and single writes. Redis cluster is configured in the L3 cache to support TB-level data storage; consistent hashing supports horizontal expansion; data sharding to avoid hot spot problems.

[0136] In addition, the embodiments of the present application propose a distributed computing framework, a task allocation strategy based on the size of the candidate information set and dynamic sharding of computational complexity; a polling + weighted strategy, a load balancing algorithm that allocates tasks according to node performance; and a fault-tolerant recovery mechanism that automatically retry failed tasks and supports breakpoint resumption.

[0137] The present application also optimizes the database connection pool. It proposes thread-local storage and independent connections for each thread to avoid lock contention. It also proposes a connection reuse strategy, keeping idle connections alive and dynamically expanding capacity during busy periods. It also proposes connection health checks, regular connection status checks, and automatic reconnection.

[0138] In addition, the present embodiment also proposes performance monitoring and tuning measures. These measures analyze real-time performance indicators such as QPS, latency distribution, error rate, and resource utilization, and employ an automatic tuning mechanism that automatically adjusts parameters based on load. Furthermore, an early warning system is implemented to provide timely alerts and automatic recovery when system performance anomalies occur.

[0139] As can be seen from the above, the embodiments of this application propose a high-performance optimization architecture with a distributed parallel computing framework, a three-level cache optimization system, database connection pool optimization, and coexistence of performance monitoring and tuning. The use of this system architecture enables the author name disambiguation method proposed in the technical solution of this application to be calculated in a more efficient manner, optimizes the query and result storage patterns and strategies, and the combination of a stable and reliable system and real-time monitoring and alarm mechanisms enhances users' sensitivity to system anomalies.

[0140] Figure 5 This is a diagram of the implementation architecture of an author name disambiguation method provided in an embodiment of the present application. Figure 5 As shown, the implementation architecture involves five stages: Phase 1: Intelligent initialization and network construction.

[0141] Phase 2: Construction and filtering of intelligent candidate information sets.

[0142] The third stage: innovative multi-dimensional similarity calculation.

[0143] Stage 4: Adaptive decision fusion mechanism.

[0144] Phase 5: High-performance optimized architecture.

[0145] In the first phase, it involves the initialization of a two-level cache system, the construction of a multi-level cooperation network, the construction of an institutional hierarchy identification system, a country code standardization mapping table, and an institutional geographic location and discipline distribution feature library. In addition, a clustering algorithm for institutional aliases is implemented.

[0146] The second stage involves the input parsing of multi-dimensional author information, hierarchical filtering strategies, optimization of intelligent query methods, etc.

[0147] The third stage involves enhanced identity code matching algorithms, deep name similarity calculations, intelligent organization matching algorithms, in-depth analysis of the topological structure of collaborative networks, intelligent analysis of the time dimension, and intelligent extended matching of aliases.

[0148] In the fourth stage, it involves the fusion mechanism of innovative dynamic weight optimization methods, multi-level rule engine, parallel optimal matching search, and mapping and verification of matching results.

[0149] The fifth stage involves distributed parallel computing framework, three-level cache optimization system, database connection pool optimization, and performance monitoring and tuning.

[0150] Based on the analysis and introduction above, the existing technologies have the following problems: (1) Problem of insufficient depth of network analysis: Although existing technologies take collaborative networks into consideration, they only stay at the level of simply counting co-collaborators and fail to explore the network's topological structural characteristics (such as degree centrality, clustering coefficient, shortest path) and dynamic evolution patterns, resulting in limited ability to identify complex academic relationships.

[0151] (2) The problem of static similarity fusion mechanism: Existing methods use fixed weights for multi-dimensional fusion, which cannot be dynamically adjusted according to data quality differences, dimension credibility and specific application scenarios. The effect is significantly reduced when processing incomplete data or special scenarios.

[0152] (3) The original problem of computational efficiency optimization strategy: Existing technologies lack candidate information set pre-filtering algorithms and hierarchical cache optimization mechanisms. When faced with millions of author data, the computational complexity grows exponentially and cannot meet real-time application requirements.

[0153] (4) Problems with rough name standardization: Existing methods do not adequately address details such as compound surname recognition, Unicode character standardization, and multilingual name format unification, resulting in a significant drop in accuracy when processing international academic data.

[0154] (5) Problems with simplifying the institution matching algorithm: Existing technologies have not established an institutional hierarchy identification system and standardized mapping of country codes, and are unable to effectively handle the diverse expressions and hierarchical relationships of institutional names.

[0155] This application proposes innovative technical approaches to address the above issues: (1) Deep network structure analysis technology: Construct a multi-level academic cooperation network model, calculate the topological characteristics of nodes such as degree centrality, clustering coefficient, betweenness centrality, etc. through the graph neural network algorithm, and realize deep similarity analysis based on network structure.

[0156] (2) Adaptive multi-dimensional fusion algorithm: Design a data quality assessment model and a dynamic weight adjustment mechanism to adjust the fusion weight in real time based on the integrity, credibility and historical matching success rate of data in each dimension to improve the overall matching accuracy.

[0157] (3) Hierarchical intelligent cache optimization system: Create a three-level cache architecture of in-memory LRU cache + SQLite persistent cache + distributed cache, combined with query prediction algorithm and hot data identification mechanism to achieve millisecond-level response.

[0158] (4) Multilingual intelligent standardization engine: Establish a name standardization system with compound surname recognition algorithms, Unicode standardization processing and unified multilingual formats, and support mixed processing of Chinese, English and other languages.

[0159] The core technical highlights of this application's technical solution are as follows: 1. A pioneering algorithm for deep network topology analysis: Unlike existing technologies that simply count collaborators, this algorithm implements network structure similarity calculation based on graph neural networks, including advanced network features such as degree centrality, clustering coefficient, and structural hole analysis, filling a technological gap in academic deep network analysis. 2. Groundbreaking Adaptive Weight Fusion Mechanism: An innovative data quality assessment model and dynamic weight adjustment algorithm optimize fusion weights in real time based on the integrity, credibility, and historical success rate of data in each dimension, outperforming fixed weight methods. 3. Original three-level intelligent cache architecture: The unique three-level cache system of in-memory LRU + SQLite persistence + distributed cache, combined with query prediction and hotspot identification, achieves a 99.5% cache hit rate and millisecond-level response speed; 4. The first multilingual intelligent standardization engine: Establishes a complete system for compound surname recognition, Unicode standardization, and multilingual format unification, supporting mixed processing of Chinese and English and institutional standardization in over 200 countries / regions; 5. Innovative multi-level special rule engine: Designed with four types of intelligent decision-making rules, including forced matching, suppression rules, enhancement rules, and risk identification, it enables accurate judgment and anomaly identification in complex scenarios. 6. Unique distributed parallel computing framework: This framework implements a complete distributed architecture with task sharding, load balancing, and fault-tolerance recovery. It supports real-time processing of millions of author data, significantly improving computing efficiency compared to traditional methods. 7. Original Network Evolution Time Series Analysis: This is the first time that time series network analysis has been introduced into author disambiguation. Author identities are identified through the temporal evolution patterns of collaborative relationships, solving the problem of traditional methods ignoring temporal dynamics.

[0160] Based on the method introduced in the above embodiment, this application also proposes an author name disambiguation device. Figure 6 This is a schematic diagram of the structure of the author name disambiguation device, such as Figure 6 The device shown comprises: The information acquisition module 61 is used to acquire a piece of author information to be analyzed and a candidate information set; the candidate information set includes multiple pieces of author information; the author information to be analyzed includes a list of collaborators; A multi-dimensional similarity mining module 62 is configured to perform multi-dimensional similarity mining on the plurality of author information and the author information to be analyzed, respectively; wherein a collaboration network is one of the plurality of dimensions, wherein nodes in the collaboration network represent authors, and lines between nodes represent collaboration relationships; and similarity mining of the collaboration network includes at least one of the following: direct collaboration relationship similarity mining based on the collaboration network, indirect network relationship similarity mining based on the collaboration network, network structure topology feature similarity mining based on the collaboration network, or temporal network evolution similarity mining based on the collaboration network; A multi-dimensional similarity fusion module 63 is used to fuse the results of multi-dimensional similarity mining between the same author information in the candidate information set and the author information to be analyzed to obtain a multi-dimensional similarity fusion result; The matching module 64 is configured to obtain a matching result between the author to be analyzed and the authors in the candidate information set based on a fusion result of the multiple dimensions of similarity between the multiple pieces of author information and the author information to be analyzed.

[0161] In an optional implementation, the multi-dimensional similarity mining module 62 is specifically configured to: Analyzing the network structure topology characteristics of the first collaboration network to obtain the local clustering coefficient, structural hole location, and network centrality of the first collaboration network; the first collaboration network is a collaboration network about the author to be analyzed; Analyze the network structure topology characteristics of the second collaboration network to obtain the local clustering coefficient, structural hole location, and network centrality of the second collaboration network; the second collaboration network is a collaboration network about the target author; Calculating the local clustering similarity between the target author and the author to be analyzed based on the local clustering coefficient of the first collaboration network and the local clustering coefficient of the second collaboration network; Based on the structural hole positions of the first collaboration network and the structural hole positions of the second collaboration network, calculating the structural hole position similarity between the target author and the author to be analyzed; Based on the network centrality of the first collaboration network and the network centrality of the second collaboration network, calculating the network centrality similarity between the target author and the author to be analyzed; Based on the local clustering similarity, the structural hole position similarity and the network centrality similarity, a similarity mining result of the network structure topology characteristics between the target author and the author to be analyzed is calculated.

[0162] In an optional implementation, the multi-dimensional similarity mining module 62 is specifically configured to: Performing a time window analysis based on a first collaboration network to obtain a first time window analysis result and identifying a dynamic change pattern of the first collaboration network; the first collaboration network is a collaboration network of the author to be analyzed; Performing a time window analysis based on a second collaboration network to obtain a second time window analysis result, and identifying a dynamic change pattern of the second collaboration network; the second collaboration network is a collaboration network about the target author; Based on the first time window analysis result and the second time window analysis result, obtaining the network time window similarity between the target author and the author to be analyzed; Obtaining a similarity between the network dynamic change patterns of the target author and the author to be analyzed based on the dynamic change pattern of the first collaboration network and the dynamic change pattern of the second collaboration network; According to the network time window similarity and the network dynamic change pattern similarity, a temporal network evolution similarity mining result between the target author and the author to be analyzed is calculated.

[0163] In an optional implementation, the multi-dimensional similarity mining module 62 is specifically configured to: Calculating a first collaboration intensity and a first collaboration density between the author to be analyzed and each author in the collaborator list based on a first collaboration network; wherein the first collaboration network is a collaboration network of the author to be analyzed; Calculating a second collaboration intensity and a second collaboration density between the target author and each author in the collaborator list based on a second collaboration network; wherein the second collaboration network is a collaboration network about the target author; Obtaining a similarity in cooperation intensity between the target author and the author to be analyzed based on the first cooperation intensity and the second cooperation intensity of the author to be analyzed and the target author with the same author in the collaborator list; Obtaining a collaboration density similarity between the target author and the author to be analyzed based on a first collaboration density and a second collaboration density between the author to be analyzed and the target author and the same author in the collaborator list; According to the cooperation intensity similarity and the cooperation density similarity, a direct cooperation relationship similarity mining result between the target author and the author to be analyzed is calculated.

[0164] In an optional implementation, the multi-dimensional similarity mining module 62 is specifically configured to: Identifying second-degree connected nodes and third-degree connected nodes of the author to be analyzed from a first collaboration network; the first collaboration network is a collaboration network about the author to be analyzed; Identifying second-degree connected nodes and third-degree connected nodes of the target author from a second collaboration network; the second collaboration network is a collaboration network about the target author; Obtaining a second-degree connection similarity between the target author and the author to be analyzed based on the second-degree connection nodes of the author to be analyzed and the second-degree connection nodes of the target author; Obtaining a three-degree connection similarity between the target author and the author to be analyzed based on the three-degree connection nodes of the author to be analyzed and the three-degree connection nodes of the target author; According to the second-degree connection similarity and the third-degree connection similarity, an indirect network relationship similarity mining result between the target author and the author to be analyzed is calculated.

[0165] In an optional implementation, similarity mining of the collaboration network includes: similarity mining of direct collaboration relationships based on the collaboration network, similarity mining of indirect network relationships based on the collaboration network, similarity mining of network structure topology features based on the collaboration network, and similarity mining of temporal network evolution based on the collaboration network; The multi-dimensional similarity mining module 62 is also used to obtain the results of cooperative network similarity mining based on the network structure topology feature similarity mining results, the temporal network evolution similarity mining results, the direct cooperative relationship similarity mining results and the indirect network relationship similarity mining results.

[0166] In an optional implementation, the multiple dimensions further include: identification code, name, organization, time, and alias; The multi-dimensional similarity fusion module 63 is specifically used to: The results of identity code similarity mining, name similarity mining, organization similarity mining, collaboration network similarity mining, time similarity mining and alias similarity mining between the same author information in the candidate information set and the author information to be analyzed are fused according to the weights of the corresponding dimensions to obtain multiple dimensional similarity fusion results between the same author information and the author information to be analyzed.

[0167] In an optional implementation, the author name disambiguation device further includes: The basic weight acquisition module is used to obtain the basic weights corresponding to the six dimensions of identity identification code, name, organization, cooperation network, time and alias; the sum of the basic weights corresponding to the six dimensions is 1; The weight adaptive adjustment module is used to dynamically adjust the basic weights corresponding to the six dimensions according to the weight dynamic adjustment scheme corresponding to the conditions met, if the author information meets the weight adaptive adjustment conditions based on data quality, or the author information meets the weight adaptive adjustment conditions based on scenarios; The multi-dimensional similarity fusion module 63 is specifically used to: fuse the results of identity code similarity mining, name similarity mining, organization similarity mining, cooperation network similarity mining, time similarity mining and alias similarity mining between the same author information in the candidate information set and the author information to be analyzed according to the adjusted weights of the corresponding dimensions.

[0168] In an optional implementation, the multi-dimensional similarity fusion module 63 is specifically configured to: The results of identity code similarity mining, name similarity mining, organization similarity mining, collaboration network similarity mining, time similarity mining, and alias similarity mining between the same author information in the candidate information set and the author information to be analyzed are fused according to the weights of the corresponding dimensions to obtain a preliminary fusion result; Determining whether the preliminary fusion result needs to be revised based on the multi-level rule engine; if it is determined that revision is required, obtaining a revised result of the preliminary fusion result based on the multi-level rule engine, and using the revised result as the multi-dimensional similarity fusion result; if it is determined that revision is not required, using the preliminary fusion result as the multi-dimensional similarity fusion result; The multi-level rule engine includes at least two categories of rules: Enforce matching rules, suppress rules, enhance rules, or risk identification rules.

[0169] In an optional implementation, the multi-dimensional similarity mining module 62 is configured to: Performing standardization preprocessing on the institution information in the target author information and the institution information in the author information to be analyzed respectively; the standardization preprocessing includes: removing the prefix of the institution type, extracting the core keywords of the institution, and applying the country code for standardization mapping; The Jaccard similarity and fuzzy matching similarity were calculated for the keywords of the two sets of institutional information after standardization preprocessing. Calculating preliminary organization similarity based on the Jaccard similarity and the fuzzy matching similarity; Based on the matching of the geographical locations and the matching of the institutional levels of the two sets of institutional information, the preliminary similarity of the institutions is optimized, and the optimized similarity result is used as the result of the institutional similarity mining between the target author and the author to be analyzed.

[0170] In an optional implementation, the author name disambiguation apparatus further includes a building module for: Establishing an institutional hierarchy identification system in which universities are ranked higher than research institutes, which are ranked higher than departments; Constructing a standardized mapping table for country codes; the standardized mapping table for country codes supports multiple expressions for more than 200 countries or regions; Construct a database of institutional geographical location and discipline distribution characteristics; The application of the country code for standardized mapping is specifically: performing standardized mapping on the country code in the institution information based on the country code standardized mapping table; The matching of the geographical locations of the two sets of institution information is determined based on analysis of the geographical locations of the institutions and the subject distribution feature database; The matching of the institutional hierarchies of the two sets of institutional information is determined based on analysis of the institutional hierarchy identification system.

[0171] In an optional implementation, the multi-dimensional similarity mining module 62 is configured to: Performing Unicode normalization processing on the name information in the target author information and the name information in the author information to be analyzed respectively; Recognize compound surnames based on two sets of name information after Unicode standardization; Based on the surname's exact match weight, surname's fuzzy match weight, full name token sorting match weight, and initial sequence match weight, a multi-level matching result for the two sets of name information is calculated. Cultural background matching is used for the two sets of name information; cultural background matching includes: pinyin conversion and initial consonant matching for Chinese names; Based on the multi-level matching results and the cultural background matching results, a result of name similarity mining between the target author and the author to be analyzed is generated.

[0172] In an optional implementation, the multi-dimensional similarity mining module 62 is configured to: Calculate the publication year similarity using a Gaussian decay function based on the publication year in the target author information and the publication year in the author information to be analyzed; Inferring the academic career stage of the target author based on the first publication year in the target author information, and inferring the academic career stage of the author to be analyzed based on the first publication year in the author information to be analyzed, performing career trajectory analysis based on the two groups of academic career stages, and obtaining career trajectory similarity analysis results; Calculate the overlap of the active publication period of the works based on the publication year in the target author information and the publication year in the author information to be analyzed; Calculating the publication frequency similarity between the target author and the author to be analyzed based on the annual publication number in the target author information and the annual publication number in the author information to be analyzed; According to the similarity of the publication years, the similarity analysis results of the career trajectories, the overlap of the active publication periods of the works and the similarity of the publication frequencies, the results of time similarity mining between the target author and the author to be analyzed are generated.

[0173] In an optional implementation, the multi-dimensional similarity mining module 62 is configured to: generating an alias for the name information in the target author information based on an author alias knowledge base, and generating an alias for the name information in the author information to be analyzed based on the author alias knowledge base; Perform extended matching based on the two sets of generated aliases to obtain alias matching results; A confidence evaluation is performed on the alias matching result based on usage frequency, time distribution and source credibility, and an alias similarity mining result between the target author and the author to be analyzed is generated based on the evaluated confidence and the alias matching result.

[0174] In an optional implementation, the multi-dimensional similarity mining module 62 is configured to: Performing a complete match between the ORCID in the target author information and the ORCID in the author information to be analyzed to obtain a first match score; Performing fuzzy matching on the Scopus ID in the target author information and the Scopus ID in the author information to be analyzed using an edit distance algorithm to obtain a second matching score; Based on the documents jointly published by the target author and the author to be analyzed, an indirect ID association is established between the target author and the author to be analyzed to obtain a third matching score; Based on the source, verification status, and usage frequency of the identity identification code, credibility weights are respectively configured for the first matching score, the second matching score, and the third matching score, and based on the configured credibility weights and the three matching scores, the results of identity identification code similarity mining between the target author and the author to be analyzed are generated.

[0175] In an optional implementation, the information acquisition module 61 is specifically configured to: In the author information database, author information is matched based on the last name and the initial of the author information to be analyzed to obtain a first matching information set that meets the matching condition; In the first matching information set, author information is matched based on the institution keyword and year range of the author information to be analyzed to obtain a second matching information set that meets the matching conditions; In the second matching information set, author information matching is performed based on the identity identification code of the author information to be analyzed, and a third matching information set that meets the matching conditions is obtained as the candidate information set.

[0176] Based on the author name disambiguation method and apparatus provided in the aforementioned embodiments, the present application further provides an author name disambiguation device, the device comprising: a processor and a memory communicatively connected to each other; The memory stores a computer program; The processor is configured to run the computer program to implement the author name disambiguation method of any implementation manner described in the method embodiment.

[0177] In addition, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored. When the computer program is processed, it implements the steps of the author name disambiguation method as described in the method embodiment.

[0178] It should be noted that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device and equipment embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments. The device and equipment embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components indicated as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without expending creative work.

[0179] The above is merely one specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A method for disambiguating author names, characterized in that: include: Get an author information to be analyzed and a candidate information set; The candidate information set includes multiple pieces of author information; The author information to be analyzed includes a list of collaborators; Performing multiple-dimensional similarity mining on the multiple pieces of author information and the author information to be analyzed, respectively; wherein the collaboration network is one of the multiple dimensions, wherein the authors are represented by nodes in the collaboration network, and the collaboration relationships are represented by lines between the nodes; the similarity mining of the collaboration network includes at least one of the following: direct collaboration relationship similarity mining based on the collaboration network, indirect network relationship similarity mining based on the collaboration network, network structure topology feature similarity mining based on the collaboration network, or temporal network evolution similarity mining based on the collaboration network; fusing the results of multi-dimensional similarity mining between the same author information in the candidate information set and the author information to be analyzed to obtain a multi-dimensional similarity fusion result; According to the fusion results of the multiple dimensions of similarity between the multiple pieces of author information and the author information to be analyzed, a matching result between the author to be analyzed and the authors in the candidate information set is obtained.

2. The method according to claim 1, characterized in that Performing similarity mining of network structure topology features based on a collaborative network on the target author information in the candidate information set and the author information to be analyzed, including: Analyzing the network structure topology characteristics of the first collaboration network to obtain the local clustering coefficient, structural hole location, and network centrality of the first collaboration network; the first collaboration network is a collaboration network about the author to be analyzed; Analyze the network structure topology characteristics of the second collaboration network to obtain the local clustering coefficient, structural hole location, and network centrality of the second collaboration network; the second collaboration network is a collaboration network about the target author; Calculating the local clustering similarity between the target author and the author to be analyzed based on the local clustering coefficient of the first collaboration network and the local clustering coefficient of the second collaboration network; Based on the structural hole positions of the first collaboration network and the structural hole positions of the second collaboration network, calculating the structural hole position similarity between the target author and the author to be analyzed; Based on the network centrality of the first collaboration network and the network centrality of the second collaboration network, calculating the network centrality similarity between the target author and the author to be analyzed; Based on the local clustering similarity, the structural hole position similarity and the network centrality similarity, a similarity mining result of the network structure topology characteristics between the target author and the author to be analyzed is calculated.

3. The method according to claim 1, characterized in that Performing time series network evolution similarity mining based on a collaboration network on the target author information in the candidate information set and the author information to be analyzed, including: Performing a time window analysis based on a first collaboration network to obtain a first time window analysis result and identifying a dynamic change pattern of the first collaboration network; the first collaboration network is a collaboration network of the author to be analyzed; Performing a time window analysis based on a second collaboration network to obtain a second time window analysis result, and identifying a dynamic change pattern of the second collaboration network; the second collaboration network is a collaboration network about the target author; Based on the first time window analysis result and the second time window analysis result, obtaining the network time window similarity between the target author and the author to be analyzed; Obtaining a similarity between the network dynamic change patterns of the target author and the author to be analyzed based on the dynamic change pattern of the first collaboration network and the dynamic change pattern of the second collaboration network; According to the network time window similarity and the network dynamic change pattern similarity, a temporal network evolution similarity mining result between the target author and the author to be analyzed is calculated.

4. The method according to claim 1, wherein Performing direct cooperation relationship similarity mining based on the cooperation network on the target author information in the candidate information set and the author information to be analyzed, including: Calculating a first collaboration intensity and a first collaboration density between the author to be analyzed and each author in the collaborator list based on a first collaboration network; wherein the first collaboration network is a collaboration network of the author to be analyzed; Calculating a second collaboration intensity and a second collaboration density between the target author and each author in the collaborator list based on a second collaboration network; wherein the second collaboration network is a collaboration network about the target author; Obtaining a similarity in cooperation intensity between the target author and the author to be analyzed based on the first cooperation intensity and the second cooperation intensity of the author to be analyzed and the target author with the same author in the collaborator list; Obtaining a collaboration density similarity between the target author and the author to be analyzed based on a first collaboration density and a second collaboration density between the author to be analyzed and the target author and the same author in the collaborator list; According to the cooperation intensity similarity and the cooperation density similarity, a direct cooperation relationship similarity mining result between the target author and the author to be analyzed is calculated.

5. The method according to claim 1, wherein Performing indirect network relationship similarity mining based on a collaborative network on the target author information in the candidate information set and the author information to be analyzed, including: Identifying second-degree connected nodes and third-degree connected nodes of the author to be analyzed from a first collaboration network; the first collaboration network is a collaboration network about the author to be analyzed; Identifying second-degree connected nodes and third-degree connected nodes of the target author from a second collaboration network; the second collaboration network is a collaboration network about the target author; Obtaining a second-degree connection similarity between the target author and the author to be analyzed based on the second-degree connection nodes of the author to be analyzed and the second-degree connection nodes of the target author; Obtaining a three-degree connection similarity between the target author and the author to be analyzed based on the three-degree connection nodes of the author to be analyzed and the three-degree connection nodes of the target author; According to the second-degree connection similarity and the third-degree connection similarity, an indirect network relationship similarity mining result between the target author and the author to be analyzed is calculated.

6. The method according to claim 1, characterized in that Similarity mining of collaboration networks includes: similarity mining of direct collaboration relationships based on collaboration networks, similarity mining of indirect network relationships based on collaboration networks, similarity mining of network structure topology features based on collaboration networks, and similarity mining of temporal network evolution based on collaboration networks. The method further includes: fusing the similarity mining results of the cooperative network based on the similarity mining results of the network structure topology characteristics, the similarity mining results of the temporal network evolution, the similarity mining results of the direct cooperative relationship and the similarity mining results of the indirect network relationship to obtain the similarity mining results of the cooperative network.

7. The method according to claim 1, characterized in that Multiple dimensions also include: identification code, name, organization, time, and alias; The fusion processing of the results of the multi-dimensional similarity mining between the same author information in the candidate information set and the author information to be analyzed to obtain the multi-dimensional similarity fusion result includes: The results of identity code similarity mining, name similarity mining, organization similarity mining, collaboration network similarity mining, time similarity mining and alias similarity mining between the same author information in the candidate information set and the author information to be analyzed are fused according to the weights of the corresponding dimensions to obtain multiple dimensional similarity fusion results between the same author information and the author information to be analyzed.

8. The method according to claim 7, characterized in that The method further comprises: Obtain the basic weights corresponding to the six dimensions of identity identification code, name, organization, cooperation network, time and alias; the sum of the basic weights corresponding to the six dimensions is 1; If the author information meets the weight adaptive adjustment conditions based on data quality, or the author information meets the weight adaptive adjustment conditions based on scenarios, the basic weights corresponding to the six dimensions will be dynamically adjusted according to the weight dynamic adjustment scheme corresponding to the conditions met; The results of the identity code similarity mining, name similarity mining, organization similarity mining, collaboration network similarity mining, time similarity mining, and alias similarity mining between the same author information in the candidate information set and the author information to be analyzed are integrated according to the weights of the corresponding dimensions, specifically including: The results of identity code similarity mining, name similarity mining, organization similarity mining, collaboration network similarity mining, time similarity mining and alias similarity mining between the same author information in the candidate information set and the author information to be analyzed are fused according to the adjusted weights of the corresponding dimensions.

9. The method according to claim 8, characterized in that The weight adaptive adjustment condition based on data quality includes: one or more conditions among the first condition, the second condition and the third condition; The first condition is: the credibility of the identity identification code information in the author information is greater than a preset credibility threshold; The second condition is: the name information in the author information is incomplete; The third condition is: the institutional information in the author information is missing; The weight adjustment scheme corresponding to the first condition is: the weight corresponding to the identity identification code is increased, and the weights corresponding to other dimensions are proportionally reduced; The weight adjustment plan corresponding to the second condition is: increase the weight corresponding to the cooperation network and reduce the weights corresponding to other dimensions proportionally; The weight adjustment scheme corresponding to the third condition is: the sum of the weights of the time and alias dimensions is increased, and the weights corresponding to other dimensions are reduced proportionally.

10. The method according to claim 8, characterized in that The scenario-based weight adaptive adjustment condition includes: one or more of the fourth condition, the fifth condition and the sixth condition; The fourth condition is that the author's information meets the requirements of international cooperation; The fifth condition is that the author information conforms to the interdisciplinary research scenario; The sixth condition is: the author's information is identified as an emerging scholar; The weight adjustment plan corresponding to the fourth condition is: increase the weight corresponding to the institution and reduce the weights corresponding to other dimensions proportionally; The weight adjustment solution corresponding to the fifth condition is: increase the weight corresponding to the cooperation network and reduce the weights corresponding to other dimensions proportionally; The weight adjustment scheme corresponding to the sixth condition is: increase the weight of time and reduce the weights corresponding to other dimensions proportionally.

11. The method according to claim 7, characterized in that The identification code similarity mining results, name similarity mining results, organization similarity mining results, collaboration network similarity mining results, time similarity mining results, and alias similarity mining results between the same piece of author information in the candidate information set and the author information to be analyzed are fused according to the weights of the corresponding dimensions to obtain a fusion result of similarities in multiple dimensions between the same piece of author information and the author information to be analyzed, including: The results of identity code similarity mining, name similarity mining, organization similarity mining, collaboration network similarity mining, time similarity mining, and alias similarity mining between the same author information in the candidate information set and the author information to be analyzed are fused according to the weights of the corresponding dimensions to obtain a preliminary fusion result; Determining whether the preliminary fusion result needs to be revised based on the multi-level rule engine; if it is determined that revision is required, obtaining a revised result of the preliminary fusion result based on the multi-level rule engine, and using the revised result as the multi-dimensional similarity fusion result; if it is determined that revision is not required, using the preliminary fusion result as the multi-dimensional similarity fusion result; The multi-level rule engine includes at least two categories of rules: Enforce matching rules, suppress rules, enhance rules, or risk identification rules.

12. The method according to claim 11, characterized in that The mandatory matching rule includes: at least one of a first rule and a second rule; The first rule is: if the identity identification code is completely matched, the lower limit of the similarity fusion result of multiple dimensions is controlled; The second rule is: if the names and organizations are the same and there are more than three co-collaborators, the lower limit of the similarity fusion results of multiple dimensions is controlled.

13. The method according to claim 11, characterized in that The suppression rule includes at least one of a third rule, a fourth rule, and a fifth rule; The third rule is: if the result of name similarity mining is less than a first preset value, then the upper limit of the similarity fusion result of multiple dimensions is controlled; The fourth rule is: if the institution country does not match and there is no history of international cooperation, then the upper limit of the similarity fusion results of multiple dimensions is controlled; The fifth rule is: if the difference in publication years exceeds 15 years and there is no ongoing collaboration, the upper limit of the similarity fusion results of multiple dimensions is controlled.

14. The method according to claim 11, characterized in that The enhanced rule includes at least one of the sixth rule, the seventh rule and the eighth rule; The sixth rule is: if the similarity mining results of three or more dimensions all exceed the second preset value, then the value is increased based on the preliminary fusion result; The seventh rule is: if the network centrality similarity of the collaboration network exceeds a third preset value and the organizations match, then the value is adjusted upward based on the preliminary fusion result; The eighth rule is: if the similarity of the publication patterns exceeds the fourth preset value, the value will be adjusted upward based on the preliminary fusion result; the publication pattern is a behavioral pattern formed by comprehensively considering the characteristics of multiple time dimensions including publication frequency, distribution of active publication periods, evolution of career trajectories, and annual output rhythm.

15. The method according to claim 11, characterized in that The risk identification rules include at least one of the following: If an abnormally high score is detected in the preliminary fusion result, a secondary verification mechanism is automatically triggered to check whether there are any data quality issues and identify whether there is any malicious matching behavior; Based on detected data quality issues or identified malicious matching behavior, perform at least one of the following actions: Lower the confidence of the match, mark it for manual review, or reject the match.

16. The method according to claim 7, characterized in that Performing institution-based similarity mining on the target author information in the candidate information set and the author information to be analyzed, including: Performing standardization preprocessing on the institution information in the target author information and the institution information in the author information to be analyzed respectively; the standardization preprocessing includes: removing the prefix of the institution type, extracting the core keywords of the institution, and applying the country code for standardization mapping; The Jaccard similarity and fuzzy matching similarity were calculated for the keywords of the two sets of institutional information after standardization preprocessing. Calculating preliminary organization similarity based on the Jaccard similarity and the fuzzy matching similarity; Based on the matching of the geographical locations and the matching of the institutional levels of the two sets of institutional information, the preliminary similarity of the institutions is optimized, and the optimized similarity result is used as the result of the institutional similarity mining between the target author and the author to be analyzed.

17. The method according to claim 16, characterized in that The method further comprises: Establishing an institutional hierarchy identification system in which universities are ranked higher than research institutes, which are ranked higher than departments; Constructing a standardized mapping table for country codes; the standardized mapping table for country codes supports multiple expressions for more than 200 countries or regions; Construct a database of institutional geographical location and discipline distribution characteristics; The application of the country code for standardized mapping is specifically: performing standardized mapping on the country code in the institution information based on the country code standardized mapping table; The matching of the geographical locations of the two sets of institution information is determined based on analysis of the geographical locations of the institutions and the subject distribution feature database; The matching of the institutional hierarchies of the two sets of institutional information is determined based on analysis of the institutional hierarchy identification system.

18. The method according to claim 7, characterized in that Performing name-based similarity mining on the target author information in the candidate information set and the author information to be analyzed, including: Performing Unicode normalization processing on the name information in the target author information and the name information in the author information to be analyzed respectively; Recognize compound surnames based on two sets of name information after Unicode standardization; Based on the surname's exact match weight, surname's fuzzy match weight, full name token sorting match weight, and initial sequence match weight, a multi-level matching result for the two sets of name information is calculated. Cultural background matching is used for the two sets of name information; cultural background matching includes: pinyin conversion and initial consonant matching for Chinese names; Based on the multi-level matching results and the cultural background matching results, a result of name similarity mining between the target author and the author to be analyzed is generated.

19. The method according to claim 7, characterized in that Performing time-based similarity mining on the target author information in the candidate information set and the author information to be analyzed includes: Calculate the publication year similarity using a Gaussian decay function based on the publication year in the target author information and the publication year in the author information to be analyzed; Inferring the academic career stage of the target author based on the first publication year in the target author information, and inferring the academic career stage of the author to be analyzed based on the first publication year in the author information to be analyzed, performing career trajectory analysis based on the two groups of academic career stages, and obtaining career trajectory similarity analysis results; Calculate the overlap of the active publication period of the works based on the publication year in the target author information and the publication year in the author information to be analyzed; Calculate the publication frequency similarity between the target author and the author to be analyzed based on the annual publication number in the target author information and the annual publication number in the author information to be analyzed; According to the similarity of the publication years, the similarity analysis results of the career trajectories, the overlap of the active publication periods of the works and the similarity of the publication frequencies, the results of time similarity mining between the target author and the author to be analyzed are generated.

20. The method according to claim 7, wherein Performing similarity mining based on aliases on the target author information in the candidate information set and the author information to be analyzed, including: generating an alias for the name information in the target author information based on an author alias knowledge base, and generating an alias for the name information in the author information to be analyzed based on the author alias knowledge base; Perform extended matching based on the two sets of generated aliases to obtain alias matching results; A confidence evaluation is performed on the alias matching result based on usage frequency, time distribution and source credibility, and an alias similarity mining result between the target author and the author to be analyzed is generated based on the evaluated confidence and the alias matching result.

21. The method according to claim 7, characterized in that Performing similarity mining based on identity identification codes on the target author information in the candidate information set and the author information to be analyzed, including: Performing a complete match between the ORCID in the target author information and the ORCID in the author information to be analyzed to obtain a first match score; Performing fuzzy matching on the Scopus ID in the target author information and the Scopus ID in the author information to be analyzed using an edit distance algorithm to obtain a second matching score; Based on the documents jointly published by the target author and the author to be analyzed, an indirect ID association is established between the target author and the author to be analyzed to obtain a third matching score; Based on the source, verification status, and usage frequency of the identity identification code, credibility weights are respectively configured for the first matching score, the second matching score, and the third matching score, and based on the configured credibility weights and the three matching scores, the results of identity identification code similarity mining between the target author and the author to be analyzed are generated.

22. The method according to claim 1, wherein The candidate information set is obtained through a three-level screening strategy based on surname and initials, institution keywords and year range, and identity identification code; obtaining the candidate information set includes: In the author information database, author information is matched based on the last name and the initial of the author information to be analyzed to obtain a first matching information set that meets the matching condition; In the first matching information set, author information is matched based on the institution keyword and year range of the author information to be analyzed to obtain a second matching information set that meets the matching conditions; In the second matching information set, author information matching is performed based on the identity identification code of the author information to be analyzed, and a third matching information set that meets the matching conditions is obtained as the candidate information set.

23. An author name disambiguation device, characterized in that: include: The information acquisition module is used to obtain a piece of author information to be analyzed and a candidate information set; The candidate information set includes multiple pieces of author information; The author information to be analyzed includes a list of collaborators; a multi-dimensional similarity mining module for performing multi-dimensional similarity mining on the multiple pieces of author information and the author information to be analyzed; wherein the collaborative network is one of the multiple dimensions, wherein the authors are represented by nodes in the collaborative network and the collaborative relationships are represented by lines between the nodes; and similarity mining of the collaborative network includes at least one of the following: direct collaborative relationship similarity mining based on the collaborative network, indirect network relationship similarity mining based on the collaborative network, network structure topology feature similarity mining based on the collaborative network, or temporal network evolution similarity mining based on the collaborative network; A multi-dimensional similarity fusion module is used to fuse the results of multi-dimensional similarity mining between the same author information in the candidate information set and the author information to be analyzed to obtain a multi-dimensional similarity fusion result; The matching module is used to obtain a matching result between the author to be analyzed and the authors in the candidate information set based on the fusion results of the multiple dimensions of similarity between the multiple author information and the author information to be analyzed.

24. An author name disambiguation device, characterized in that include: a processor and memory communicatively connected to each other; The memory stores a computer program; The processor is configured to run the computer program to implement the author name disambiguation method according to any one of claims 1 to 22.

25. A computer-readable storage medium, characterized in that A computer program is stored, and when the computer program is processed, the steps of the author name disambiguation method according to any one of claims 1 to 22 are implemented.

Citation Information

Patent Citations

  • Name duplication disambiguation method of Chinese literature authors

    CN105653590A

  • Vertical domain entity disambiguation method fusing topic model and convolutional neural network

    CN112069826A

  • Literature author name duplication disambiguation method and literature author name duplication disambiguation construction system

    CN112131872A

  • Talent information database disambiguation system

    CN112487825A

  • Knowledge name disambiguation method oriented to specific scientific research tasks

    CN118689915A