A method and system for correcting historical errors to prevent excessive splitting of scholar paper databases
By reconstructing the scholar's name and using deep learning models to identify entity information, and combining the integrated learning model to judge the alignment of the paper cluster, the problem of oversplit errors in the scholar's paper library is solved, the correction speed and accuracy are improved, and it is suitable for a variety of big data platforms.
Patent Information
- Application Number
- CN202210199458.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-01-19
- Filing Date
- 2022-03-02
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-03-02
AI Technical Summary
Existing scholar paper libraries are prone to over-splitting errors in the process of eliminating author names, resulting in excessive splitting of paper clusters, affecting the recall and accuracy of incremental algorithms.
By reconstructing the scholar's name, quickly obtain the target scholar's paper library and the paper cluster to be allocated; use the BERT-Bi-LSTM-CRF model to identify the author's relevant information and entity information in the abstract; use the entity information and matching degree to calculate the similarity characteristics of the candidate alignment paper cluster and the target scholar's paper library, judge the alignment through the integrated learning model, and merge or manually label to update the model accuracy.
Without losing recall and accuracy, the split errors have been effectively corrected, which improves the correction speed, accuracy and recall rate, and is suitable for different types of technology big data system platforms.
Smart Images

Figure CN114707574B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of information processing technology, and in particular to a method and system for correcting historical errors to prevent excessive splitting of scholar paper libraries. Background Art
[0002] At present, over-splitting of scholar papers is a common problem encountered by the scholar paper database during operation. This problem comes from the over-splitting error scenario that occurs during the operation of the cold start disambiguation and incremental disambiguation algorithms designed by the scholar paper database to eliminate the ambiguity of the same name of different authors. The over-splitting scenario generates a large number of fragment clusters, that is, multiple scholar libraries of the same scholar. The fragment clusters will continue to increase with the increase of system papers, which will lead to a significant reduction in the recall and precision of the incremental algorithm.
[0003] In the related art, in order to avoid over-splitting of scholars' papers, improvements are usually made in the cold start stage and incremental stage of the paper. That is, most of the improvements are made to the name disambiguation algorithm in response to historical errors of over-splitting caused by the name disambiguation algorithm. For example, in the incremental disambiguation algorithm, the INC algorithm directly merges all the over-splitting fragments by checking whether the similarity between the over-splitting fragments is greater than the specified threshold, but this method is prone to incorrect mergers. The MINDi algorithm specifically studies the merging of fragments. When multiple candidate authors meet the specified conditions, it further calculates the distance between the candidate authors and merges the two closest candidate authors. However, this method has shortcomings in accuracy and recall.
[0004] Therefore, in the above scheme for preventing over-splitting by improving the name disambiguation algorithm, under the principle of purity priority, multiple papers by the same author will inevitably be divided into multiple clusters, resulting in over-splitting errors, and the accuracy of over-splitting error correction is low. Summary of the invention
[0005] The present application aims to solve one of the technical problems in the related art at least to some extent.
[0006] To this end, the first purpose of this application is to propose a method for correcting historical errors that prevent over-splitting of scholar paper libraries. This method can correct over-splitting errors generated during the operation of the disambiguation system without losing recall and precision, thereby improving the speed, accuracy and recall rate of over-splitting error correction, and can be applied to different types of scientific and technological big data system platforms, with high versatility.
[0007] The second purpose of this application is to propose a historical error correction system to prevent excessive splitting of scholar paper libraries.
[0008] A third object of the present application is to provide a non-transitory computer-readable storage medium.
[0009] To achieve the above-mentioned purpose, the first embodiment of the present application proposes a method for correcting historical errors in preventing excessive splitting of scholar paper libraries, comprising the following steps:
[0010] By reconstructing the names of scholars, the target scholar's paper database and the corresponding multiple paper clusters to be assigned can be quickly obtained;
[0011] Directly matching the target scholar's paper library and the paper cluster to be assigned according to information that can uniquely identify the author, and directly aligning the paper clusters that are directly matched successfully;
[0012] For the paper clusters that were not successfully matched directly, the bidirectional encoding representation-bidirectional long short-term memory network-conditional random field BERT-Bi-LSTM-CRF model based on the converter is used to identify the author-related information and entity information in the abstracts of the papers in the paper cluster and the target scholar's paper library;
[0013] Calculate the matching degree of the author's institution information and journal information included in the target scholar's paper database and the papers to be matched in the candidate aligned paper cluster respectively;
[0014] Calculate the similarity features of each candidate aligned paper cluster and the target scholar's paper library respectively based on the entity information and the matching degree, construct a corresponding similarity vector based on all similarity features corresponding to each candidate aligned paper cluster, train a corresponding ensemble learning model based on the similarity vector, and judge whether each candidate aligned paper cluster is aligned with the target scholar's paper library through the trained ensemble learning model;
[0015] The candidate aligned paper clusters determined as aligned by the trained ensemble learning model are merged with the target scholar's paper library, and the paper clusters determined as unaligned are manually labeled. According to the labeling results, the misjudged paper clusters are put back into the training set to update the model accuracy.
[0016] Optionally, in one embodiment of the present application, identifying the author-related information and entity information in the abstracts of the papers in the paper cluster and the target scholar paper library includes:
[0017] S11: Determine pre-named entity information, where the entity information includes pre-defined proper noun information;
[0018] S12: vectorizing the author-related information and each word of each sentence in the abstract through a converter-based bidirectional encoding representation BERT model to generate a word vector sequence;
[0019] S13: Input the word vector sequence into a bidirectional long short-term memory network BiLSTM model to extract the context information of the words, and generate the sequence information of the author-related information and the abstract;
[0020] S14: Calculate the conditional probability of each sequence information through the conditional random field CRF model, and output the sequence with the maximum probability corresponding to different entity information;
[0021] S15: Repeat steps S12 to S14 to train the BERT-Bi-LSTM-CRF model by adjusting the structure and hyperparameters of the model;
[0022] S16: Extract the pre-named entity information through the trained BERT-Bi-LSTM-CRF model.
[0023] Optionally, in one embodiment of the present application, the matching degree of the author's institution information and the journal information included in the target scholar's paper library and the papers to be matched in the candidate aligned paper cluster are calculated respectively, including:
[0024] Directly perform string matching on the to-be-matched string containing the author's institution information or the journal information in the to-be-matched paper;
[0025] Perform word segmentation on two to-be-matched strings whose matching degree is below a preset matching degree threshold, and extract a corresponding keyword sequence from each word segmentation sequence;
[0026] Obtaining the bidirectional word order of each of the word segmentation sequences and each of the keyword sequences respectively through the bidirectional long short-term memory network BiLSTM model;
[0027] Calculate the offset of the bidirectional word order of the word segmentation sequence of the author institution or the journal and the bidirectional word order of the keyword sequence of the two strings to be matched;
[0028] Determining similarity vectors of the two to-be-matched strings based on the bidirectional word order of each of the segmentation sequences, the bidirectional word order of each of the keyword sequences, the offset of the bidirectional word order of the segmentation sequences, and the offset of the bidirectional word order of the keyword sequences;
[0029] The similarity vector is input into a preset fully connected layer to calculate the matching degree of the author's institution information or the journal information of the two strings to be matched.
[0030] Optionally, in one embodiment of the present application, the similarity features of each candidate aligned paper cluster and the target scholar paper library are calculated respectively in combination with the entity information and the matching degree, including: according to whether the number of papers contained in each candidate aligned paper cluster exceeds a preset quantity threshold, all candidate aligned paper clusters are divided into a first part and a second part, wherein the number of papers contained in the candidate aligned paper clusters of the first part is greater than or equal to the quantity threshold, and the number of papers contained in the candidate aligned paper clusters of the second part is less than the quantity threshold; for each candidate aligned paper cluster belonging to the first part, the common author similarity, common institution similarity, common journal similarity, word frequency similarity of abstracts and titles, Jaccard similarity, author scarcity, paper publication time overlap and paper proper noun similarity are calculated with the target scholar paper library; for each candidate aligned paper cluster belonging to the second part, the similarity of the author's institution of the paper, the similarity of the journal information in the paper, the author's scarcity, the paper proper noun similarity and the similarity between paper embeddings are calculated with the target scholar paper library.
[0031] Optionally, in one embodiment of the present application, whether each of the candidate aligned paper clusters is aligned with the target scholar's paper library is determined by a trained integrated learning model, including: scoring the degree of alignment between the corresponding candidate aligned paper clusters and the target scholar's paper library through the trained integrated learning model to obtain an alignment degree score for each of the candidate aligned paper clusters; comparing the alignment degree score of any candidate aligned paper cluster with a preset alignment threshold, and if the alignment degree score is greater than the alignment threshold, determining that any of the candidate aligned paper clusters is aligned with the target scholar's paper library.
[0032] Optionally, in one embodiment of the present application, the paper clusters determined to be misaligned are manually labeled, and the misjudged paper clusters are put back into the training set according to the labeling results, including: selecting the target misaligned paper cluster with the highest alignment score among all misaligned paper clusters; manually labeling the target misaligned paper cluster to determine whether the target misaligned paper cluster is assigned to the target scholar's paper library; if it is determined that the target misaligned paper cluster is assigned to the target scholar's paper library, then inputting the target misaligned paper cluster into the training set of the BERT-Bi-LSTM-CRF model and the integrated learning model.
[0033] Optionally, in one embodiment of the present application, the scholar's name is reconstructed, including: reconstructing the scholar's name through local sensitive hashing (LSH), wherein the scholar's name that has a Chinese name is first distinguished by the Chinese name.
[0034] Optionally, in one embodiment of the present application, before reconstructing the scholar's name, it also includes: performing data cleaning on the acquired data containing the target scholar's paper library and the corresponding multiple paper clusters to be assigned, and the data cleaning includes: deleting blank fields and hypertext markup speech HTML tag text information in the data.
[0035] To achieve the above purpose, the second embodiment of the present application proposes a historical error correction system for preventing excessive splitting of scholar paper libraries, including the following modules:
[0036] A construction module is used to quickly obtain the target scholar's paper library and the corresponding multiple paper clusters to be assigned by reconstructing the scholar's name;
[0037] A matching module, used for directly matching the target scholar's paper library and the paper cluster to be assigned according to information that can uniquely identify the author, and directly aligning the paper clusters that are directly matched successfully;
[0038] An identification module is used to identify the author-related information and entity information in the abstracts of the paper clusters that have not been successfully matched directly, using a converter-based bidirectional encoding representation-bidirectional long short-term memory network-conditional random field BERT-Bi-LSTM-CRF model;
[0039] The first calculation module is used to respectively calculate the matching degree of the author's institution information and the journal information included in the target scholar's paper library and the papers to be matched in the candidate alignment paper cluster;
[0040] The second calculation module is used to calculate the similarity features of each candidate aligned paper cluster and the target scholar paper library in combination with the entity information and the matching degree, construct a corresponding similarity vector based on all similarity features corresponding to each candidate aligned paper cluster, and train a corresponding ensemble learning model according to the similarity vector, and judge whether each candidate aligned paper cluster is aligned with the target scholar paper library through the trained ensemble learning model;
[0041] The labeling module is used to merge the candidate aligned paper clusters determined as aligned by the trained ensemble learning model with the target scholar's paper library, and manually label the paper clusters determined as unaligned, and put the misjudged paper clusters back into the training set according to the labeling results to update the model accuracy.
[0042] The technical solution provided by the embodiment of the present application brings at least the following beneficial effects: Based on the idea of integration, the present application solves the defects of the current name disambiguation system under the science and technology big data platform through the method of over-splitting historical error correction. This method can correct the over-splitting errors generated during the operation of the disambiguation system without losing recall and accuracy through five operation processes: rule matching, entity information extraction, paper institution and journal name matching, scholar library alignment and manual annotation. It can perform historical error correction on data at the billion level, improve the speed, accuracy and recall rate of over-splitting error correction, and can be applied to different types of science and technology big data system platforms. It has high versatility and is conducive to improving the name disambiguation system of digital library and other platforms.
[0043] In order to implement the above-mentioned embodiments, the third aspect embodiment of the present application also proposes a non-temporary computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the historical error correction method for preventing excessive splitting of scholar paper libraries in the above-mentioned embodiment is implemented.
[0044] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, wherein
[0046] Figure 1 A flowchart of a method for correcting historical errors to prevent excessive splitting of scholar paper libraries proposed in an embodiment of the present application;
[0047] Figure 2 A flowchart of a specific entity information identification method proposed in an embodiment of the present application;
[0048] Figure 3 A schematic diagram of the structure of a specific historical error correction system for preventing excessive splitting of scholar paper libraries proposed in an embodiment of the present application;
[0049] Figure 4 A flowchart of a specific method for correcting historical errors in preventing excessive splitting of scholar paper libraries proposed in an embodiment of the present application;
[0050] Figure 5 A structural diagram of a historical error correction system for preventing excessive splitting of scholar paper libraries proposed in an embodiment of the present application. DETAILED DESCRIPTION
[0051] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and should not be construed as limiting the present invention.
[0052] The following describes a historical error correction method and system for preventing excessive splitting of scholar paper libraries proposed in an embodiment of the present invention with reference to the accompanying drawings.
[0053] Figure 1 A flowchart of a method for correcting historical errors to prevent excessive splitting of scholar paper libraries proposed in an embodiment of the present application is shown in FIG. Figure 1 The method comprises the following steps:
[0054] Step 101, by reconstructing the scholar's name, quickly obtain the target scholar's paper library and the corresponding multiple paper clusters to be assigned.
[0055] Specifically, for data obtained from a scientific and technological big data platform used by the historical error correction method of this application, such as Aminer or other types of scholar paper libraries, the scholar names are first reconstructed to quickly construct an alternative set to be aligned. The alternative set includes the target scholar paper library and multiple paper clusters corresponding to the paper library whose allocation needs to be determined.
[0056] In the embodiment of the present application, after determining the alternative set, a batch of scholar paper libraries with the same name in a specified format are obtained. There are several papers in the scholar paper library. The number of papers is determined according to the actual situation. The information contained in each paper includes information such as title, author, journal, publication time, keywords, abstract, etc. Among them, the information contained in the author includes information such as the author's name, the author's published papers, and the author's related information. In the subsequent historical error correction method for preventing the scholar paper library from being over-splitting by the present application, the aligned scholar library pairs can be output, that is, each scholar library belonging to the same author.
[0057] In specific implementation, in one embodiment of the present application, the scholar names can be reconstructed by using Locality-Sensitive Hashing (LSH), wherein the scholar names with Chinese names are first distinguished by the Chinese names and stored using a remote dictionary service (Remote Dictionary Server, redis for short), so as to quickly match the scholar paper library to be aligned and its paper cluster to be assigned.
[0058] Specifically, since locality-sensitive hashing can be used for approximate nearest neighbor fast search of massive high-dimensional data, reconstructing the scholar names through locality-sensitive hashing (LSH) can quickly match the paper library and its paper clusters to be assigned. This process reduces the time complexity of a single match from O(n) to O(1), improving the execution efficiency. Moreover, the mapping relationship between the constructed hash value hash and the target paper scholar library id is cached in the redis database to further accelerate the access speed. Thus, an alternative set can be quickly constructed.
[0059] It should be noted that in this embodiment, when constructing the scholar names, Chinese and English names are unified into English. For example, "Li Hua" is converted to "hua li". In addition, for those with Chinese names, the Chinese names should be used for differentiation first. For example, "Li Hua" and "Li Hua" are unambiguous in Chinese but ambiguous in English. Therefore, differentiation is carried out through Chinese first before conversion to further eliminate ambiguity.
[0060] It should also be noted that since the directly obtained data may contain data unrelated to error correction, in an embodiment of the present application, before reconstructing the scholar names, it further includes: cleaning the data obtained, which includes the target scholar paper library and the corresponding multiple paper clusters to be assigned. Data cleaning includes: deleting blank fields and hypertext markup language (HTML) tag text information in the data. For example, the data in AMiner mostly comes from public data on the Internet, so the data needs to be cleaned, deleting the fields with empty values and cleaning the HTML tag text information in the text.
[0061] Step 102, directly match the target scholar paper library and the paper clusters to be assigned according to the information that can uniquely identify the author, and directly align the paper clusters with successful direct matches.
[0062] Among them, the information that can uniquely identify the author is the information that can uniquely identify the author to distinguish from other authors. For example, it can include email and Open Researcher and Contributor ID (ORCID for short), etc., which are not limited here.
[0063] In the embodiment of the present application, the target scholar paper library and the paper clusters to be assigned are directly matched through the information that can uniquely identify the author, such as email and orcid. If a successful match can be made according to this method, the subsequent processing steps are skipped, and the paper clusters with successful matches are directly aligned with the target scholar paper library. For the paper clusters that cannot be directly matched among the multiple paper clusters to be assigned, the subsequent processing steps are executed.
[0064] Step 103, for the paper clusters that are not directly matched successfully, the author-related information and entity information in the abstracts of the papers in the paper clusters and the target scholar's paper library are identified through the converter-based bidirectional encoding representation-bidirectional long short-term memory network-conditional random field BERT-Bi-LSTM-CRF model.
[0065] Among them, entity information includes information such as author institutions, collaborators and research directions contained in contact, as well as proper noun information contained in abstract.
[0066] Specifically, for the paper clusters that were not directly matched successfully in the previous step, pre-define the entity information in contact and abstract, and then use the bidirectional encoder representation from transformers (BERT)-bidirectional long short-term memory network (Bi-LSTM)-conditional random field (CRF) model, i.e., the BERT-Bi-LSTM-CRF model, to pre-train the predefined nouns using a preset data set, such as the Science Citation Index (SCI) data set. Through the trained BERT-Bi-LSTM-CRF model, extract the author-related information and entity information in the abstract of the papers in the paper clusters that were not directly matched successfully and the target scholar's paper library.
[0067] In order to more clearly illustrate the specific implementation method of identifying the relevant entity information in the contact and abstract of the present application, a specific entity information identification method proposed in this embodiment is exemplified below. Figure 2 As shown, the method comprises the following steps:
[0068] S11: Determine pre-named entity information, where the entity information includes pre-defined proper noun information.
[0069] Specifically, determine to extract a series of information such as author institution, collaborators, research direction and publication journal from contact, and extract proper noun information from abstract. In this step, name the above entity information. For example, when naming proper nouns, the intrinsic object structure can be defined as [I-ScientificTerm] through the following text.
[0070] In[O]this[O]paper[O],[O]a[O]novel[O]method[O]to[O]learn[O]the[O]intrinsic[I-ScientificTerm]object[I-ScientificTerm]structure[I-ScientificTerm]for[O]robust[O]visual[O]tracking[O]is[O]proposed[O]. Where [O] can be characters that do not need to be processed, such as spaces.
[0071] S12: The author-related information and each word of each sentence in the abstract are vectorized through the transformer-based bidirectional encoding representation BERT model to generate a word vector sequence.
[0072] Specifically, BERT is used to vectorize the words in a sentence. In this example, BERT can be controlled to use two unsupervised tasks, Mask LM and Next Sentence Prediction, for pre-training to generate features that can describe the character level, word level, sentence level, and inter-sentence level.
[0073] S13: Input the word vector sequence into the BiLSTM model to extract the context information of the words and generate sequence information of author-related information and abstracts.
[0074] Specifically, the word vector sequence generated by S12 is input into BiLSTM to further perceive the contextual information of the word context. Thus, the contextual information of the word and the sequence information of the text are extracted. The sequence information is obtained based on the current word information and context information.
[0075] S14: Calculate the conditional probability of each sequence information through the conditional random field CRF model, and output the sequence with the maximum probability corresponding to different entity information.
[0076] Specifically, the sequence expression with the maximum probability is output through the CRF layer. The conditional probability formula of the CRF model is described as follows:
[0077]
[0078] The formula can be used to calculate the probability that each sequence information corresponds to different entity information. The meaning of each parameter can be referred to the explanation in the relevant technology and will not be repeated here. Then, by comparing the size of the probability value, the sequence with the maximum probability corresponding to each entity information is determined.
[0079] S15: Repeat steps S12 to S14 to train the BERT-Bi-LSTM-CRF model by adjusting the structure and hyperparameters of the model.
[0080] Specifically, the above steps S12 to S14 can be a training cycle of the BERT-Bi-LSTM-CRF model. The BERT-Bi-LSTM-CRF model is trained by repeatedly executing steps S12 to S14 and adjusting the structure of the model, as well as hyperparameters such as the loss function, activation function and learning rate.
[0081] S16: Extract pre-named entity information through the trained BERT-Bi-LSTM-CRF model.
[0082] Specifically, the entity information named in S11 is extracted through the trained model, and the acquired information is integrated into the existing information to provide a data basis for the subsequent feature vector construction.
[0083] Step 104 , respectively calculate the matching degree of the author's institution information and the journal information included in the target scholar's paper database and the papers to be matched in the candidate aligned paper cluster.
[0084] It should be noted that, taking the ambiguity of org as an example, for the two author institutions "Yanshan University" and "Yanshan University Information Science and Engineering", if a direct judgment method such as string matching is used directly, it can be determined that the two are two institutions that cannot be matched. In this example, it can be seen that there is an inclusion relationship between the two, that is, in actual applications, when org does not match, it may be that the author's institution in the paper is only written at the school level but not the college, causing ambiguity. In addition, the same venue may also have different representations. To this end, the embodiment of the present application also matches the author's institution information and journal information included in the papers to be matched in the target scholar's paper library and the candidate aligned paper cluster through a relevant matching algorithm. Among them, the candidate aligned paper cluster is the paper cluster that was not directly matched successfully in step 103 and still needs to be subsequently corrected.
[0085] In specific implementation, as a possible implementation method, the matching method may include the following steps: first, directly performing string matching on the to-be-matched strings containing the author's institution information or journal information in the to-be-matched papers; performing word segmentation processing on the two to-be-matched strings whose matching degree is below a preset matching degree threshold, and extracting the corresponding keyword sequence from each word segmentation sequence; obtaining the bidirectional word order of each word segmentation sequence and each keyword sequence respectively through a bidirectional long short-term memory network BiLSTM model; calculating the offset of the bidirectional word order of the word segmentation sequence of the author's institution or journal and the bidirectional word order of the keyword sequence of the two to-be-matched strings; determining the similarity vectors of the two to-be-matched strings based on the bidirectional word order of each word segmentation sequence, the bidirectional word order of each keyword sequence, the offset of the bidirectional word order of the word segmentation sequence and the offset of the bidirectional word order of the keyword sequence; and inputting the similarity vector into a preset fully connected layer to calculate the matching degree of the author's institution information or journal information of the two to-be-matched strings.
[0086] For example, when matching the org in the target scholar's paper database and any two papers to be matched in the candidate aligned paper cluster (the same applies to venue), the direct matching method is first used to directly perform string matching, that is, matching each character one by one to determine whether the two are exactly the same. If they are completely matched, the matching degree between the two is 1.0.
[0087] For those that do not meet the first step, continue to take the following steps to obtain their matching degree. First, the strings containing org in the two papers, that is, the two strings to be matched, are i ,s j Perform word segmentation processing to obtain v i and v j , and from v i Extract keyword sequence k from i , from v j Extract keyword sequence k from j Then, use BiLSTM to obtain v i ,k i ,v j ,k j The bidirectional word order of Among them, it can be calculated by the following formula:
[0088]
[0089]
[0090] Then, calculate the v between the orgs of the two strings to be matched (i and j) (1) and k (1) The offset d ij , which can be calculated by the following formula:
[0091]
[0092] Then, we get the similarity vector s of the two strings i and j to be matched. ij , and put the similar vectors into a fully connected layer, and use the neural network to predict the similarity between the two vectors to obtain the matching degree y between the two. ij This step is expressed in the following formula:
[0093]
[0094] y ij = linear(s ij ).
[0095] It should be noted that the org and venue matched in this step can be obtained from the entity information identified in step 103.
[0096] Step 105, combining entity information and matching degree to calculate the similarity features of each candidate aligned paper cluster and the target scholar's paper library, constructing a corresponding similarity vector based on all similarity features corresponding to each candidate aligned paper cluster, and training a corresponding ensemble learning model based on the similarity vector, and judging whether each candidate aligned paper cluster is aligned with the target scholar's paper library through the trained ensemble learning model.
[0097] The integrated learning model may be a neural network model integrating an XGBoost model and a gradient boosting decision tree (GBDT) model.
[0098] In one embodiment of the present application, in order to improve the precision and recall rate of the historical error correction method for preventing the over-splitting of the scholar paper library of the present application, the candidate alignment paper clusters are subdivided into two parts according to the number of papers contained in the candidate alignment paper clusters. That is, according to whether the number of papers contained in each candidate alignment paper cluster exceeds the preset quantity threshold, all candidate alignment paper clusters are divided into a first part and a second part, wherein the number of papers contained in the candidate alignment paper clusters in the first part is greater than or equal to the quantity threshold, and the number of papers contained in the candidate alignment paper clusters in the second part is less than the quantity threshold, and the paper clusters in the two parts are aligned separately. As an example, the preset quantity threshold can be 5, that is, the candidate clusters with more than or equal to 5 papers and the candidate clusters with less than 5 papers are aligned respectively.
[0099] Specifically, for each candidate aligned paper cluster belonging to the first part, similarity features such as co-author similarity, co-institution similarity, co-journal similarity, word frequency similarity of abstracts and titles, Jaccard similarity, author scarcity, paper publication time overlap, and paper proper noun similarity are calculated between the cluster and the target scholar's paper library.
[0100] For each candidate aligned paper cluster belonging to the second part, similarity features such as the similarity of the author institutions of the papers between the target scholar's paper library, the similarity of the journal information in the papers, the scarcity of the authors, the similarity of the paper proper nouns, and the similarity between the paper embeddings are calculated.
[0101] In the specific implementation, as the first example, for candidate alignment paper clusters with more than or equal to 5 papers, any one is selected, and the above similarity features are calculated from the two perspectives of the target scholar's paper library and the candidate alignment paper cluster. When calculating the similarity features with the target scholar's paper library, for the co-author (CoAuthor) similarity, al and ar are calculated by the following formula:
[0102]
[0103]
[0104] For the common organization (CoOrg) similarity, ol and or are calculated by the following formulas, wherein when determining whether it is a common organization, the org matching algorithm described in step S104 can be used, which will not be repeated here.
[0105]
[0106]
[0107] For the common journal (CoVenue) similarity, vl and vr are calculated by the following formula, wherein, when determining whether it is a common institution, the venue matching algorithm described in step S104 can be used, which will not be repeated here.
[0108]
[0109]
[0110] For the word frequency similarity between abstract and title, we first need to remove the stop words in the text of abstract+title and obtain its segmentation. The cosine similarity and Jaccard similarity of abstract+title between the scholar library and the candidate cluster are calculated by the following formula: c_s and j_s respectively:
[0111]
[0112]
[0113] The scarcity of the author (r_s) is calculated by the following formula:
[0114] r_s = Sigmoid (number of co-authors)
[0115] For the overlap of paper publication time, tl and tr are calculated by the following formula:
[0116]
[0117]
[0118] For the similarity of proper nouns in the paper (pt_s), the method for obtaining the proper nouns in the paper can be obtained by the method in step S103 and calculated by the following formula:
[0119]
[0120] Furthermore, a corresponding similarity vector is constructed based on all similarity features corresponding to each candidate aligned paper cluster, and a corresponding ensemble learning model is trained based on the similarity vector. The trained ensemble learning model is used to determine whether each candidate aligned paper cluster is aligned with the target scholar's paper library.
[0121] In one embodiment of the present application, a trained ensemble learning model is used to determine whether each candidate aligned paper cluster is aligned with a target scholar's paper library, including the following steps: first, the trained ensemble learning model is used to score the degree of alignment between the corresponding candidate aligned paper cluster and the target scholar's paper library to obtain an alignment degree score for each candidate aligned paper cluster, and then the alignment degree score of the candidate aligned paper cluster is compared with a preset alignment threshold; if the alignment degree score is greater than the alignment threshold, the candidate aligned paper cluster is determined to be aligned with the target scholar's paper library.
[0122] Continuing with the above example, for the candidate aligned paper cluster, a similarity vector v_s = [al, ar, ol, or, vl, vr, c_s, j_r, r_s, tl, tr, pt_s] and its label label∈[0,1] are constructed based on the similarity features calculated above. The constructed vector is then put into the XGBoost&GBDT ensemble learning model for training. After the training is completed, the model is used to score the degree of alignment between the target scholar's paper library and the candidate aligned paper cluster, and compared with the preset alignment threshold to determine whether they are aligned.
[0123] As a second example, for candidate aligned paper clusters with less than 5 papers, when calculating the similarity features with the target scholar's paper library, the similarity (_o_s) of the institutions to which the authors of the papers between the target scholar's paper library belong is calculated using the following formula:
[0124]
[0125] It should be noted that the similarity of the institutions to which the authors of the papers between the candidate aligned paper cluster and the target scholar paper library refers to the org of all the papers in the paper cluster (the org of the paper refers to the org of the relevant authors in the paper) and the org of the scholars to be aligned in the target scholar paper library (that is, the org information extracted from the contact), where the scholar's org can be obtained by the method in step S103, and whether it is a common institution can be determined by the org matching algorithm in step S104.
[0126] The similarity (_v_s) of the journal information in the paper between the target scholar's paper library can be calculated by the following formula:
[0127]
[0128] The scholar's venue may be obtained by the method in step S103 , and whether it is a common journal may be determined by the venue matching algorithm in step S104 .
[0129] For the scarcity of authors (_r_s) and the similarity of proper nouns in papers (_pt_s), the calculation method is the same as that for candidate alignment paper clusters with more than or equal to 5 papers.
[0130] For the similarity between paper embeddings, the OAG-BERT model can be used to obtain Paper_EmbeddingE_1 of the target scholar's paper library and Paper_EmbeddingE_2 of the current candidate aligned paper cluster, and then the cosine_similarity of E_1 and E_2 is calculated to obtain the similarity vector, and a series of key digital features of the vector are extracted, including its maximum value max_, minimum value min_, median_, mean_ and standard deviation std_.
[0131] Furthermore, based on the similarity features calculated above, a similarity vector _v_s = [_o_s, _v_s, _r_s, _pt_s, max_, min_, median_, mean_, std_] and its label label∈[0,1] are constructed. The constructed vector is then put into the XGBoost&GBDT ensemble learning model for training. After the training is completed, the model is used to score the degree of alignment between the target scholar's paper library and the candidate alignment paper cluster, and compared with the preset alignment threshold to determine whether it is aligned.
[0132] Step 106, the candidate aligned paper clusters determined by the trained ensemble learning model are merged with the target scholar's paper library, and the paper clusters determined to be unaligned are manually labeled. According to the labeling results, the misjudged paper clusters are put back into the training set to update the model accuracy.
[0133] Specifically, the aligned target scholar paper libraries and candidate aligned paper clusters are merged pairwise, and the database is updated. The paper clusters that are judged as misaligned by the model are manually labeled, and the misjudged paper clusters are put back into the training set based on the labeling results.
[0134] In one embodiment of the present application, paper clusters judged to be misaligned are manually annotated, and the misjudged paper clusters are put back into the training set according to the annotation results, including: selecting a target misaligned paper cluster with the highest alignment score among all misaligned paper clusters; manually annotating the target misaligned paper cluster to determine whether the target misaligned paper cluster is assigned to the target scholar's paper library; if it is determined that the target misaligned paper cluster is assigned to the target scholar's paper library, then inputting the target misaligned paper cluster into the training set of the BERT-Bi-LSTM-CRF model and the ensemble learning model.
[0135] Specifically, for each paper cluster that the model determines to be misaligned, the Score-Top1 is taken for annotation, where score refers to the alignment score output by the ensemble learning model, and Top1 refers to the paper cluster with the highest alignment score among the paper clusters that the model determines to be misaligned, corresponding to the target scholar's paper library. If this cluster is judged to be misaligned after manual annotation, that is, Score-Top1 should be assigned to the target scholar's paper library (that is, the BERT-Bi-LSTM-CRF model and / or the ensemble learning model has produced a misjudgment phenomenon), the training samples of the model are enriched by inputting the misjudged paper cluster into the training set of the BERT-Bi-LSTM-CRF model and the ensemble learning model, and then the accuracy of the iterative model is updated. For paper clusters that are still judged to be misaligned after manual annotation, the papers in these paper clusters are put back into the set of paper clusters that have not been aligned, so as to facilitate the subsequent determination of which scholar library this cluster should be aligned with.
[0136] Therefore, the historical error correction method for preventing over-splitting of scholar paper libraries in the embodiment of the present application determines whether the target scholar paper library is aligned with multiple paper clusters to be assigned, calculates the matching degree of org and venue in the papers to be matched, and merges the paper libraries that were previously mistakenly split into multiple ones belonging to the same author, so as to correct the historical errors of over-splitting in the scholar paper library.
[0137] In summary, the historical error correction method for preventing over-splitting of scholar paper libraries in the embodiment of the present application, based on the idea of integration, solves the defects of the current name disambiguation system under the science and technology big data platform through the method of over-splitting historical error correction. This method can correct the over-splitting errors generated during the operation of the disambiguation system without losing recall and accuracy through five operation processes: rule matching, entity information extraction, paper institution and journal name matching, scholar library alignment and manual annotation. It can perform historical error correction on data at the billion level, improve the speed, accuracy and recall rate of over-splitting error correction, and can be applied to different types of science and technology big data system platforms, with high versatility, which is conducive to improving the name disambiguation system of platforms such as digital libraries.
[0138] In order to more clearly illustrate the historical error correction method for preventing the scholar paper library from being over-splitting in the embodiment of the present application, the following is combined with Figure 3 and Figure 4 , and a specific example of system development in practical application is used to illustrate:
[0139] In this embodiment, firstly, a Figure 3 The system shown includes a paper cluster alignment subsystem 100, a redis database 200 and a distributed file storage subsystem Mongo300, wherein the paper cluster alignment subsystem 100 integrates a rule matching module (Rule Alignment) 110, an entity information extraction module (Entity Recognition) 120, a scholar library paper organization and journal name disambiguation module 130 (Org&Venue Alignment), a scholar library alignment module (PaperCluster Alignment) 140 and a manual annotation module (Check) to form an automated process operation. The modules in the system are equivalent to containers of different categories. One or more algorithms of a specified category can be placed in a container of the same category. The final result of the container is determined by the AND or logic between the internal algorithms. The model algorithms in the entity information extraction module 120 and the scholar library paper organization and journal name disambiguation module 130 can be pre-trained in an offline manner. The distributed file storage subsystem 300 may include multiple distributed file storage databases, for example, Figure 3A first distributed file storage database 310 and a second distributed file storage database 320 are shown.
[0140] Through the modules in the system performing their respective functions, Figure 4 The historical error correction method for preventing the excessive splitting of scholar paper databases is shown in FIG. Figure 4 As shown, the method comprises the following steps:
[0141] Step S401: Input the scholar paper database and the paper cluster to be assigned.
[0142] In this step, the scholar names are reconstructed by using locality sensitive hashing (LSH) before input and stored using redis, so as to quickly match the scholar's paper library to be aligned and its paper cluster to be assigned.
[0143] Step S402: Determine whether direct matching of rules can be adopted, if so, execute step S403, if not, execute step S404.
[0144] Step S403: Direct alignment.
[0145] In this step, the information that can uniquely identify the author, such as email and orcid, is used to directly perform alignment and matching. For the paper clusters to be assigned that can be directly matched, the paper clusters are directly aligned with the scholar paper database, and the process ends.
[0146] Step S404: Identify related entity information in contact and abstract.
[0147] In this step, the entity information in the contact and abstract is predefined, and the information such as the author organization, collaborators, and research direction can be extracted from the contact, and the proper noun information can be extracted from the abstract. BERT-BiLSTM-CTR is used to pre-train the predefined nouns using the sci dataset, and its hyperparameters are adjusted to the optimal level to further extract the predefined proper nouns.
[0148] Step S405: Match org and venue in the matching papers.
[0149] This step is to determine whether org and venue are one or have a relationship. In this step, string matching is first used for direct simple matching. For those that cannot be simply matched, the BiLSTM model is used to extract bidirectional sequence information. Based on the extracted information, a linear layer is trained to obtain the similarity between org and venue. In this process, some org and venue come from step S404.
[0150] Step S406: Extract features and take appropriate thresholds for alignment based on model scores.
[0151] In this step, the data set is divided, and the candidate alignment scholar library is processed separately according to whether the number of papers under it is greater than or equal to 5. After extracting the corresponding series of features, the two types of feature matrices constructed are respectively put into the integrated learning model composed of XGboost and GBDT, and the threshold is selected for alignment. In this process, the similarity calculation of Org and Venue adopts the method defined in step S405. The method for obtaining the proper nouns of the paper adopts the method defined in step S404.
[0152] Step S407: Manually label the paper clusters that are judged to be misaligned by the model.
[0153] In this step, for samples predicted to be non-aligned, their Score-Top1 is taken for annotation to determine whether they are predicted incorrectly. The samples judged to be incorrectly judged are put back into the training set for training to update the iterative algorithm, and step S406 is re-executed. For paper clusters that are still judged to be non-aligned after manual annotation, the papers in these paper clusters are put back into the set of non-aligned paper clusters, and the process ends.
[0154] In order to implement the above embodiment, the present application also proposes a historical error correction system to prevent excessive splitting of scholar paper libraries. Figure 5 This is a schematic diagram of the structure of a historical error correction system for preventing excessive splitting of scholar paper libraries proposed in an embodiment of the present application.
[0155] like Figure 5 As shown, the system includes: a construction module 100 , a matching module 200 , a recognition module 300 , a first calculation module 400 , a second calculation module 500 and a labeling module 600 .
[0156] Among them, the construction module 100 is used to quickly obtain the target scholar's paper library and the corresponding multiple paper clusters to be assigned by reconstructing the scholar's name.
[0157] The matching module 200 is used to directly match the target scholar's paper library and the paper cluster to be assigned based on the information that can uniquely identify the author, and directly align the paper clusters that are directly matched successfully.
[0158] The identification module 300 is used to identify the author-related information and entity information in the abstracts of the paper clusters and the papers in the paper library of the target scholar for the paper clusters that have not been directly matched successfully, through the bidirectional encoding representation-bidirectional long short-term memory network-conditional random field BERT-Bi-LSTM-CRF model based on the converter.
[0159] The first calculation module 400 is used to respectively calculate the matching degree of the author's institution information and the journal information included in the papers to be matched in the target scholar's paper database and the candidate aligned paper cluster.
[0160] The second calculation module 500 is used to calculate the similarity features of each candidate aligned paper cluster and the target scholar's paper library based on the entity information and the matching degree, construct a corresponding similarity vector based on all the similarity features corresponding to each candidate aligned paper cluster, and train the corresponding integrated learning model according to the similarity vector, and judge whether each candidate aligned paper cluster is aligned with the target scholar's paper library through the trained integrated learning model.
[0161] The labeling module 600 is used to merge the candidate aligned paper clusters determined as aligned by the trained ensemble learning model with the target scholar's paper library, and manually label the paper clusters determined as unaligned, and put the misjudged paper clusters back into the training set based on the labeling results to update the model accuracy.
[0162] Optionally, in one embodiment of the present application, the recognition module 300 specifically includes: a determination unit, which is used to determine pre-named entity information, and the entity information includes pre-defined proper noun information; a conversion unit, which is used to vectorize the author-related information and each word of each sentence in the abstract through a bidirectional encoding representation BERT model based on a converter, and generate a word vector sequence; a generation unit, which is used to input the word vector sequence into a bidirectional long short-term memory network BiLSTM model to extract the context information of the words, and generate sequence information of the author-related information and the abstract; a calculation unit, which is used to calculate the conditional probability of each of the sequence information through a conditional random field CRF model, and output a sequence with the maximum probability corresponding to different entity information; a model training unit, which is used to control the conversion unit, the generation unit and the calculation unit to repeatedly perform their own functions, and train the BERT-Bi-LSTM-CRF model by adjusting the structure and hyperparameters of the model; an extraction unit, which is used to extract the pre-named entity information through the trained BERT-Bi-LSTM-CRF model.
[0163] In one embodiment of the present application, the first calculation module 400 is specifically used to: directly perform string matching on the to-be-matched strings containing the author's institution information or journal information in the to-be-matched papers; perform word segmentation processing on the two to-be-matched strings whose matching degree is below a preset matching degree threshold, and extract the corresponding keyword sequence from each word segmentation sequence; obtain the bidirectional word order of each word segmentation sequence and each keyword sequence respectively through a bidirectional long short-term memory network BiLSTM model; calculate the offset of the bidirectional word order of the word segmentation sequence of the author's institution or journal and the bidirectional word order of the keyword sequence of the two to-be-matched strings; determine the similarity vector of the two to-be-matched strings based on the bidirectional word order of each of the word segmentation sequences, the bidirectional word order of each keyword sequence, the offset of the bidirectional word order of the word segmentation sequence and the offset of the bidirectional word order of the keyword sequence; input the similarity vector into a preset fully connected layer to calculate the matching degree of the author's institution information or journal information of the two to-be-matched strings.
[0164] Optionally, in one embodiment of the present application, the second calculation module 500 is specifically used to: divide all candidate aligned paper clusters into a first part and a second part according to whether the number of papers contained in each candidate aligned paper cluster exceeds a preset quantity threshold, wherein the number of papers contained in the candidate aligned paper clusters in the first part is greater than or equal to the quantity threshold, and the number of papers contained in the candidate aligned paper clusters in the second part is less than the quantity threshold; for each candidate aligned paper cluster belonging to the first part, calculate the common author similarity, common institution similarity, common journal similarity, word frequency similarity of abstracts and titles, Jaccard similarity, author scarcity, paper publication time overlap and paper proper noun similarity with the target scholar paper library; for each candidate aligned paper cluster belonging to the second part, calculate the similarity of the author's institution of the paper, the similarity of the journal information in the paper, the author's scarcity, the paper proper noun similarity and the similarity between paper embeddings with the target scholar paper library.
[0165] Optionally, in one embodiment of the present application, the second calculation module 500 is also used to: score the degree of alignment between the corresponding candidate aligned paper clusters and the target scholar's paper library through the trained integrated learning model to obtain the alignment degree score of each candidate aligned paper cluster; compare the alignment degree score of any candidate aligned paper cluster with a preset alignment threshold, and if the alignment degree score is greater than the alignment threshold, then determine that any candidate aligned paper cluster is aligned with the target scholar's paper library.
[0166] Optionally, in one embodiment of the present application, the second calculation module 500 is also used to: select a target unaligned paper cluster with the highest alignment score among all unaligned paper clusters; manually annotate the target unaligned paper cluster to determine whether the target unaligned paper cluster is assigned to the target scholar's paper library; if it is determined that the target unaligned paper cluster is assigned to the target scholar's paper library, then input the target unaligned paper cluster into the training set of the BERT-Bi-LSTM-CRF model and the integrated learning model.
[0167] Optionally, in one embodiment of the present application, the construction module 100 is specifically used to: reconstruct the scholar's name by using local sensitive hashing (LSH), wherein the scholar's name that has a Chinese name is first distinguished by the Chinese name.
[0168] Optionally, in one embodiment of the present application, the construction module 100 is also used to: perform data cleaning on the acquired data containing the target scholar's paper library and the corresponding multiple paper clusters to be assigned, and the data cleaning includes: deleting blank fields and hypertext markup speech HTML tag text information in the data.
[0169] In summary, the historical error correction system for preventing over-splitting of scholar paper libraries in the embodiment of the present application, based on the idea of integration, solves the defects of the current name disambiguation system under the science and technology big data platform through the method of over-splitting historical error correction. The system includes five parts: rule matching, entity information extraction, paper institution and journal name matching, scholar library alignment and manual annotation. It can correct the over-splitting errors generated during the operation of the disambiguation system without losing recall and accuracy. It can perform historical error correction on data at the billion level, improve the speed, accuracy and recall rate of over-splitting error correction, and can be applied to different types of science and technology big data system platforms. It has high versatility and is conducive to improving the name disambiguation system of platforms such as digital libraries.
[0170] In order to implement the above-mentioned embodiments, the present invention also proposes a non-temporary computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, a historical error correction method for preventing excessive splitting of a scholar's paper library described in the first aspect of the embodiment of the present application is implemented.
[0171] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, if schematic expressions of the above terms are used in multiple embodiments or examples, it does not mean that these embodiments or examples are the same. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, unless they are contradictory.
[0172] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0173] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.
[0174] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute the instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purpose of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways if necessary, and then stored in a computer memory.
[0175] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0176] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.
[0177] In addition, each functional unit in each embodiment of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0178] The storage medium mentioned above may be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limiting the present application. A person of ordinary skill in the art may change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A historical error correction method to prevent over-splitting of scholar paper databases, It is characterized in that The following steps are involved: By reconstructing the names of scholars, the target scholar's paper database and the corresponding multiple paper clusters to be assigned can be quickly obtained; Directly matching the target scholar's paper library and the paper cluster to be assigned according to information that can uniquely identify the author, and directly aligning the paper clusters that are directly matched successfully; For the paper clusters that were not successfully matched directly, the bidirectional encoding representation-bidirectional long short-term memory network-conditional random field BERT-Bi-LSTM-CRF model based on the converter is used to identify the author-related information and entity information in the abstracts of the papers in the paper cluster and the target scholar's paper library; Calculate the matching degree of the author's institution information and journal information included in the target scholar's paper database and the papers to be matched in the candidate aligned paper cluster respectively; The similarity features between each candidate aligned paper cluster and the target scholar's paper library are calculated respectively in combination with the entity information and the matching degree, a corresponding similarity vector is constructed based on all the similarity features corresponding to each candidate aligned paper cluster, and a corresponding integrated learning model is trained according to the similarity vector, and whether each candidate aligned paper cluster is aligned with the target scholar's paper library is judged by the trained integrated learning model; wherein, the similarity features between each candidate aligned paper cluster and the target scholar's paper library are calculated respectively in combination with the entity information and the matching degree, including: according to whether the number of papers contained in each candidate aligned paper cluster exceeds a preset number threshold, all candidate aligned paper clusters are divided into a first part and a second part, wherein the first part The number of papers contained in the candidate aligned paper clusters of the first part is greater than or equal to the number threshold, and the number of papers contained in the candidate aligned paper clusters of the second part is less than the number threshold; for each candidate aligned paper cluster belonging to the first part, the common author similarity, common institution similarity, common journal similarity, word frequency similarity of abstracts and titles, Jaccard similarity, author scarcity, paper publication time overlap and paper proper noun similarity between the candidate aligned paper clusters and the target scholar paper library are calculated; for each candidate aligned paper cluster belonging to the second part, the similarity of the institutions to which the authors of the papers belong, the similarity of the journal information in the papers, the scarcity of the authors, the similarity of the proper nouns of the papers and the similarity between the paper embeddings are calculated between the candidate aligned paper clusters and the target scholar paper library; The candidate aligned paper clusters determined as aligned by the trained ensemble learning model are merged with the target scholar's paper library, and the paper clusters determined as unaligned are manually labeled. According to the labeling results, the misjudged paper clusters are put back into the training set to update the model accuracy.
2. The method according to claim 1, It is characterized in that The identification of the author-related information and entity information in the abstract of the papers in the paper cluster and the target scholar's paper library includes: S11: Determine pre-named entity information, where the entity information includes pre-defined proper noun information; S12: vectorizing the author-related information and each word of each sentence in the abstract through a converter-based bidirectional encoding representation BERT model to generate a word vector sequence; S13: Input the word vector sequence into a bidirectional long short-term memory network BiLSTM model to extract the context information of the words, and generate the sequence information of the author-related information and the abstract; S14: Calculate the conditional probability of each sequence information through the conditional random field CRF model, and output the sequence with the maximum probability corresponding to different entity information; S15: Repeat steps S12 to S14 to train the BERT-Bi-LSTM-CRF model by adjusting the structure and hyperparameters of the model; S16: Extract the pre-named entity information through the trained BERT-Bi-LSTM-CRF model.
3. The method according to claim 1 or 2, It is characterized in that The step of respectively calculating the matching degree of the author's institution information and the journal information included in the target scholar's paper database and the papers to be matched in the candidate aligned paper clusters includes: Directly perform string matching on the to-be-matched string containing the author's institution information or the journal information in the to-be-matched paper; Perform word segmentation on two to-be-matched strings whose matching degree is below a preset matching degree threshold, and extract a corresponding keyword sequence from each word segmentation sequence; Obtaining the bidirectional word order of each of the word segmentation sequences and each of the keyword sequences respectively through the bidirectional long short-term memory network BiLSTM model; Calculate the offset of the bidirectional word order of the word segmentation sequence of the author institution or the journal and the bidirectional word order of the keyword sequence of the two strings to be matched; Determining similarity vectors of the two to-be-matched strings based on the bidirectional word order of each of the segmentation sequences, the bidirectional word order of each of the keyword sequences, the offset of the bidirectional word order of the segmentation sequences, and the offset of the bidirectional word order of the keyword sequences; The similarity vector is input into a preset fully connected layer to calculate the matching degree of the author's institution information or the journal information of the two strings to be matched.
4. The method according to claim 1, It is characterized in that The integrated learning model completed through training determines whether each candidate aligned paper cluster is aligned with the target scholar's paper library, including: Scoring the alignment degree between the corresponding candidate aligned paper cluster and the target scholar's paper library through the trained integrated learning model to obtain the alignment degree score of each candidate aligned paper cluster; The alignment degree score of any candidate aligned paper cluster is compared with a preset alignment threshold. If the alignment degree score is greater than the alignment threshold, it is determined that any candidate aligned paper cluster is aligned with the target scholar's paper library.
5. The method according to claim 4, It is characterized in that Manually labeling the paper clusters that are determined to be misaligned, and re-placing the misjudged paper clusters into the training set according to the labeling results, includes: Selecting a target unaligned paper cluster with the highest alignment score from all unaligned paper clusters; Manually annotate the target unaligned paper clusters to determine whether the target unaligned paper clusters are assigned to the target scholar paper library; If it is determined that the target unaligned paper cluster is assigned to the target scholar paper library, the target unaligned paper cluster is input into the training set of the BERT-Bi-LSTM-CRF model and the integrated learning model.
6. The method according to claim 1, It is characterized in that The reconstruction of the names of scholars includes: The names of scholars are reconstructed by using local sensitive hashing (LSH), wherein the names of scholars that have Chinese names are first distinguished by using the Chinese names.
7. The method according to claim 1, It is characterized in that Before reconstructing the names of scholars, it also includes: The acquired data including the target scholar's paper library and the corresponding multiple paper clusters to be assigned are cleaned, and the data cleaning includes: deleting blank fields and hypertext markup speech (HTML) tag text information in the data.
8. A historical error correction system to prevent over-splitting of scholar paper repositories, It is characterized in that include: A construction module is used to quickly obtain the target scholar's paper library and the corresponding multiple paper clusters to be assigned by reconstructing the scholar's name; A matching module, used for directly matching the target scholar's paper library and the paper cluster to be assigned according to information that can uniquely identify the author, and directly aligning the paper clusters that are directly matched successfully; An identification module is used to identify the author-related information and entity information in the abstracts of the paper clusters that have not been successfully matched directly, using a converter-based bidirectional encoding representation-bidirectional long short-term memory network-conditional random field BERT-Bi-LSTM-CRF model; The first calculation module is used to respectively calculate the matching degree of the author's institution information and the journal information included in the target scholar's paper library and the papers to be matched in the candidate alignment paper cluster; The second calculation module is used to calculate the similarity features of each candidate aligned paper cluster and the target scholar's paper library based on the entity information and the matching degree, construct a corresponding similarity vector based on all the similarity features corresponding to each candidate aligned paper cluster, and train the corresponding integrated learning model according to the similarity vector, and judge whether each candidate aligned paper cluster is aligned with the target scholar's paper library through the trained integrated learning model; wherein, the second calculation module is specifically used to: divide all candidate aligned paper clusters into a first part and a second part according to whether the number of papers contained in each candidate aligned paper cluster exceeds a preset number threshold, wherein the candidate aligned paper clusters in the first part contain papers The number is greater than or equal to the number threshold, and the number of papers contained in the candidate aligned paper clusters of the second part is less than the number threshold; for each candidate aligned paper cluster belonging to the first part, calculate the common author similarity, common institution similarity, common journal similarity, word frequency similarity of abstracts and titles, Jaccard similarity, author scarcity, paper publication time overlap and paper proper noun similarity with the target scholar paper library; for each candidate aligned paper cluster belonging to the second part, calculate the similarity of the institutions to which the authors of the papers belong, the similarity of the journal information in the papers, the scarcity of the authors, the similarity of the paper proper nouns and the similarity between the paper embeddings with the target scholar paper library; The labeling module is used to merge the candidate aligned paper clusters determined as aligned by the trained ensemble learning model with the target scholar's paper library, and manually label the paper clusters determined as unaligned, and put the misjudged paper clusters back into the training set according to the labeling results to update the model accuracy.
9. A non-transitory computer-readable storage medium having stored thereon a computer program, It is characterized in that When the computer program is executed by a processor, a historical error correction method for preventing excessive splitting of scholar paper libraries is implemented as described in any one of claims 1-7.
Citation Information
Patent Citations
Homonymous author disambiguation method based on network representation and semantic representation
CN111191466A
Scholar name disambiguation method and device, storage medium and terminal
CN111581949A