Mechanism name normalization method based on deep learning

Through a deep learning-based method, combined with multi-grained feature embedding and context encoding, the problem of insufficient accuracy and automation of organization name standardization in the prior art is solved, and high accuracy and high automation of organization name standardization is achieved, which is suitable for large-scale diversified data processing.

CN120235146AInactive Publication Date: 2025-07-01INSPUR SOFTWARE CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510712714.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-07-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing institutional name standardization method has limited effect when processing large-scale and diversified data, making it difficult to deal with spelling errors, abbreviations and variants, and has a low level of automation and intelligence, so it is impossible to fully capture the semantics and context of institutional names.

Method used

A deep learning-based method is adopted to achieve standardization of organization names through steps such as address retrieval and context selection, multi-grained feature embedding, context encoding and bidirectional matching, similarity calculation and synonym judgment, training objectives and synonym relationship recognition and normalization.

Benefits of technology

It significantly improves the accuracy, degree of automation and ability to process large-scale data institution name norms, can accurately capture character-level and word-level characteristics of organization name, comprehensively capture context semantics, improve synonymous relationship recognition accuracy, and greatly reduce the need for manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235146A_ABST
    Figure CN120235146A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of deep learning and natural language processing, and particularly provides a deep learning-based organization name standardization method, which comprises the following steps of S1, address retrieval and context selection; s2, multi-granularity feature embedding is carried out; s3, performing context coding and bidirectional matching; s4, performing similarity calculation and synonymous judgment; s5, training a target; and S6, identifying and normalizing the synonymous relationship. Compared with the prior art, the organization names in different forms can be standardized automatically and accurately, and the consistency and processing efficiency of the organization names in literature data are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of deep learning and natural language processing, and specifically provides a method for standardizing institutional names based on deep learning. Background Art

[0002] Academic evaluation is a process of evaluating and measuring researchers, institutions, or academic fields, aiming to assess their contributions and influences in the academic community. As a globally renowned academic citation database, the literature data in Web of Science (WOS) is an important data source for academic evaluation. However, since institutional names may exist in various forms in the literature, such as spelling mistakes, abbreviations, variants, etc., this has seriously affected the accuracy of data processing and academic evaluation. Existing methods for standardizing institutional names mostly adopt rule-based or traditional machine learning methods, but these methods have limited effects when dealing with large-scale and diverse data and are difficult to meet the actual application requirements.

[0003] Although the existing technologies have made certain progress in standardizing institutional names, they still face challenges of diversity and complexity, and are difficult to handle spelling mistakes, abbreviations, and variants in different literatures; the context-dependence problem causes the existing methods to be unable to comprehensively capture the semantics and context of institutional names; in terms of data scale and generalization ability, the existing methods are limited in performance when dealing with large-scale data and are difficult to handle unseen institutional names; in addition, the existing methods have low levels of automation and intelligence, rely on a large amount of manual intervention and rule-making, and are difficult to meet the actual application requirements.

[0004] The method based on string similarity compares the string similarity of organization names through edit distance, Jaccard similarity or cosine similarity. It is simple and easy to implement, has strong processing ability for spelling mistakes and small changes, but has poor processing ability for names that are semantically the same but have different forms or different organization names with high literal similarity; The statistical-based method uses context information and statistical features (such as TF-IDF, Naive Bayes) to judge whether they are the same organization. It adapts to literature data in different fields and languages, but relies on manual selection of statistical features and a large amount of data and is easily affected by data quality; The rule-based method relies on preset rules (such as string matching, abbreviation expansion, geographical location information matching) to eliminate ambiguity. It is simple and intuitive and has good processing ability for obvious situations, but rule formulation is complex, requires manual participation, and has poor processing ability for complex situations; The entity-linking-based method links organization names to entities in a knowledge base (such as Wikipedia) to eliminate ambiguity, uses rich knowledge base information to process complex situations, and can automatically construct organization specification files, but requires high-quality knowledge base support and has poor processing ability for new organizations not in the knowledge base; The deep learning-based method uses a deep learning model to learn the semantic features and context information of organization names, and realizes organization disambiguation through vector similarity. It automatically learns and extracts features and has strong processing ability for complex situations, but requires a large amount of labeled data, and model training and tuning are complex. Summary of the Invention

[0005] The present invention aims at the deficiencies of the above-mentioned prior art and provides a highly practical deep learning-based method for standardizing organization names.

[0006] The technical solution adopted by the present invention to solve its technical problems is as follows: A deep learning-based method for standardizing organization names has the following steps: S1. Address retrieval and context selection; S2. Multi-granularity feature embedding; S3. Context encoding and bidirectional matching; S4. Similarity calculation and synonym judgment; S5. Training objective; S6. Synonym relationship recognition and standardization.

[0007] Furthermore, in step S1, for each entity in the candidate organization entity set, the most similar P context fragments are obtained from the associated corpus through a context retriever to obtain a multi-dimensional context information representation; The context information of each organization name is measured and selected through a string similarity method, thereby constructing a context set to be matched.

[0008] Furthermore, in step S2, it includes: S2-1, Character-level feature extraction; S2-2, Word-level feature extraction; S2-3, Feature fusion.

[0009] Further, in step S2-1, character-level feature extraction is performed using Char-CNN. Assuming the character table is , each character is represented as a vector of length d, and the character sequence of a certain word is represented as , then the process of character-level feature embedding of this word is: Let the convolutional kernel be , where k represents the width of the convolutional kernel; The eigenvalue calculation formula is as follows: ; Among them, represents the character sequence in the window, and b is the bias term; Through the max-pooling operation, the global character-level feature representation is extracted.

[0010] Further, in step S2-2, the word-level features are implemented using a combination of Word2Vec and TF-IDF. The Word2Vec captures the semantic relationships of words, and the TF-IDF performs weight adjustment to enhance the features. The length of the word vector is set to d, and the final word-level embedding representation is: .

[0011] Further, in step S2-3, the Highway network is adopted. The Highway network has a gating mechanism and can adaptively control the information flow. The calculation formula of the Highway network is: ; Among them: is the transformation of the input; is the transmission gate, which controls the information flow through the Sigmoid function; x is the input feature, and are weight matrices.

[0012] Further, in step S3, the context encoding adopts the BiLSTM model to encode the context of each institution name. The BiLSTM captures the context information and generates a global context embedding vector. The specific steps are: Let the forward and backward hidden states of the BiLSTM be and , respectively. Then the hidden state representation at each time step is: ; This encoding enables the model to obtain the semantic information of the organization name in the context and improve the ability to recognize synonymous relationships; In the multi-context information matching layer, by calculating the similarity between different contexts, the accuracy of synonymous recognition is enhanced. For two organization entities to be matched and of the context fragment sets, a similarity matrix S is defined: ; where and respectively represent and the i-th and j-th context fragments of, and through normalization operations, the comprehensive context similarity is calculated.

[0013] Furthermore, in step S4, after obtaining the context embedding, cosine similarity is used to calculate the synonymous relationship of the organization name. The final similarity calculation formula is: ; where and are respectively the comprehensive context representations of organizations and . The cosine similarity is used to judge the similarity score, and those exceeding the preset threshold are determined as synonymous organization names.

[0014] Furthermore, in step S5, the model is optimized using the Siamese loss function, which is defined as follows: ; where: represents the label value, taking 1 when it is synonymous and 0 when it is not synonymous; represents the similarity distance, is the margin value margin, which is used to distinguish the boundary between synonymous and non-synonymous.

[0015] Furthermore, in step S6, according to the output result of the BiLSTM model, by setting the similarity threshold, the synonymous organization names of the same entity in different documents are identified; The organization names with similarity exceeding the threshold are normalized to generate a unified representation of the organization name, and the accuracy and stability of the model are further improved through iterative optimization.

[0016] Compared with the prior art, a method for normalizing organization names based on deep learning of the present invention has the following outstanding beneficial effects: The method for standardizing organization names based on deep learning proposed by the present invention significantly improves the accuracy, automation level, and the ability to process large-scale data of organization name standardization.

[0017] (1) By combining the Character Convolutional Neural Network (Char-CNN) and Word2Vec technology, the present invention can accurately capture the character-level and word-level features of organization names, and particularly demonstrates strong robustness when dealing with spelling mistakes, abbreviations, and variants. Compared with traditional rule-based or machine learning methods, the deep learning model of the present invention automatically learns semantic and context features, enhancing the recognition ability for complex organization names.

[0018] (2) Introducing the Bidirectional Long Short-Term Memory Network (BiLSTM) and multiple context matching layers can comprehensively capture the semantic and context information of organization names in different contexts. This multi-context processing method enables the model to maintain consistent recognition effects in different documents, significantly improving the recognition accuracy of synonymous organization names and solving the problem of weak context dependence in the prior art.

[0019] (3) The present invention realizes highly automated organization name standardization through a deep learning model, greatly reducing the need for manual rule formulation and intervention. The automated processing enables this method to perform excellently in large-scale literature data, being able to quickly adapt to literature data in different languages and fields and enhancing the data processing efficiency.

[0020] (4) The present invention uses a deep learning model, which is trained with a large amount of labeled data and has good generalization ability. This method can still maintain high accuracy when dealing with unseen organization names, solving the problem that the performance of existing methods is limited when dealing with large-scale and diverse data.

[0021] (5) The present invention can process and standardize organization names in large-scale academic literature, effectively improving the standardization and consistency of organization information in academic databases, and enhancing the data quality and processing efficiency of scientific research evaluation systems and academic database management systems. Brief Description of the Drawings

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0023] Figure 1 It is a schematic flowchart of a method for standardizing organization names based on deep learning; Figure 2Schematic diagram of the Char-CNN structure in a deep learning-based institutional name normalization method; Figure 3 Schematic diagram of the Highway network structure in a deep learning-based institutional name normalization method. Detailed implementation manners

[0024] To enable those skilled in the art to better understand the solution of the present invention, the present invention will be further described in detail below in conjunction with specific implementation manners. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the scope of protection of the present invention.

[0025] The following gives an optimal embodiment: Embodiment 1: As Figure 1 、 Figure 2 and Figure 3 shown, a deep learning-based institutional name normalization method in this embodiment has the following steps: S1. Address retrieval and context selection; For each entity in the candidate institutional entity set, this method obtains the P most similar context fragments from the associated corpus through a context retriever to obtain a multi-dimensional context information representation. The context information of each institutional name is measured and selected through a string similarity method (such as edit distance, Jaccard similarity), so as to construct a context set to be matched.

[0026] S2. Multi-granularity feature embedding; The multi-granularity feature embedding layer combines character-level features and word-level semantic features, and can identify spelling mistakes, abbreviations, and semantic relationships at the same time.

[0027] The specific implementation is as follows: S2-1. Character-level feature extraction; Character-level feature extraction is performed using Char-CNN (character-level convolutional neural network). Assume that the character table is , each character is represented as a vector of length d, and the character sequence of a certain word is represented as , then the process of character-level feature embedding of this word is: Let the convolution kernel be , where k represents the width of the convolution kernel.

[0028] The eigenvalue calculation formula is as follows: ; Among them, Represents the character sequence in the window, and b is the bias term.

[0029] Through the max-pooling operation, the global character-level feature representation is extracted.

[0030] S2-2, Word-level feature extraction; The word-level features are implemented by the combination of Word2Vec and TF-IDF. Word2Vec captures the semantic relationships of words but lacks weight information, so the TF-IDF weights are adjusted to enhance the features. The length of the word vector is set to d. The final word-level embedding representation is: ; S2-3, Feature fusion; To further integrate the character-level and word-level features, the present invention adopts the Highway Network, which has a gating mechanism and can adaptively control the information flow. The calculation formula of the Highway Network is: ; Where: Is the transformation of the input; Is the transmission gate, which controls the information flow through the Sigmoid function; x is the input feature (obtained by concatenating the character-level and word-level features), And Are the weight matrices.

[0031] S3, Context encoding and bidirectional matching; The context encoding adopts the BiLSTM (Bidirectional Long Short-Term Memory Network) model to encode the context of each institution name. BiLSTM can capture the context information and generate a global context embedding vector. The specific steps are: Let the forward and backward hidden states of BiLSTM be And , then the hidden state representation at each time step is: ; This encoding enables the model to obtain the semantic information of the institution name in the context and improves the ability to identify synonymous relationships.

[0032] In the multi-context information matching layer, the synonymous recognition accuracy is enhanced by calculating the similarity between different contexts. For the context fragment sets of two institution entities And To be matched, the similarity matrix S is defined as: ; Where and respectively represent and the i-th and j-th context fragments of

[0033] S4. Similarity calculation and synonym judgment; After obtaining the context embeddings, cosine similarity is used to calculate the synonym relationship of the organization names. The final similarity calculation formula is: ; where and are respectively the comprehensive context representations of organizations and . The cosine similarity is used to judge the similarity score, and those exceeding the preset threshold are determined as synonymous organization names.

[0034] S5. Training objective; The model is optimized using the Siamese loss function, which is defined as follows: ; where: represents the label value, taking 1 for synonymous and 0 for non-synonymous; represents the similarity distance, is the margin value, used to distinguish the boundary between synonymous and non-synonymous.

[0035] This loss function is optimized on positive and negative example samples, making the distance between synonymous relationship samples smaller and the distance between non-synonymous relationship samples larger, thereby improving the classification accuracy of the model.

[0036] S6. Synonym relationship recognition and normalization; According to the output results of the BiLSTM model, by setting the similarity threshold, the synonymous organization names of the same entity in different documents are recognized.

[0037] The organization names with similarity exceeding the threshold are normalized to generate a unified representation of the organization names, and the accuracy and stability of the model are further improved through iterative optimization.

[0038] Example 2: In this embodiment, the original literature data containing the institution names is first preprocessed to ensure the effectiveness and accuracy of the subsequent model. Literature data is extracted from academic databases such as Web of Science (WOS), and irrelevant characters such as punctuation marks, line breaks, and special symbols are removed. Spelling mistakes are corrected using simple string matching and dictionary-based methods. Appropriate word segmentation tools (such as Jieba, NLTK) are selected according to the language in the literature to segment Chinese and English institution names to ensure the accuracy of the segmentation results.

[0039] Character-level feature extraction is performed. The institution name is converted into a character sequence, and a character convolutional neural network (Char-CNN) is used to perform convolutional operations on the character sequence. The settings of the convolutional kernel size and number enable the network to capture local patterns and features in the character sequence, and it performs excellently especially when dealing with spelling mistakes and variants.

[0040] The Word2Vec model is used to generate the word vector representation of the institution name, and the word vectors are processed by TF-IDF weighting to highlight the words that are important in the literature data. Word2Vec captures the semantic relationships of the words. For example, "Peking University" and "PKU" are represented as similar vectors.

[0041] After the feature extraction is completed, the present invention uses a Highway network to fuse character-level and word-level features. The Highway network controls the information flow through a gating mechanism, and the formula is as follows: ; where t is the transmission gate, is the weight matrix, is the bias, and is the ReLU or Tanh activation function. Through the feature fusion of the Highway network, the present invention can comprehensively integrate character-level and word-level information and enhance the semantic representation of the institution name.

[0042] A bidirectional long short-term memory network (BiLSTM) is used to perform bidirectional processing on the context information of the institution name. The BiLSTM network can capture the semantic and contextual information of the institution name in the context of the literature. To further improve the model performance, a multi-context matching layer is added, and the similarity of institution names in different literatures is calculated through cosine similarity. The model is trained using the labeled institution name data, and the model parameters are optimized through the backpropagation algorithm to enhance the ability of institution name recognition and normalization.

[0043] After the model training is completed, the present invention calculates the similarity between institution names through cosine similarity. According to the set similarity threshold, if the similarity of two institution names exceeds the threshold, they are determined to be a synonymous relationship. The Siamese network structure is adopted, and the following loss function is used to optimize the similarity calculation: ; Among them, represents the loss of synonymous relationship, represents the loss of non-synonymous relationship.

[0044] Finally, according to the synonymous relationship recognition result, the institutional names with synonymous relationships are normalized to a unified representation. For example, it is recognized that "Peking University", "Beida", "Peking University" etc. are the same institution, and they are uniformly normalized to "Peking University". Through the iterative optimization method, the accuracy and consistency of the model in the normalization process are gradually improved.

[0045] The above specific implementation manners are only specific cases of the present invention. The patent protection scope of the present invention includes but is not limited to the above specific implementation manners. Any technical solution that conforms to the above specific implementation manners of the present invention and any appropriate changes or substitutions made by any person of ordinary skill in the art shall fall within the patent protection scope of the present invention.

[0046] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principle and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for standardizing organization names based on deep learning, characterized in that, It has the following steps: S1. Address retrieval and context selection; S2. Multi-granularity feature embedding; S3. Context encoding and bidirectional matching; S4. Similarity calculation and synonym judgment; S5. Training objective; S6. Synonym relationship recognition and normalization.

2. The method for standardizing organization names based on deep learning according to claim 1, wherein In step S1, for each entity in the candidate institutional entity set, the most similar P context fragments are obtained from the associated corpus through the context retriever to obtain a multi-dimensional context information representation; The context information of each institutional name is measured and selected by the string similarity method, thereby constructing a context set to be matched.

3. The method for standardizing organization names based on deep learning according to claim 2, wherein, In step S2, it includes: S2-1. Character-level feature extraction; S2-2. Word-level feature extraction; S2-3. Feature fusion.

4. The method for standardizing institutional names based on deep learning according to claim 3, characterized in that In step S2-1, character-level feature extraction is performed using Char-CNN. Assuming the character set is , each character is represented as a vector of length d, and the character sequence of a certain word is represented as , then the process of character-level feature embedding of this word is as follows: Let the convolution kernel be , where k represents the width of the convolution kernel; The feature value calculation formula is as follows: ; Among them, represents the character sequence in the window, and b is the bias term; Through the max pooling operation, the global character-level feature representation is extracted.

5. The method for standardizing the name of an institution based on deep learning according to claim 4, wherein In step S2-2, the word-level features are implemented by a combination of Word2Vec and TF-IDF. The rd2Vec captures the semantic relationships of words, and the TF-IDF adjusts the weights to enhance the features. The length of the word vector is set to d, and the final word-level embedding representation is: 。 6. The method for standardizing organization names based on deep learning according to claim 5, wherein In step S2-3, the Highway network is adopted. The Highway network has a gating mechanism and can adaptively control the information flow. The calculation formula of the Highway network is: ; Where: is a transformation of the input; is a transmission gate that controls the flow of information through the Sigmoid function; x is the input feature, and is the weight matrix.

7. A method for standardizing organization names based on deep learning according to claim 6, characterized in that, In step S3, the context encoding adopts the BiLSTM model to encode the context of each institutional name. The BiLSTM captures the context information and generates a global context embedding vector. The specific steps are: Let the forward and backward hidden states of the BiLSTM be and , respectively. Then the hidden state at each time step is represented as: ; This encoding enables the model to obtain the semantic information of the institutional name in the context and improves the synonym relationship recognition ability; In the multi-context information matching layer, the accuracy of synonymous recognition is enhanced by calculating the similarity between different contexts. For two institutional entities to be matched and of the context fragment sets, a similarity matrix S is defined as follows: ; wherein and respectively represent and the i-th and j-th context segments of, and the comprehensive context similarity is calculated through a normalization operation.

8. A method for standardizing organization names based on deep learning according to claim 7, characterized in that In step S4, after obtaining the context embedding, the cosine similarity is used to calculate the synonym relationship of the institutional name. The final similarity calculation formula is: ; Among them and are the comprehensive context representations of the institutions and respectively. The cosine similarity is used to judge the similarity score, and those exceeding the preset threshold are determined as synonymous institution names.

9. A method for standardizing organization names based on deep learning according to claim 8, characterized in that In step S5, the model is optimized by the Siamese loss function, and its definition is as follows: ; Where: Indicates the label value. When the value is 1, it indicates synonymy; when the value is 0, it indicates non-synonymy. Indicates the similarity distance, is the margin value, which is used to distinguish the boundary between synonyms and non-synonyms.

10. The method for standardizing organization names based on deep learning according to claim 9, characterized in that, In step S6, according to the output result of the BiLSTM model, by setting the similarity threshold, the synonymous institutional names of the same entity in different documents are identified; The institutional names with similarity exceeding the threshold are normalized to generate a unified institutional name representation, and the accuracy and stability of the model are further improved through iterative optimization.

Citation Information

Patent Citations

  • Chinese sentence semantic intelligent matching method and device based on multi-granularity fusion model

    CN111310438A

  • Regulation retelling matching method and system based on deep learning and cosine similarity matching

    CN119598994A

  • Medical institution name governance method based on comparative learning

    CN119761311A

  • Medical text big data intelligent labeling and knowledge graph construction method and system

    CN119851968A

  • Reading type examination question generation system and method based on commonsense reasoning

    WO2023225858A1