Method and apparatus for knowledge fusion of scientific concept system
Patent Information
- Application Number
- CN202310190487.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-23
- Publication Date
- 2026-08-28
- Estimated Expiration
- 2043-02-23
AI Technical Summary
[0005]为此,本申请的第一个目的在于提出一种对科技概念体系进行知识融合的方法,解决了现有方法领域涵盖窄、分类不够准确的技术问题,通过整合不同粒度和领域的分类数据构建一个种类丰富且质量较高的树形结构,能够在融合多种数据源的同时提高科技概念分类的准确性
[0042]本申请实施例的对科技概念体系进行知识融合的方法、装置、计算机设备和非临时性计算机可读存储介质,解决了现有方法领域涵盖窄、分类不够准确的技术问题,通过整合不同粒度和领域的分类数据构建一个种类丰富且质量较高的树形结构,能够在融合多种数据源的同时提高科技概念分类的准确性。
Smart Images

Figure CN116304037B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of technology of integrating scientific and technological concept systems, and in particular to a method and apparatus for integrating scientific and technological concept systems. Background Technology
[0002] Concept trees abstract domain-specific knowledge into a hierarchical structure according to certain standards, serving as the foundation for text structuring and management. Concept trees are widely used in various applications, such as representing disciplinary dependencies using natural science systems, managing journal articles, and recommending peer experts. Most concept trees are constructed by experts, relying on human intervention. Methods that organize concept trees by mining the hierarchical relationships between words in text suffer from inconsistent granularity and are not applicable to specific domains. Automated concept tree construction methods within a domain can be divided into two categories: one involves completing or expanding an existing system to create new subordinate concepts; the other involves generating a tree-like structure of a certain number of levels from top to bottom based on the contextual features of text and concept associations, even without a pre-existing system. Concept tree completion uses a self-supervised approach to train and predict on an existing system. While self-supervised training can predict hierarchical relationships of concepts well, it cannot judge the granularity and quality of concepts.
[0003] The existing science and technology classification system suffers from narrow field coverage and inaccurate classification, while also lacking the ability to be continuously updated, thus failing to effectively support data governance and downstream applications. Summary of the Invention
[0004] This application aims to at least partially address one of the technical problems in the related art.
[0005] Therefore, the first objective of this application is to propose a method for knowledge fusion of scientific and technological concept systems, which solves the technical problems of narrow domain coverage and inaccurate classification in existing methods. By integrating classification data of different granularities and domains to construct a rich and high-quality tree structure, it can improve the accuracy of scientific and technological concept classification while integrating multiple data sources.
[0006] The second objective of this application is to propose a device for knowledge fusion of a system of scientific and technological concepts.
[0007] The third objective of this application is to propose a computer device.
[0008] The fourth objective of this application is to provide a non-transitory computer-readable storage medium.
[0009] To achieve the above objectives, the first aspect of this application proposes a method for knowledge fusion of scientific and technological concept systems, comprising: acquiring multiple scientific and technological concept systems and dividing the multiple scientific and technological concept systems to obtain a target scientific and technological concept system and multiple scientific and technological concept systems to be fused, wherein the structure of the scientific and technological concept system is a tree-like hierarchical structure composed of different nodes; identifying nodes to be split by splitting rules in the target scientific and technological concept system and the scientific and technological concept systems to be fused, splitting the concept names of the nodes to be split according to manual annotation rules, and updating the nodes of the target scientific and technological concept system and the scientific and technological concept systems to be fused according to the splitting results; detecting the nodes of the updated target scientific and technological concept system and the scientific and technological concept systems to be fused by similarity calculation to obtain nodes with concept names with the same meaning, and fusion of the updated target scientific and technological concept system and the scientific and technological concept systems to be fused according to the nodes with concept names with the same meaning; calculating the confidence of the nodes of the scientific and technological concept systems to be fused and the parent nodes of the fused target scientific and technological concept system, and fusion according to the confidence to obtain the fused scientific and technological concept system.
[0010] Optionally, in one embodiment of this application, the nodes of the target technology concept system and the technology concept system to be merged are identified by splitting rules to obtain the nodes to be split, the concept names of the nodes to be split are split according to manual annotation rules, and the nodes of the target technology concept system and the technology concept system to be merged are updated according to the splitting results, including:
[0011] Obtain the names of the concepts to be split and analyze them to obtain the splitting rules;
[0012] Based on the splitting rules, the nodes of the target technology concept system and the technology concept system to be integrated are identified through regular expression analysis to obtain the nodes to be split.
[0013] The concept names of the nodes to be split are split according to the manual annotation rules to obtain the splitting results, and the structure of the target technology concept system and the technology concept system to be integrated is adjusted according to the splitting results.
[0014] Optionally, in one embodiment of this application, nodes of the updated target technology concept system and the technology concept system to be merged are detected through similarity calculation to obtain nodes with concept names having the same meaning, including:
[0015] Calculate the character similarity, vector similarity, and structural similarity between nodes in the updated target technology concept system and the technology concept system to be merged;
[0016] By linearly combining character similarity, vector similarity, and structural similarity, the similarity between nodes in the updated target technology concept system and the technology concept system to be merged is obtained. Nodes with similarity greater than a first preset threshold are regarded as nodes with concept names having the same meaning.
[0017] Optionally, in one embodiment of this application, before calculating the confidence scores of the nodes of the technology concept system to be fused and the parent nodes of the fused target technology concept system, and performing fusion based on the confidence scores to obtain the fused technology concept system, the following steps are included:
[0018] The concept-to-relationship prediction model uses semantic vector representations for the nodes of the technology concept system to be integrated and the target technology concept system after integration.
[0019] The nodes of the merged target technology concept system and the technology concept system to be merged are retrieved for higher-level and lower-level information respectively, and information fragments of the nodes are constructed.
[0020] Optionally, in one embodiment of this application, the confidence scores of the nodes of the technological concept system to be fused and the parent nodes of the fused target technological concept system are calculated, and fusion is performed based on the confidence scores to obtain the fused technological concept system, including:
[0021] The hierarchical relationship between the nodes of the technology concept system to be merged and the parent node of the target technology concept system after merging is predicted by the concept pair relationship prediction model, and the corresponding probability values are obtained.
[0022] The parent node of the fused target technology concept system with a probability value greater than the second preset threshold is selected as the candidate node of the node of the technology concept system to be fused, and the confidence between the node of the technology concept system to be fused and the candidate node is calculated based on the information fragment of the node.
[0023] Select the candidate node with the highest confidence level and greater than the third preset threshold to attach the node of the technology concept system to be integrated, and obtain the integrated technology concept system.
[0024] Optionally, in one embodiment of this application, the information fragment of a node includes the node, the node's parent node, and the node's child nodes. Calculating the confidence level between nodes and candidate nodes in the technology concept system to be fused based on the node's information fragment includes:
[0025] A score matrix is constructed based on the information fragments of the nodes and candidate nodes in the technology concept system to be integrated.
[0026] Calculate the fraction matrix according to the first calculation formula, and then adjust the fraction matrix.
[0027] The average score of the adjusted score matrix is used as the confidence level between the nodes and candidate nodes of the technology concept system to be integrated.
[0028] The first calculation formula is expressed as follows:
[0029]
[0030] Among them, s k,b This represents the value at the k-th row and b-th column in the fraction matrix, match(s) k ,s b ) represents the string similarity between nodes, sim(s) k ,s b ) represents the similarity of concept names between nodes, count(s) k ,s b ) represents the conceptual name of the node. k and s b The number of identical characters between two lines.
[0031] To achieve the above objectives, a second aspect of this application provides an apparatus for knowledge fusion of a scientific and technological concept system, comprising:
[0032] The acquisition module is used to acquire multiple technology concept systems and divide them into a target technology concept system and multiple technology concept systems to be integrated. The structure of the technology concept system is a tree-like hierarchical structure composed of different nodes.
[0033] The splitting module is used to identify the nodes of the target technology concept system and the technology concept system to be integrated through splitting rules, obtain the nodes to be split, split the concept names of the nodes to be split according to the manual annotation rules, and update the nodes of the target technology concept system and the technology concept system to be integrated based on the splitting results.
[0034] The first fusion module is used to detect the nodes of the updated target technology concept system and the technology concept system to be fused through similarity calculation, obtain the nodes with concept names with the same meaning, and fuse the updated target technology concept system and the technology concept system to be fused based on the nodes with concept names with the same meaning.
[0035] The second fusion module is used to calculate the confidence of the nodes of the technology concept system to be fused and the parent nodes of the target technology concept system after fusion, and to perform fusion based on the confidence to obtain the fused technology concept system.
[0036] Optionally, in one embodiment of this application, the split module is specifically used for:
[0037] Obtain the names of the concepts to be split and analyze them to obtain the splitting rules;
[0038] Based on the splitting rules, the nodes of the target technology concept system and the technology concept system to be integrated are identified through regular expression analysis to obtain the nodes to be split.
[0039] The concept names of the nodes to be split are split according to the manual annotation rules to obtain the splitting results, and the structure of the target technology concept system and the technology concept system to be integrated is adjusted according to the splitting results.
[0040] To achieve the above objectives, a third aspect of this application provides a computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the method for knowledge fusion of the scientific and technological concept system described in the above embodiments.
[0041] To achieve the above objectives, a fourth aspect of this application provides a non-transitory computer-readable storage medium that, when the instructions in the storage medium are executed by a processor, enables the execution of a method for knowledge fusion of a system of scientific and technological concepts.
[0042] The method, apparatus, computer equipment, and non-transitory computer-readable storage medium for knowledge fusion of scientific and technological concept systems in this application solve the technical problems of narrow domain coverage and inaccurate classification in existing methods. By integrating classification data of different granularities and domains to construct a tree structure with rich variety and high quality, it can improve the accuracy of scientific and technological concept classification while integrating multiple data sources.
[0043] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0044] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0045] Figure 1 This is a flowchart illustrating a method for knowledge fusion of a scientific and technological concept system provided in Embodiment 1 of this application;
[0046] Figure 2 This is another flowchart illustrating the method for knowledge fusion of a scientific and technological concept system according to an embodiment of this application;
[0047] Figure 3 An example diagram of a scientific and technological concept system for the method of knowledge fusion of a scientific and technological concept system according to an embodiment of this application;
[0048] Figure 4This is an example diagram of the concept to be decomposed for the method of knowledge fusion of a scientific and technological concept system according to an embodiment of this application.
[0049] Figure 5 This is a concept alignment information association diagram for the method of knowledge fusion of a scientific and technological concept system according to an embodiment of this application;
[0050] Figure 6 This is a diagram illustrating the concept pair relationship prediction model structure of the method for knowledge fusion of a scientific and technological concept system according to an embodiment of this application.
[0051] Figure 7 This is a concept completion node information association diagram for the method of knowledge fusion of scientific and technological concept system in this application embodiment;
[0052] Figure 8 This is an example diagram of the fusion score matrix of the method for knowledge fusion of a scientific and technological concept system according to an embodiment of this application;
[0053] Figure 9 This is a schematic diagram of a device for knowledge fusion of a scientific and technological concept system provided in Embodiment 2 of this application. Detailed Implementation
[0054] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.
[0055] A science and technology concept system is a hierarchical structure organized according to certain logical rules, based on the fields, methods, and technologies of scientific research in the real world. Existing science and technology classification systems suffer from narrow field coverage, inaccurate classification, lack of continuous updating capabilities, and inability to effectively support data governance and downstream applications. This application designs a method for knowledge fusion of science and technology concept systems, aiming to construct a rich and high-quality tree structure by integrating classification data of different granularities and fields. The main challenge of knowledge fusion lies in how to solve the concept alignment and hierarchical relationship identification between different classification systems. This application can consider the accuracy of science and technology concept classification while integrating multiple data sources.
[0056] The main technical architecture of this application consists of three parts: standardizing category names, aligning classification concepts, and completing the classification system. Standardizing category names involves splitting and correcting the input raw data to ensure the atomicity of categories. Aligning classification concepts involves discovering the correspondences between concepts in different classification systems and collecting concepts with the same meaning. Completing the classification system involves finding the most likely parent node of the classification nodes in the technology system to be integrated within the existing structure and attaching them there, thereby increasing the depth and breadth of the technology concept system.
[0057] The following describes, with reference to the accompanying drawings, a method and apparatus for knowledge fusion of a scientific and technological concept system according to embodiments of this application.
[0058] Figure 1 This is a flowchart illustrating a method for knowledge fusion of a scientific and technological concept system provided in Embodiment 1 of this application.
[0059] like Figure 1 As shown, the method for knowledge integration of a scientific and technological concept system includes the following steps:
[0060] Step 101: Obtain multiple technology concept systems and divide them into a target technology concept system and multiple technology concept systems to be integrated. The structure of the technology concept system is a tree-like hierarchical structure composed of different nodes.
[0061] Step 102: Identify the nodes of the target technology concept system and the technology concept system to be integrated by splitting rules to obtain the nodes to be split, split the concept names of the nodes to be split according to the manual annotation rules, and update the nodes of the target technology concept system and the technology concept system to be integrated according to the splitting results.
[0062] Step 103: Detect the nodes of the updated target technology concept system and the technology concept system to be merged through similarity calculation, obtain the nodes with concept names with the same meaning, and merge the updated target technology concept system and the technology concept system to be merged based on the nodes with concept names with the same meaning.
[0063] Step 104: Calculate the confidence of the nodes of the technology concept system to be merged and the parent nodes of the target technology concept system after merging, and merge them according to the confidence to obtain the merged technology concept system.
[0064] The method for knowledge fusion of scientific and technological concept systems in this application embodiment involves acquiring multiple scientific and technological concept systems, dividing them into a target scientific and technological concept system and multiple scientific and technological concept systems to be fused, wherein the structure of the scientific and technological concept system is a tree-like hierarchical structure composed of different nodes; identifying nodes to be split using splitting rules in the target scientific and technological concept system and the scientific and technological concept systems to be fused, splitting the concept names of the nodes to be split according to manual annotation rules, and updating the nodes of the target scientific and technological concept system and the scientific and technological concept systems to be fused based on the splitting results; detecting the nodes of the updated target scientific and technological concept system and the scientific and technological concept systems to be fused through similarity calculation, obtaining nodes with concept names with the same meaning, and fusion of the updated target scientific and technological concept system and the scientific and technological concept systems to be fused based on the nodes with concept names with the same meaning; calculating the confidence of the nodes of the scientific and technological concept systems to be fused and the parent nodes of the merged target scientific and technological concept system, and fusion based on the confidence, to obtain the merged scientific and technological concept system. Therefore, it can address the technical problems of existing methods having narrow scope and inaccurate classification. By integrating classification data of different granularities and fields to construct a rich and high-quality tree structure, it can improve the accuracy of scientific and technological concept classification while integrating multiple data sources.
[0065] To ensure the informativeness and domain relevance of the concepts, this application selects multiple publicly available science and technology classifications for constructing the concept system. While concept tree construction, through clustering, can distinguish certain hierarchical relationships between concepts, different classifications may occur due to the selection of model parameters, requiring continuous adjustments to obtain a better concept system. Considering that the science and technology system should encompass different expressions of concepts and cover a broad range of disciplines, this application employs an unsupervised approach to align and complete existing concept systems.
[0066] like Figure 2 As shown, the knowledge fusion algorithm for the science and technology concept system first inputs a multi-source science and technology concept system; then it standardizes the category names, that is, it splits and corrects the input raw data; then it aligns the classification concepts, that is, it finds the correspondence between concepts in different classification systems to obtain concepts with the same meaning; finally, it completes the classification system, that is, it finds the most likely parent node of the classification node in the science and technology system to be fused in the constructed structure and attaches it to it to obtain the fused science and technology concept system.
[0067] A system of scientific and technological concepts categorizes real-world scientific systems according to their inherent connections based on certain principles, and expresses them in a logically consistent manner. A partial structure of the system of scientific and technological concepts is as follows: Figure 3 As shown, it is a tree-like hierarchical structure composed of different nodes, each node having zero or more child nodes.
[0068] The multi-source technology concept system obtained in this application consists of multiple technology concept systems with different granularities and fields.
[0069] In this embodiment, the acquired multiple technology concept systems are divided into a target technology concept system and multiple technology concept systems to be integrated. The technology concept system with the most nodes and the most authoritative source among the multiple technology concept systems is designated as the target technology concept system (an existing concept classification system), while the other technology concept systems are designated as technology concept systems to be integrated.
[0070] Furthermore, in this embodiment of the application, the nodes of the target technology concept system and the technology concept system to be merged are identified by splitting rules to obtain the nodes to be split, the concept names of the nodes to be split are split according to manual annotation rules, and the nodes of the target technology concept system and the technology concept system to be merged are updated according to the splitting results, including:
[0071] Obtain the names of the concepts to be split and analyze them to obtain the splitting rules;
[0072] Based on the splitting rules, the nodes of the target technology concept system and the technology concept system to be integrated are identified through regular expression analysis to obtain the nodes to be split.
[0073] The concept names of the nodes to be split are split according to the manual annotation rules to obtain the splitting results, and the structure of the target technology concept system and the technology concept system to be integrated is adjusted according to the splitting results.
[0074] The classification system for scientific and technological concepts contains instances of concept compounding. For example, there might be a sub-tree for "Computer Graphics and Virtual Reality," such as... Figure 4 As shown, the concept of "Computer Graphics and Virtual Reality" can be broken down into "Computer Graphics" and "Virtual Reality".
[0075] In this embodiment, the characteristics of the concept to be split are first analyzed, and splitting rules are defined to detect the node to be split. Then, the concept name is split, completed and corrected by a combination of program judgment and manual inspection. Finally, the original nodes in the concept system are updated according to the splitting results and the ownership of the sub-concepts.
[0076] Standardizing concept names first requires detecting the concepts to be split. This application analyzes and statistically examines the characteristics of concepts that need to be split, collecting and organizing some characteristic characters, such as AND, OR, AND, and punctuation marks such as commas and pauses. Then, regular expression analysis is used to identify concepts that may need to be split. After detecting the concepts to be split according to custom rules, the splitting results need to be manually completed and corrected. The rules for manual annotation include: high-level concepts are not split; if sub-concepts cannot be separated after splitting, they are not split; for semantically similar concepts after splitting, the common prefix is retained; and concepts that have lost suffix information after splitting are supplemented.
[0077] Based on the manually labeled splitting results, the original structure of the science and technology concept system is adjusted. The specific update process includes retrieving the nodes to be split, deleting the nodes and their edges, creating new concepts after splitting, connecting the new concepts to their parent nodes, and connecting the new concepts to their child nodes.
[0078] Furthermore, in this embodiment of the application, similarity calculation is used to detect nodes in the updated target technology concept system and the technology concept system to be merged, to obtain nodes with concept names having the same meaning, including:
[0079] Calculate the character similarity, vector similarity, and structural similarity between nodes in the updated target technology concept system and the technology concept system to be merged;
[0080] By linearly combining character similarity, vector similarity, and structural similarity, the similarity between nodes in the updated target technology concept system and the technology concept system to be merged is obtained. Nodes with similarity greater than a first preset threshold are regarded as nodes with concept names having the same meaning.
[0081] This application obtains conceptual representations pointing to the same meaning through similarity calculation. The similarity calculation compares the similarity of strings, semantics, and structure between conceptual nodes in different classification systems, selecting the node with the highest score as the alignment result. This application can also use rule matching to obtain conceptual representations pointing to the same meaning. Rule matching retrieves aligned conceptual names from the concept encyclopedia description using predefined rules. This application can then merge the updated target technology concept system and the technology concept system to be merged based on the aligned conceptual names.
[0082] Alignment classification refers to discovering the correspondence between concepts in different science and technology classification systems, given node e. i and e j This application can employ multiple strategies for concept alignment, as illustrated in the diagram of information association between nodes. Figure 5 As shown.
[0083] Based on the structural similarities between the alignment concepts, this application first calculates the node e to be aligned. j With e i Father e parent Brother e sibling and child node e child Vector similarity,
[0084] sim parent (e j ) = sim(e j ,e parent )
[0085] sim siblings (e j ) = [sim(e j ,e sibling1 ),…,sim(e j ,e siblingN )]
[0086] sim children (e j ) = [sim(e j ,e child1 ),…,sim(e j ,e childM )]
[0087] Where N is the number of sibling nodes and M is the number of child nodes. Similarly, the number of child nodes can be calculated for node e. i The vector similarity of the parent node sim parent (e i ), the vector similarity sim of sibling nodes siblings (e i The vector similarity between the child nodes and the vector similarity sim children (e i ).
[0088] The alignment classification concept in this application can use three features: string similarity, vector similarity, and structural similarity. Among them, the string similarity score... string Vector similarity score vector and structural similarity score structure The calculation formula is shown below.
[0089] score string =Levenshtein(e i ,e j )
[0090] score vector =sim(e i ,e j )
[0091] score structure =3-abs(sim parent (e i )-sim parent (e j ))
[0092] -D KL (sim siblings (e i )||sim siblings (e j ))
[0093] -D KL (sim children (e i )||sim children (e j ))
[0094] Specifically, string similarity is calculated using Levenshtein distance, vector similarity is measured using cosine distance to measure the degree of difference between distributed vectors, and structural similarity is evaluated using absolute value and KL (Kullback-Leibler) divergence to assess the similarity of the probability distributions of parent nodes, sibling nodes, and child nodes between concept nodes.
[0095] This application ultimately performs a linear combination of the scores from different strategies and smooths the result using the sigmoid function, with α being a preset parameter. The node to be aligned, e... j With e i similarity score (e) i ,e j ) is represented as:
[0096] score(e i ,e j ) = score string ×sigmod(α-5)
[0097] + vector ×sigmoid(α)
[0098] + structure ×sigmod(α+5)
[0099] Furthermore, in this embodiment of the application, before calculating the confidence scores of the nodes of the technology concept system to be fused and the parent nodes of the fused target technology concept system, and performing fusion based on the confidence scores to obtain the fused technology concept system, the following steps are included:
[0100] The concept-to-relationship prediction model uses semantic vector representations for the nodes of the technology concept system to be integrated and the target technology concept system after integration.
[0101] The nodes of the merged target technology concept system and the technology concept system to be merged are retrieved for higher-level and lower-level information respectively, and information fragments of the nodes are constructed.
[0102] This application first utilizes representation learning on the existing concept classification system and represents nodes as fixed-dimensional semantic vectors to facilitate subsequent processing. Then, it retrieves superior and subordinate information of nodes in the constructed concept system and the concept to be merged to construct information fragments. Finally, it attaches the node to the node to be merged by finding the most likely parent node.
[0103] A scientific and technological concept system is a tree-like structure composed of high-level and low-level concepts, which can be defined as T=(D,R), where D=v1,v2,…,v |D|} represents the set of concept direction nodes, R = <v i ,v j >|v i ,v j ∈D} represents the set of edges between nodes in the concept direction.
[0104] This application learns the semantic representation of nodes by predicting whether a hierarchical relationship exists between concept pairs. The main structure of the concept pair relationship prediction model is as follows: Figure 6 As shown, it mainly includes two modules: a concept encoding module, in which the concept pair relationship prediction model uses OAG-BERT to learn the association features between entity pairs; and a concept pair relationship prediction module, in which the concept pair relationship prediction model uses the output of the OAG-BERT model to predict the relationship between concepts and obtain the corresponding probability distribution.
[0105] Assume concept node v i ={Tok1,Tok2,…TokN}, Concept node v j = {Tok1, Tok2, ..., TokM}, where Tok1 represents the first character in the concept node name, and N and M represent the node v. i and v j The number of characters. The concept pair relationship prediction model first uses OAG-BERT to quantize concept characters into low-dimensional dense semantic vectors, and then sends them to the classification layer to predict whether there is a relationship between concept entities. The classification layer mainly consists of two fully connected layers and one Softmax layer.
[0106] This application utilizes existing hierarchical concept relationships within the concept system to construct training positive examples, and then randomly replaces concept nodes in the positive examples to construct negative samples. To effectively prevent overfitting, an L2 regularization term is added during the training of the concept pair relationship prediction model, and the training objective is to minimize the cross-entropy loss function.
[0107]
[0108] Where n is the number of training samples, y ij It is a real label between concepts, p ij It is the probability of whether or not there is a relationship between concept pairs, λ‖θ‖ 2 It is a regularization term.
[0109] To preserve the structural information of concept nodes to the greatest extent possible, this application segments the nodes in the graph into fragments of a certain length. Given a concept node v i The corresponding information fragment is in For v i The parent node, Indicates v i The first child node, H represents v i The total number of subordinate concepts (0 <= H < 50).
[0110] In finding the most likely parent node of the node to be merged, this application considers the node itself, as well as its related superior and subordinate concepts, within the established conceptual framework. For example... Figure 7 As shown, the left side represents the established scientific and technological concept system, and the right side represents the concepts to be integrated. When searching for the most likely parent node of the concept node "Empirical Bayesian Analysis" such as "Probability and Statistics", this application will simultaneously calculate the consistency between "Empirical Bayesian Analysis" and "Novel Computation and Application", as well as between "Empirical Bayesian Analysis" and "Distribution Function", "Probability Representation", "Multivariate Statistics", etc.
[0111] Further, in this embodiment of the application, the confidence scores of the nodes of the technology concept system to be fused and the parent nodes of the fused target technology concept system are calculated, and fusion is performed based on the confidence scores to obtain the fused technology concept system, including:
[0112] The hierarchical relationship between the nodes of the technology concept system to be merged and the parent node of the target technology concept system after merging is predicted by the concept pair relationship prediction model, and the corresponding probability values are obtained.
[0113] The parent node of the fused target technology concept system with a probability value greater than the second preset threshold is selected as the candidate node of the node of the technology concept system to be fused, and the confidence between the node of the technology concept system to be fused and the candidate node is calculated based on the information fragment of the node.
[0114] Select the candidate node with the highest confidence level and greater than the third preset threshold to attach the node of the technology concept system to be integrated, and obtain the integrated technology concept system.
[0115] In this embodiment, a concept-pair relationship prediction model is used to predict the hierarchical relationship between nodes of the technology concept system to be merged and the parent node of the target technology concept system after merging, and corresponding probability values are obtained. Candidate nodes with probability values greater than a second preset threshold are obtained, and the confidence level between corresponding nodes is calculated based on the information fragments of the candidate nodes. For example, if the probability prediction value of "artificial intelligence" and "knowledge graph" is greater than 0.5, it is considered that there is a hierarchical relationship; otherwise, it is not. If the probability prediction value of "artificial intelligence" and "knowledge graph" is greater than the second preset threshold of 0.6, then the "artificial intelligence" node is selected as a candidate node for the "knowledge graph" node.
[0116] Furthermore, in this embodiment, the information fragment of a node includes the node, the node's parent node, and the node's child nodes. Calculating the confidence level between nodes and candidate nodes in the technology concept system to be fused based on the node's information fragment includes:
[0117] A score matrix is constructed based on the information fragments of the nodes and candidate nodes in the technology concept system to be integrated.
[0118] Calculate the fraction matrix according to the first calculation formula, and then adjust the fraction matrix.
[0119] The average score of the adjusted score matrix is used as the confidence level between the nodes and candidate nodes of the technology concept system to be integrated.
[0120] The first calculation formula is expressed as follows:
[0121]
[0122] Among them, s k,b This represents the value at the k-th row and b-th column in the fraction matrix, match(s) k ,s b ) represents the string similarity between nodes, sim(s) k ,s b ) represents the similarity of concept names between nodes, count(s) k ,s b ) represents the conceptual name of the node. k and s b The number of identical characters between two lines.
[0123] Assume the node to be merged is e i In the constructed conceptual framework, a certain node is e. j The corresponding node information fragment is and This application first calculates as follows: Figure 8 The fractional matrix shown is denoted as . Among them, s k,bThe calculation method for ∈S is as follows
[0124]
[0125] Where, match(s) k ,s b ) calculates similarity by comparing string sequences, sim(s k ,s b This measures the topic consistency between concept names by calculating the cosine similarity between vectors, using a fine-tuned OAG-BERT quantization method, count(s) k ,s b ) represents the statistical concept name s k and s b The number of identical characters between two lines.
[0126] After calculating the score matrix, this application adjusts the score matrix as follows: the more similar the parent of the node to be merged is to the current node in the already constructed system, the more likely they are to belong to a hierarchical relationship. Therefore, the score at position k=0, b=1 is increased. The calculation method is as follows:
[0127] s k,b =s k,b +s k,b ×s k,b
[0128] The more similar the parent concept of the node to be merged is to the node in the already constructed system, the higher the level of the node to be merged and the current node in the already constructed system. Therefore, it is less suitable to be a sub-concept of the current node, and the score for positions k=1 and b=0 is reduced. The calculation method is as follows:
[0129] s k,b =s k,b -s k,b ×s k,b
[0130] After the above adjustments, this application uses the average of the sums of all scores in the matrix as node e. i Mount to node e j The confidence level is calculated using the following formula:
[0131]
[0132] This application calculates the scores of the node to be merged and all parent nodes in the constructed conceptual system, and selects the node with the highest score that is greater than the third preset threshold as the most likely parent node.
[0133] Figure 9 This is a schematic diagram of a device for knowledge fusion of a scientific and technological concept system provided in Embodiment 2 of this application.
[0134] like Figure 9 As shown, the device for knowledge integration of a technological concept system includes:
[0135] The acquisition module 10 is used to acquire multiple technology concept systems and divide the multiple technology concept systems to obtain a target technology concept system and multiple technology concept systems to be integrated. The structure of the technology concept system is a tree-like hierarchical structure composed of different nodes.
[0136] The splitting module 20 is used to identify the nodes of the target technology concept system and the technology concept system to be integrated through splitting rules to obtain the nodes to be split, split the concept names of the nodes to be split according to the manual annotation rules, and update the nodes of the target technology concept system and the technology concept system to be integrated according to the splitting results.
[0137] The first fusion module 30 is used to detect the nodes of the updated target technology concept system and the technology concept system to be fused through similarity calculation, obtain the nodes with concept names with the same meaning, and fuse the updated target technology concept system and the technology concept system to be fused based on the nodes with concept names with the same meaning.
[0138] The second fusion module 40 is used to calculate the confidence of the nodes of the technology concept system to be fused and the parent nodes of the target technology concept system after fusion, and to perform fusion based on the confidence to obtain the fused technology concept system.
[0139] The apparatus for knowledge fusion of scientific and technological concept systems according to embodiments of this application includes an acquisition module for acquiring multiple scientific and technological concept systems and dividing the multiple scientific and technological concept systems to obtain a target scientific and technological concept system and multiple scientific and technological concept systems to be fused, wherein the structure of the scientific and technological concept system is a tree-like hierarchical structure composed of different nodes; a splitting module for identifying nodes to be split by using splitting rules in the target scientific and technological concept system and the scientific and technological concept systems to be fused, splitting the concept names of the nodes to be split according to manual annotation rules, and updating the nodes of the target scientific and technological concept system and the scientific and technological concept systems to be fused according to the splitting results; a first fusion module for detecting the nodes of the updated target scientific and technological concept system and the scientific and technological concept systems to be fused through similarity calculation to obtain nodes with concept names with the same meaning, and fusion the updated target scientific and technological concept system and the scientific and technological concept systems to be fused according to the nodes with concept names with the same meaning; and a second fusion module for calculating the confidence of the nodes of the scientific and technological concept systems to be fused and the parent nodes of the fused target scientific and technological concept system, and fusion according to the confidence to obtain the fused scientific and technological concept system. Therefore, it can address the technical problems of existing methods having narrow scope and inaccurate classification. By integrating classification data of different granularities and fields to construct a rich and high-quality tree structure, it can improve the accuracy of scientific and technological concept classification while integrating multiple data sources.
[0140] Furthermore, in this embodiment of the application, the split module is specifically used for:
[0141] Obtain the names of the concepts to be split and analyze them to obtain the splitting rules;
[0142] Based on the splitting rules, the nodes of the target technology concept system and the technology concept system to be integrated are identified through regular expression analysis to obtain the nodes to be split.
[0143] The concept names of the nodes to be split are split according to the manual annotation rules to obtain the splitting results, and the structure of the target technology concept system and the technology concept system to be integrated is adjusted according to the splitting results.
[0144] To implement the above embodiments, this application also proposes a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the method for knowledge fusion of the scientific and technological concept system described in the above embodiments.
[0145] To implement the above embodiments, this application also proposes a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the method for knowledge fusion of the scientific and technological concept system described in the above embodiments.
[0146] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0147] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0148] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.
[0149] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.
[0150] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0151] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.
[0152] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0153] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.
Claims
1. A method for knowledge integration of a scientific and technological concept system, characterized in that, Includes the following steps: Multiple technology concept systems are acquired and divided to obtain a target technology concept system and multiple technology concept systems to be integrated. The structure of the technology concept system is a tree-like hierarchical structure composed of different nodes. The nodes of the target technology concept system and the technology concept system to be integrated are identified by splitting rules to obtain the nodes to be split. The concept names of the nodes to be split are split according to manual annotation rules, and the nodes of the target technology concept system and the technology concept system to be integrated are updated according to the splitting results. The updated target technology concept system and the technology concept system to be merged are detected by similarity calculation to obtain nodes with concept names that have the same meaning, and the updated target technology concept system and the technology concept system to be merged are merged according to the nodes with concept names that have the same meaning. The confidence scores of the nodes of the technology concept system to be merged and the parent nodes of the merged target technology concept system are calculated, and fusion is performed based on the confidence scores to obtain the merged technology concept system. Specifically: the hierarchical relationship between the nodes of the technology concept system to be merged and the parent nodes of the merged target technology concept system is predicted using a concept pair relationship prediction model, and corresponding probability values are obtained; the parent nodes of the merged target technology concept system whose probability values are greater than a second preset threshold are selected as candidate nodes for the nodes of the technology concept system to be merged, and the confidence scores between the nodes of the technology concept system to be merged and the candidate nodes are calculated based on the information fragments of the nodes; the candidate nodes with the highest confidence scores and greater than a third preset threshold are selected to attach nodes of the technology concept system to be merged, thus obtaining the merged technology concept system.
2. The method as described in claim 1, characterized in that, The process of identifying nodes in the target technology concept system and the technology concept system to be merged using splitting rules to obtain nodes to be split, splitting the concept names of the nodes to be split according to manual annotation rules, and updating the nodes in the target technology concept system and the technology concept system to be merged based on the splitting results includes: Obtain the names of the concepts to be split and analyze them to obtain the splitting rules; According to the splitting rules, the nodes of the target technology concept system and the technology concept system to be integrated are identified by regular expression analysis to obtain the nodes to be split; The concept names of the nodes to be split are split according to the manual annotation rules to obtain the splitting results, and the structure of the target technology concept system and the technology concept system to be integrated is adjusted according to the splitting results.
3. The method as described in claim 1, characterized in that, The step of detecting nodes in the updated target technology concept system and the technology concept system to be merged through similarity calculation to obtain nodes with concept names having the same meaning includes: Calculate the character similarity, vector similarity, and structural similarity between nodes in the updated target technology concept system and the technology concept system to be merged; The character similarity, vector similarity, and structural similarity are linearly combined to obtain the similarity between nodes of the updated target technology concept system and the technology concept system to be merged. Nodes with similarity greater than a first preset threshold are regarded as nodes with concept names having the same meaning.
4. The method as described in claim 1, characterized in that, Before calculating the confidence scores of the nodes of the technology concept system to be merged and the parent nodes of the merged target technology concept system, and performing the fusion based on the confidence scores to obtain the merged technology concept system, the following steps are included: The nodes of the technology concept system to be integrated and the integrated target technology concept system are represented by semantic vectors using a concept pair relationship prediction model. The nodes of the merged target technology concept system and the technology concept system to be merged are retrieved for higher-level and lower-level information respectively, and information fragments of the nodes are constructed.
5. The method as described in claim 4, characterized in that, The information fragment of the node includes the node, the node's parent node, and the node's child nodes. The step of calculating the confidence level between the nodes of the technology concept system to be integrated and the candidate nodes based on the node's information fragment includes: Based on the information fragments of the nodes and candidate nodes in the technology concept system to be integrated, a score matrix is constructed; the score matrix is calculated according to the first calculation formula, and the score matrix is adjusted. The average score of the adjusted score matrix is used as the confidence level between the nodes of the technology concept system to be integrated and the candidate nodes. The first calculation formula is expressed as follows: in, This represents the value at the k-th row and b-th column in the fractional matrix. This indicates the string similarity between nodes. This indicates the similarity of concept names between nodes. The conceptual name of a node and The number of identical characters between two lines.
6. A device for knowledge integration of a scientific and technological concept system, characterized in that, include: The acquisition module is used to acquire multiple technology concept systems and divide the multiple technology concept systems to obtain a target technology concept system and multiple technology concept systems to be integrated. The structure of the technology concept system is a tree-like hierarchical structure composed of different nodes. The splitting module is used to identify nodes of the target technology concept system and the technology concept system to be integrated through splitting rules to obtain nodes to be split, split the concept names of the nodes to be split according to manual annotation rules, and update the nodes of the target technology concept system and the technology concept system to be integrated according to the splitting results. The first fusion module is used to detect the nodes of the updated target technology concept system and the technology concept system to be fused through similarity calculation, obtain the nodes with concept names with the same meaning, and fuse the updated target technology concept system and the technology concept system to be fused according to the nodes with concept names with the same meaning. The second fusion module is used to calculate the confidence scores of the nodes of the technology concept system to be fused and the parent nodes of the fused target technology concept system, and to perform fusion based on the confidence scores to obtain the fused technology concept system. Specifically: a concept pair relationship prediction model is used to predict the hierarchical relationship between the nodes of the technology concept system to be fused and the parent nodes of the fused target technology concept system, and corresponding probability values are obtained; parent nodes of the fused target technology concept system whose probability values are greater than a second preset threshold are selected as candidate nodes of the nodes of the technology concept system to be fused, and the confidence scores between the nodes of the technology concept system to be fused and the candidate nodes are calculated based on the information fragments of the nodes; the candidate node with the highest confidence score and greater than a third preset threshold is selected to attach the nodes of the technology concept system to be fused, thus obtaining the fused technology concept system.
7. The apparatus as claimed in claim 6, characterized in that, Split into modules, specifically for: Obtain the names of the concepts to be split and analyze them to obtain the splitting rules; According to the splitting rules, the nodes of the target technology concept system and the technology concept system to be integrated are identified by regular expression analysis to obtain the nodes to be split; The concept names of the nodes to be split are split according to the manual annotation rules to obtain the splitting results, and the structure of the target technology concept system and the technology concept system to be integrated is adjusted according to the splitting results.
8. A computer device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it implements the method as described in any one of claims 1-5.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-5.
Citation Information
Patent Citations
A quick knowledge comparison method and system based on a knowledge graph
CN109885693A
Heterogeneous knowledge resource intelligent fusion method
CN115391550A