Multi-source entity attribute relationship mining method based on knowledge graph

Through the multi-source entity attribute relationship mining method based on knowledge graph, the problems of entity attribute conflict and low truth value inference in multi-source data are solved, and higher data integration accuracy and reliability are achieved.

CN120069038APending Publication Date: 2025-05-30SOUTHEAST UNIV

Patent Information

Application Number
CN202510225695.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the problems of entity attribute conflicts, strong relationship implicitness, and low truth value inference efficiency in multi-source data, especially when dealing with complex correlation scenarios of multi-source heterogeneous data.

Method used

The multi-source entity attribute relationship mining method is adopted based on the knowledge graph, and the attribute relationship mining and truth value discovery are systematically realized through semantic correlation modeling and dynamic alignment technology, combined with cluster analysis and statistical optimization.

Benefits of technology

It improves the accuracy and reliability of multi-source data integration, can effectively identify potential relationships and commonalities between attributes, dynamically adjust data source weights, and improves the reliability of true value discovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069038A_ABST
    Figure CN120069038A_ABST
Patent Text Reader

Abstract

The invention provides a multi-source entity attribute relationship mining method based on a knowledge graph, and efficient mining and integration of multi-source entity attribute relationships are realized by combining the knowledge graph, community detection, semantic embedding and clustering technologies. The method comprises the following four main stages: multi-source entity attribute extraction: extracting entities and attribute information thereof from a domain text, obtaining structured attribute data through BIO labeling and feature extraction, and multi-entity alignment: vectorizing the multi-source entities by adopting a knowledge graph embedding method to realize unified identification of the same entities in different data sources, multi-source entity attribute relationship mining: constructing an intra-cluster attribute relationship knowledge graph; and truth value discovery: carrying out truth value discovery through a community detection algorithm and an Expectation-Maximization algorithm. According to the method, the attribute relationship of the multi-source entity can be accurately identified in a large-scale and complex data environment, the accuracy of multi-source data integration and the efficiency of truth value discovery are effectively improved, and the method is suitable for application scenes such as intelligent question answering and information extraction.
Need to check novelty before this filing date? Find Prior Art

Claims

1. A method for mining relations between multi-source entity attributes based on knowledge graph, characterized in that: The method comprises the following steps: A. Extracting multi-source entity attributes: First, determine the research field and collect a large amount of field-related data as source text from public data sources, domain databases, or crawled datasets. Given an entity and a pre-defined attribute name, extract the attribute value from the original text. B. In the multi-source entity alignment phase, the same entity from different data sources is identified and merged, and an entity alignment method is established, which includes an embedding module, an interaction module, and an alignment module. The embedding module is responsible for embedding the knowledge graph through knowledge representation learning technology, and the interaction module is responsible for mapping the embedding space of different knowledge graphs to the same vector space according to the aligned entity pairs. The alignment module is responsible for obtaining the entity alignment result according to the distance or similarity between entities in the vector space. By reducing high-dimensional entities and relationships, the numerical representation of low-dimensional vectors is obtained, and then the semantic relationship between two objects is characterized by the distance between two low-dimensional dense vectors. C. In the phase of mining the relationship between multi-source entity attributes, entities from different sources are clustered, entities with similar attributes are grouped into the same cluster, and K-means algorithm is used to characterize entities based on their multi-dimensional attributes. Within each cluster, by analyzing the similarity and common attributes between entities, potential attribute relationships are extracted and a knowledge graph within the cluster is constructed. Nodes represent attributes within the cluster, and edges represent relationships between different attributes. D. Use the relationship between multi-source entity attributes for truth discovery. Through the community detection algorithm, identify the communities with close relationships between attributes in the graph. If some attributes appear isolated in the entire graph or have obvious conflicts with other attributes, then these attributes are considered wrong. Combined with the weight of the data source, the similarity of attribute values ​​and the relationship between sources, assign a credibility to each attribute value. Through the expectation-maximization algorithm, continuously update the reliability of the data source and the credibility of the attribute value to converge to the final truth value.

2. The method for mining relations between multi-source entity attributes based on knowledge graph according to claim 1 is characterized in that: The specific steps of extracting multi-source entity attributes are as follows: A1. Choose a field and collect a large amount of text data for that field; A2. Mark the text data into BIO markup format, where B represents the beginning of an entity or attribute value, I represents the inside of an entity, and O represents the non-entity part, thereby obtaining a training data set. Each sample is a word sequence of a sentence and its corresponding BIO label. A3. Extract features from the text that can help the model perform classification, including vocabulary features, part-of-speech features, word morphology features, and context features, and combine these features into a feature vector for each word for classification by the machine learning algorithm; A4. Train the conditional random field model (CRF) based on the training data set. Define the input as a feature vector sequence (in the form of a list, where each list item is the feature and label of a word), and the output as the corresponding BIO label sequence. CRF will learn the probability of labeling each word through the feature vector according to the sequence context, calculate the weight of each label, and optimize the label sequence in combination with the global path information. CRF is a probabilistic graph model that calculates the joint probability distribution of label sequences. The formula is: X: input sequence; Y: tag sequence; ·f k : characteristic function; ·w k : Feature weight; i: current location; ·y i :Current location label; ·y i-1 : Previous position label; Z(X): normalization factor, A5. Tune the CRF model by adjusting the regularization parameters c1 and c2 of the CRF to balance the complexity and fit of the model. Experimentally select different combinations of features to observe changes in model performance. Use K-fold cross validation to evaluate the generalization ability of the model to ensure that the performance of the model on different data sets is roughly consistent. A6. Input new text, extract its features, and use the trained CRF model for prediction; A7. Process the output BIO tags and extract the corresponding attribute values; A8. When there are multiple candidates for the same attribute, resolve conflicts based on context information or confidence scores and retain relevant attribute values; A9. Use the classification_report tool in sklearn to evaluate the performance of the CRF model through precision, recall, and F1 score. Based on the evaluation results, further optimize the model features and hyperparameters or perform more data annotation.

3. The method for mining relations between multi-source entity attributes based on knowledge graph according to claim 1 is characterized in that: The specific steps of the multi-source entity alignment stage are as follows: B1. Use the translation model to embed the knowledge graph. Given a training set S consisting of a triple (h, l, t) of two entities h, t∈E (entity set) and a relation l∈L (relation set), construct the objective function L(h, l, t) = ||h+lt||, where |||| represents the distance between vectors. Use L1 or L2 distance, use random initialization to initialize all entities and relations into low-dimensional vectors, use stochastic gradient descent (SGD) to minimize the objective function, and obtain the final embedding vector after multiple rounds of iterations. B2. Based on the aligned entities and the mapping function, the Procrustes analysis based on the aligned entity pairs is used to map the embedding spaces of different knowledge graphs into the same space vector; A, B: Embedding vectors of two knowledge graphs; W: mapping matrix, B3. In a unified embedding space, calculate the cosine similarity between entity vectors in different knowledge graphs, set the similarity threshold based on the calculation results, calculate the cosine similarity for all entities, and generate the final entity alignment result.

4. The method for mining relations between multi-source entity attributes based on knowledge graph according to claim 1, characterized in that: The specific steps of the stage of constructing a knowledge graph to mine the relationship between multiple entity attributes are as follows: C1. Clean and normalize multi-source entity data to ensure that all attribute values ​​are comparable. In the data cleaning stage, remove duplicate entity data, check and delete records with many missing values ​​or abnormalities, and correct spelling errors and irregular naming in attribute values. In the data normalization stage, unify the names of multi-source attributes that represent the same attribute, unify the format of numerical data (such as date, price, etc.), and standardize the data when necessary to eliminate the impact of dimension. Among them, x ′ is the normalized x, C2. Vectorize the multidimensional attributes of each entity. For categorical attributes, use one-hot encoding to convert them into numerical vectors. For numerical attributes, directly use the normalized values. Then, the attributes of each entity are combined into a multidimensional vector, where x i is the value of the i-th attribute; x=[x1,x2,...,x n ] C3. Use the silhouette coefficient method to determine the optimal number of clusters K. The silhouette coefficient value range is [-1,1]. The larger the value, the better the clustering effect. a(i) is the average distance between sample i and other samples in the same cluster, b(i) is the average distance between sample i and the nearest neighbor cluster, and s(i) measures the relative difference between the compactness of sample i within its cluster and its separation from neighboring clusters. The number of clusters with the highest silhouette coefficient is selected as K. Based on the multi-dimensional attributes of the entity, the K-means clustering algorithm is applied to classify entities with similar attributes into the same cluster. The objective function is as follows, where C i is a cluster, μ i is the cluster center; First, randomly select K initial cluster centers μ i ; For each sample x j , calculate its Euclidean distance to each cluster center and assign it to the nearest cluster; Update the center of each cluster to the mean of all samples in the cluster; Repeat the above steps until the cluster assignment does not change or the objective function j converges. C4. Build a graph for each attribute in a cluster. Nodes represent attributes, and edges represent the relationship between attributes. The connected edges in the graph represent the correlation between attributes. In this paper, the frequency of co-occurrence is used to measure the correlation. The higher the frequency of attribute pairs, the closer the relationship. Among them, A and B are the two attributes of the relationship being measured, count(a∧b) refers to the number of entities that have both attribute a and attribute b in the dataset, and total entities is the total number of all entities in the dataset.

5. The method for mining relations between multi-source entity attributes based on knowledge graph according to claim 1, characterized in that: The specific steps of the truth value discovery phase using the relationship between multi-source entity attributes are as follows: D1. Use community detection algorithms to perform community detection on the knowledge graph and identify multiple communities based on the attribute relationships in the graph; D2. Assign weights according to the credibility of each data source and evaluate the credibility of each attribute through the similarity of the edges in the attribute relationship graph. Isolated or conflicting attribute values ​​can be regarded as potential erroneous data and their credibility should be lower. D3. In each iteration, the weight of the data source and the similarity of the attribute value are combined to update the credibility of each attribute value according to the credibility of the attribute value and the current true value, and the weight of each data source is updated according to the credibility of the attribute value. The closer the true value provided by the data source is to the current true value, the higher the weight is. Repeat the iteration until the credibility of the attribute value and the data source converge to a stable value. D4. According to the final credibility of the attribute value, select the attribute value with higher credibility as the final true value.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, it implements the method for mining the relationship between multi-source entity attributes based on the knowledge graph as described in any one of claims 1 to 5 above.

7. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the computer instruction is executed by a processor, the method for mining the relationship between multi-source entity attributes based on a knowledge graph as described in any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Truth value discovery method and system for knowledge verification

    CN112651505A

  • Scientific knowledge discovery method and system based on knowledge graph

    CN117786122A

  • Generating shadows for placement objects in depth estimation scene of two-dimensional image

    CN117830473A

  • Stem cell knowledge graph construction method and device, equipment and storage medium

    CN118886495A

  • True value mining method for multi-source unstructured text service in meta-universe crowdsourcing environment

    CN119203046A

Cited By

  • Multi-source heterogeneous data alignment method and system

    CN120561612A

  • Navigation notice text processing method and device based on semantic enhancement and medium

    CN120996050A