A Knowledge Fusion Method and System for Multi-Source Heterogeneous Multi-Modal Data

By preprocessing, feature extraction and modal alignment of multi-source heterogeneous multimodal data, fusion representation learning is performed in combination with graph neural networks, and the results are integrated into the knowledge graph, the problem of inefficient multi-source heterogeneous multimodal data fusion in the existing technology is solved, and efficient data fusion and knowledge structure construction are achieved.

CN118690838BActive Publication Date: 2025-05-30BEIJING SCI & TECH PATENT OFFICE

Patent Information

Application Number
CN202410735595.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-07
Publication Date
2025-05-30
Estimated Expiration
2044-06-07

AI Technical Summary

Technical Problem

The existing knowledge fusion method for multi-source heterogeneous multimodal data is inefficient and has high computational cost, making it difficult to effectively process and fuse data from different modalities in a unified manner.

Method used

Through data preprocessing and standardization, feature extraction and vector representation, feature dimensionality reduction, modal alignment, fusion representation learning using graph neural network, and finally the fused feature representation is integrated into the knowledge graph.

Benefits of technology

It improves the comparability and processing ability of data, reduces the computational complexity and cost, and realizes the effective fusion of multimodal data and the construction of knowledge structure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118690838B_ABST
    Figure CN118690838B_ABST
Patent Text Reader

Abstract

The present invention discloses a knowledge fusion method for multi-source heterogeneous multi-modal data, including: Step 1: Preprocess and standardize the data from different sources to ensure the quality and consistency of the data; Step 2: Extract features from different modal data to generate a unified vector representation; Step 3: Reduce the high-dimensional feature vectors to a low-dimensional space to reduce the computational overhead; Step 4: Align the data of different modalities so that they can be represented and processed in the same semantic space; Step 5: Use a graph neural network for fusion representation learning to generate a comprehensive feature representation; Step 6: Integrate the fused feature representation into the knowledge graph to form a unified knowledge structure. The present invention has significant advantages in improving the efficiency of data processing and fusion and reducing the computational cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data fusion, and particularly relates to a knowledge fusion method and system for multi-source heterogeneous multi-modal data. Background Art

[0002] The knowledge fusion of multi-source heterogeneous multi-modal data is a complex and important research topic, which involves extracting and integrating information from data of different sources, formats, and modalities to construct a more comprehensive and accurate knowledge graph or intelligent system. In the real world, data often comes from different sources and modalities, including various forms such as text, images, audio, etc. There are rich associations and information among these data, but due to different modalities and representation methods, unified processing and fusion are required. However, existing fusion methods generally have problems of low efficiency and high computational costs. Summary of the Invention

[0003] To solve the above problems, the present invention provides a knowledge fusion method and system for multi-source heterogeneous multi-modal data.

[0004] To achieve the above object, the technical solutions adopted by the present invention are as follows:

[0005] On the one hand, the present application discloses a knowledge fusion method for multi-source heterogeneous multi-modal data, including the following steps:

[0006] Step 1: Preprocess and standardize the data from different sources to ensure the quality and consistency of the data;

[0007] Step 2: Extract features from data of different modalities to generate a unified vector representation;

[0008] Step 3: Reduce the high-dimensional feature vectors to a low-dimensional space to reduce the computational overhead;

[0009] Step 4: Align the data of different modalities so that they are represented and processed in the same semantic space;

[0010] Step 5: Use a graph neural network for fusion representation learning to generate a comprehensive feature representation;

[0011] Step 6: Integrate the fused feature representation into the knowledge graph to form a unified knowledge structure.

[0012] Furthermore: The said Step 1 includes:

[0013] The preprocessing includes: removing noise, spelling correction, filling missing values, and uniformly converting data in different formats into a structured format;

[0014] The standardization includes: normalizing numerical data to make its numerical range consistent.

[0015] Further: Step 2 includes:

[0016] Decompose the text into a sequence of words {w 1 , w 2 ,..., w n};

[0017] Map each word w i to a vector v i ;

[0018] Calculate the sentence vector by the averaging method

[0019] Normalize the sentence vector

[0020] The generated unified vector is denoted as v final .

[0021] Where n is the number of words in the sentence, v sentence is the vector representation of the sentence, μ is the mean of the vector, σ is the standard deviation of the vector, and v normalized is the vector representation after normalization.

[0022] Further: Step 3 includes:

[0023] Calculate the similarity matrix between high-dimensional feature vectors using the following formula:

[0024]

[0025] Where p ij is the conditional probability between samples x i and x j , and σ i is the distance to its nearest neighbor of sample x i .

[0026] Optimize the mapping positions of the samples by minimizing the KL divergence between the similarities between samples in the high-dimensional space and the similarities between samples in the low-dimensional space:

[0027]

[0028] Where q ij is the conditional probability between samples y i and y j ;

[0029] Calculate the KL divergence between the conditional probability distribution P in the high-dimensional space and the conditional probability distribution Q in the low-dimensional space:

[0030]

[0031] Use the position of the sample in the low-dimensional space as the feature representation after dimensionality reduction.

[0032] Further: Step 4 includes:

[0033] Build an embedding space into which the feature vectors of each modality can be mapped:

[0034] z = [z 1 , z 2 ,..., z m

[0035] where z i represents the i-th dimensional vector in the shared space;

[0036] Define a loss function to minimize the differences between different modalities:

[0037]

[0038] where x i and y i represent the feature vectors of different modalities respectively, and l(·) is a loss function used to measure the differences between two vectors;

[0039] Through optimization algorithms such as gradient descent, minimize the loss function and update the model parameters so that the feature vectors of different modalities are gradually aligned in the shared space:

[0040] θ * = arg min θ L(θ)

[0041] where θ represents the model parameters;

[0042] The obtained model parameters θ * are used to map the feature vectors of different modalities into the shared space to achieve modality alignment:

[0043] z i = f(x i ; θ * )

[0044] where f(·) is a mapping function that maps the feature vector x i to the vector z i in the shared space.

[0045] Further: Step 5 includes:

[0046] Construct a heterogeneous graph according to the relationships between the data. In the heterogeneous graph, the nodes represent different data entities, and the edges represent the relationships between different entities;

[0047] Initialize the features of each node in the graph, and use the feature representation after modality alignment as the initial feature of the node:

[0048]

[0049] where z i is the initial feature representation of the i-th node;

[0050] Perform message passing and aggregation on the node features through multiple graph convolutional layers to achieve fused representation learning:

[0051]

[0052] where represents the feature representation of the i-th node at the l-th layer, N(i) represents the set of neighbor nodes of node i, c ij is the number of edges between node i and node j, W (l) is the weight matrix of the l-th layer, and σ(i) is the activation function;

[0053] After passing through multiple graph convolutional and graph pooling layers, the obtained node feature representation is the fused comprehensive feature representation.

[0054] Furthermore, the step 6 includes:

[0055] Use NER technology to identify entities in the text data and link them to the corresponding nodes in the knowledge graph;

[0056] Use a relation extraction model to extract the relationships between entities from the text data;

[0057] Construct the identified entities and their relationships into the nodes and edges of the knowledge graph.

[0058] On the other hand, the present application discloses a knowledge fusion system for multi-source heterogeneous multi-modal data, including:

[0059] Data preprocessing and standardization module: Preprocess and standardize data from different sources to ensure data quality and consistency;

[0060] Feature extraction and vector representation module: Extract features from data of different modalities to generate a unified vector representation;

[0061] Feature dimensionality reduction module: Reduce the high-dimensional feature vectors to a low-dimensional space to reduce computational overhead;

[0062] Modality alignment module: Align data of different modalities so that they can be represented and processed in the same semantic space;

[0063] Fusion Representation Learning Module: Use graph neural networks for fusion representation learning to generate comprehensive feature representations;

[0064] Knowledge Graph Construction Module: Integrate the fused feature representations into a knowledge graph to form a unified knowledge structure.

[0065] Compared with the prior art, the technical progress achieved by the present invention is as follows:

[0066] The present invention combines various data modalities such as text, images, and audio, fully utilizes the rich information of different modality data, and can describe data features more comprehensively. Through steps such as data preprocessing, feature extraction, and vector representation, the present invention uniformly represents data from different sources as feature vectors, enabling data of different modalities to be represented and processed in a unified semantic space, improving the comparability and processability of data. The present invention uses multi-modal neural networks or graph neural networks for fusion representation learning, which can effectively fuse data of different modalities to generate comprehensive feature representations, helping to extract higher-level semantic information. The present invention integrates the fused feature representations into a knowledge graph to form a unified knowledge structure, which helps to better understand the entities and relationships in the data and provides support for knowledge reasoning and applications. The present invention can reduce the feature dimension and computational complexity, improve the efficiency of data processing and fusion, and reduce the computational cost through technical means such as feature dimensionality reduction, attention mechanism, and graph convolutional neural network.

[0067] In summary, the present invention has significant advantages in improving the efficiency of data processing and fusion and reducing the computational cost. By comprehensively utilizing multi-modal information, unified data representation, fusion representation learning, and knowledge graph construction, etc., it can more comprehensively understand and utilize multi-source heterogeneous data and provide better data support for various application scenarios. Brief Description of the Drawings

[0068] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention.

[0069] In the drawings:

[0070] Figure 1 is a flowchart of the present invention. Detailed Embodiments

[0071] The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present invention will be described below in conjunction with the drawings.

[0072] Embodiment 1

[0073] As Figure 1As shown, the present invention discloses a knowledge fusion method for multi-source heterogeneous multi-modal data, including the following steps:

[0074] Step 1: Preprocess and standardize the data from different sources to ensure the quality and consistency of the data.

[0075] 1.1 Data cleaning

[0076] Data cleaning is the first step of preprocessing, aiming to remove noise and errors and fill in missing values, including:

[0077] Removing noise: such as removing HTML tags, special characters, etc. from the text.

[0078] Spelling correction: Check and correct spelling mistakes in the text.

[0079] Filling in missing values: If there are missing parts in the text, methods such as interpolation or context prediction can be used to fill them in.

[0080] In this embodiment, it is completed using the following code:

[0081]

[0082]

[0083] 1.2 Format conversion

[0084] Format conversion is to uniformly convert data in different formats into a structured format for subsequent processing. For text data, it can be converted into a table format, with each row representing a record and the columns representing different features.

[0085] In this embodiment, it is completed using the following code:

[0086]

[0087]

[0088] 1.3 Standardization processing

[0089] Standardization processing is to normalize or standardize numerical data to make its numerical range consistent. For text data, word embedding is usually used to convert it into a vector representation and then perform standardization, including:

[0090] In this embodiment, GloVe is used for word embedding:

[0091]

[0092]

[0093] Through the above steps, the text data has been preprocessed and standardized, including data cleaning, format conversion, and numerical normalization, which provides clean and consistent input data for subsequent feature extraction and fusion.

[0094] Step 2: Extract features from data of different modalities to generate a unified vector representation.

[0095] In Step 1, the data has been preprocessed and standardized. Next, how to perform feature extraction and vector representation will be described in detail.

[0096] 2.1 Feature Extraction and Vector Representation of Text Data

[0097] The goal of text data feature extraction is to convert the original text into a vector representation with a fixed dimension for subsequent fusion and analysis. The following steps will detail how to achieve this goal.

[0098] Text tokenization: Decompose the text into a sequence of words. Let the text be text, and the result after tokenization be {w 1 , w 2 ,..., w n}.

[0099] Word vector mapping: Map each word to a vector. For each word w i , its word vector representation is v i :

[0100] Sentence vector representation: Combine the word vectors into a sentence vector using the average method:

[0101]

[0102] where n is the number of words in the sentence, and v sentence is the vector representation of the sentence.

[0103] Vector normalization: Normalize the sentence vector so that vectors of different sentences have a consistent scale. The normalization formula:

[0104]

[0105] where μ is the mean of the vector, σ is the standard deviation of the vector, and v normalized is the vector representation after normalization.

[0106] Result representation: The final generated unified vector representation is:

[0107] v final = v normalized

[0108] Through the above steps, the text data is converted into a standardized vector representation, preparing for subsequent modality alignment and fusion. This process ensures that text data from different sources is represented in the same vector space, facilitating unified processing and analysis.

[0109] 2.2 Image data feature extraction:

[0110] The extraction of image data features is achieved through a Convolutional Neural Network (CNN). The following are the specific steps:

[0111] Convolutional and pooling layers: Multiple convolutional and pooling layers are used to gradually extract local features of the image and reduce the dimensionality of the features.

[0112] Fully connected layer: A fully connected layer is added after the convolutional layer to map the features extracted by the convolutional layer into a vector space of a fixed dimension.

[0113] Feature vectorization: The feature vector output by the fully connected layer is used as the vector representation of the image.

[0114] 2.3 Audio data feature extraction:

[0115] The extraction of audio data features is usually carried out using Mel Frequency Cepstral Coefficients (MFCC). The following are the specific steps:

[0116] Framing: The audio signal is divided into several frames, usually with a frame length of 20 - 30 milliseconds.

[0117] Windowing: The signal of each frame is windowed, usually using a Hamming window or a Hanning window, etc.

[0118] Fourier transform: The windowed signal is subjected to a Fourier transform to convert the time-domain signal to the frequency domain.

[0119] Mel filter bank: The frequency-domain signal is passed through a set of Mel filters to obtain the energy of each filter channel.

[0120] Logarithmic transformation: The logarithm of the energy of each filter channel is taken to obtain the logarithmic energy spectrum.

[0121] Discrete Cosine Transform (DCT): The logarithmic energy spectrum is subjected to a Discrete Cosine Transform to obtain the Mel Frequency Cepstral Coefficients (MFCC).

[0122] Feature vectorization: The MFCC coefficients are used as the vector representation of the audio.

[0123] Through the above steps, the features of image data and audio data can be extracted respectively to obtain a unified vector representation for subsequent fusion and processing.

[0124] Step 3: Reduce the high-dimensional feature vectors to a low-dimensional space to reduce the computational overhead.

[0125] In the feature extraction and vector representation step, the data of different modalities has been converted into high-dimensional feature vectors. Next, non-linear dimensionality reduction is performed to reduce the computational overhead and retain the main information.

[0126] 3.1 Similarity calculation

[0127] First, calculate the similarity matrix between high-dimensional feature vectors:

[0128]

[0129] where p ij is the conditional probability between samples x i and x j , and σ i is the distance between sample x i and its nearest neighbor.

[0130] 3.2 Low-dimensional space mapping

[0131] In the low-dimensional space, the mapping positions of samples are optimized by minimizing the KL divergence between the similarities of samples in the high-dimensional space and the similarities of samples in the low-dimensional space.

[0132]

[0133] where q ij is the conditional probability between samples y i and y j .

[0134] 3.3 KL divergence calculation

[0135] Calculate the KL divergence between the conditional probability distribution P in the high-dimensional space and the conditional probability distribution Q in the low-dimensional space:

[0136]

[0137] Take the positions of samples in the low-dimensional space as the feature representation after dimensionality reduction.

[0138] 3.4 Gradient descent optimization

[0139] Minimize the KL divergence through the gradient descent algorithm to optimize the positions of samples in the low-dimensional space.

[0140] 3.5 Result representation

[0141] Finally, take the positions of samples in the low-dimensional space as the feature representation after dimensionality reduction.

[0142] Through the above method, the high-dimensional feature vectors can be reduced to a two-dimensional or three-dimensional space, preserving the local relationships between samples, which is suitable for visualization and analysis. This process significantly reduces the computational overhead while retaining the main information, improving the efficiency.

[0143] Step 4: Align the data of different modalities so that they can be represented and processed in the same semantic space.

[0144] In the feature extraction and vector representation steps, the feature vector representations of different modalities have been obtained. Next, these feature vectors will be aligned into a unified semantic space for subsequent fusion and processing.

[0145] 4.1 Shared Embedding Space Modeling

[0146] First, establish a shared embedding space into which the feature vectors of each modality can be mapped:

[0147] z = [z 1 , z 2 ,..., z m

[0148] where z i represents the i-th dimensional vector in the shared space.

[0149] 4.2 Loss Function Definition

[0150] Define a loss function to minimize the differences between different modalities and encourage them to have consistent representations in the shared space:

[0151]

[0152] where x i and y i respectively represent the feature vectors of different modalities, and l(·) is a loss function used to measure the difference between two vectors.

[0153] 4.3 Parameter Optimization

[0154] Through optimization algorithms such as gradient descent, minimize the loss function and update the model parameters so that the feature vectors of different modalities are gradually aligned in the shared space.

[0155] θ * = arg min θ L(θ)

[0156] where θ represents the model parameters.

[0157] 4.4 Alignment Result Representation

[0158] Finally, the obtained model parameters θ *It can be used to map feature vectors of different modalities into a shared space to achieve modality alignment.

[0159] z i = f(x i ; θ * )

[0160] where f(·) is a mapping function that maps the feature vector x i to the vector z in the shared space i .

[0161] Through the above steps, data of different modalities can be aligned into a unified semantic space for representation and processing. This process can improve the correlation between data of different modalities and promote the fusion and utilization of cross-modal information.

[0162] Step 5: Use a graph neural network for fused representation learning to generate a comprehensive feature representation.

[0163] In the modality alignment step, data of different modalities have been mapped into a unified semantic space. Next, fused representation learning is performed to generate a comprehensive feature representation.

[0164] 5.1 Construct a heterogeneous graph

[0165] First, construct a heterogeneous graph according to the relationships between the data. In this graph, nodes represent different data entities, and edges represent the relationships between different entities.

[0166] 5.2 Node feature initialization

[0167] Initialize the features of each node in the graph, and use the feature representation after modality alignment as the initial feature of the node:

[0168]

[0169] where z i is the initial feature representation of the i-th node.

[0170] 5.3 Graph convolutional layer

[0171] Perform message passing and aggregation on the node features through multiple graph convolutional layers to achieve fused representation learning:

[0172]

[0173] where represents the feature representation of the i-th node at the l-th layer, N(i) represents the set of neighbor nodes of node i, c ij is the number of edges between node i and node j, W (l) is the weight matrix at the l-th layer, and σ(·) is the activation function.

[0174] 5.4 Graph Pooling Layer

[0175] During the fused representation learning process, a graph pooling layer can be added to further aggregate and reduce the dimensionality of node features to extract global features.

[0176] 5.5 Result Representation

[0177] After passing through multiple graph convolutional and graph pooling layers, the resulting node feature representation is the fused comprehensive feature representation.

[0178] Through the above steps, fused representation learning can be performed on the features after modality alignment to generate a comprehensive feature representation. This process can effectively capture the complex relationships between different modality data, extract more informative feature representations, and provide better inputs for subsequent tasks.

[0179] Step 6: Integrate the fused feature representations into the knowledge graph to form a unified knowledge structure.

[0180] In the fused representation learning step, a unified feature representation has been obtained. Next, these feature representations are integrated into the knowledge graph to form a unified knowledge structure.

[0181] First, perform entity recognition and linking:

[0182] Entity recognition and linking is the process of identifying entities in the data and linking them to the corresponding nodes in the knowledge graph. The following are the implementation steps:

[0183] 6.1 Entity Recognition

[0184] Use named entity recognition (NER) technology to identify entities from text data, such as people, locations, organizations, etc.

[0185] 6.2 Entity Linking

[0186] Link the identified entities to the corresponding nodes in the knowledge graph through entity linking technology. Methods based on entity descriptions or context information can be used for entity linking.

[0187] Then perform relation extraction:

[0188] Relation extraction is to extract the relationships between entities from text data to construct the edges of the knowledge graph. The following are the implementation steps:

[0189] 6.4 Relation Extraction Model

[0190] Relationship extraction models can be used to extract the relationships between entities from text data, and relationship extraction models based on rules, statistics, deep learning, etc. can be used.

[0191] 6.5 Constructing a knowledge graph

[0192] Construct the extracted entities and their relationships into the nodes and edges of the knowledge graph. Nodes represent entities, and edges represent the relationships between entities.

[0193] Through the above steps, the fused feature representations are integrated into the knowledge graph to form a unified knowledge structure. This process can help better understand the entities and relationships in the data and provide support for subsequent knowledge reasoning and applications.

[0194] Embodiment 2

[0195] A knowledge fusion system for multi-source heterogeneous multi-modal data, including:

[0196] Data preprocessing and standardization module: Preprocess and standardize data from different sources to ensure data quality and consistency;

[0197] Feature extraction and vector representation module: Extract features from data of different modalities to generate unified vector representations;

[0198] Feature dimensionality reduction module: Reduce the high-dimensional feature vectors to a low-dimensional space to reduce computational overhead;

[0199] Modal alignment module: Align data of different modalities so that they can be represented and processed in the same semantic space;

[0200] Fusion representation learning module: Use graph neural networks for fusion representation learning to generate comprehensive feature representations;

[0201] Knowledge graph construction module: Integrate the fused feature representations into the knowledge graph to form a unified knowledge structure.

[0202] The above-mentioned various modules are used to implement the corresponding functions in the embodiments.

[0203] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the scope of protection of the claims of the present invention.

Claims

1. A knowledge fusion method for multi-source heterogeneous multi-modal data, characterized in that: The steps include: Step 1: Preprocess and standardize data from different sources to ensure data quality and consistency; Step 2: Extract features from data of different modalities and generate a unified vector representation; Step 3: Reduce the high-dimensional feature vector to a low-dimensional space to reduce computational overhead; Step 4: Align data from different modalities so that they can be represented and processed in the same semantic space; Step 5: Use graph neural network to perform fusion representation learning and generate comprehensive feature representation; Step 6: Integrate the fused feature representation into the knowledge graph to form a unified knowledge structure; The step 1 comprises: The preprocessing includes: removing noise, correcting spelling, filling missing values, and converting data in different formats into a structured format; The standardization includes: normalizing the numerical data to make its numerical range consistent; The step 2 comprises: Decompose the text into a sequence of words {w1,w2,...,w n }; Each word w i Mapped to vector v i ; Calculate sentence vectors by averaging Normalize the sentence vector The resulting uniform vector is denoted as v final ; Where n is the number of words in the sentence, v sentence is the vector representation of the sentence, μ is the mean of the vector, σ is the standard deviation of the vector, and v normalized is the normalized vector representation; The step 3 comprises: The similarity matrix between high-dimensional feature vectors is calculated using the following formula: Among them, p ij is the sample x i and x j The conditional probability between i Sample x i The distance to its nearest neighbor; The mapping position of samples is optimized by minimizing the KL divergence between the similarity between samples in high-dimensional space and the similarity between samples in low-dimensional space: Among them, q ij is the sample y i and j The conditional probability between Calculate the KL divergence between the conditional probability distribution P in the high-dimensional space and the conditional probability distribution Q in the low-dimensional space: The position of the sample in the low-dimensional space is used as the feature representation after dimensionality reduction; The step 4 comprises: Create an embedding space into which the feature vectors of each modality can be mapped: z=[z1,z2,...,z m ] Among them, z i Represents the i-th dimension vector in the shared space; Define the loss function to minimize the difference between different modalities: Among them, x i and i They represent the feature vectors of different modes respectively, and l(·) is the loss function used to measure the difference between two vectors; Through the gradient descent optimization algorithm, the loss function is minimized and the model parameters are updated so that the feature vectors of different modes are gradually aligned in the shared space: i * =argmin θ L(θ) Among them, θ represents the model parameters; The obtained model parameters θ * Used to map the eigenvectors of different modes into a shared space to achieve mode alignment: z i =f(x i ;θ * ) Among them, f(·) is the mapping function, which transforms the feature vector x i The vector z mapped to the shared space i .

2. The method for knowledge fusion of multi-source heterogeneous multi-modal data according to claim 1, characterized in that: The step 5 comprises: Construct a heterogeneous graph based on the relationship between data. In the heterogeneous graph, nodes represent different data entities, and edges represent the relationship between different entities. Initialize the features of each node in the graph and use the feature representation after mode alignment as the initial feature of the node: Among them, z i is the initial feature representation of the i-th node; Through multi-layer graph convolutional layers, node features are passed and aggregated to achieve fusion representation learning: in, represents the feature representation of the i-th node in the l-th layer, N(i) represents the set of neighbor nodes of node i, c ij is the number of edges between node i and node j, W (l) is the weight matrix of the lth layer, σ(·) is the activation function; After multiple layers of graph convolution and graph pooling layers, the obtained node feature representation is the fused comprehensive feature representation.

3. The method for knowledge fusion of multi-source heterogeneous multi-modal data according to claim 2, characterized in that: The step 6 comprises: Use NER technology to identify entities in text data and link them to corresponding nodes in the knowledge graph; Use the relation extraction model to extract the relationship between entities from text data; The identified entities and their relationships are constructed into nodes and edges of the knowledge graph.

4. A knowledge fusion system for multi-source heterogeneous multi-modal data, characterized in that: include: Data preprocessing and standardization module: preprocess and standardize data from different sources to ensure data quality and consistency; Feature extraction and vector representation module: extract features from data of different modalities and generate a unified vector representation; Feature dimensionality reduction module: reduces the high-dimensional feature vector to a low-dimensional space to reduce computational overhead; Modality alignment module: aligns data from different modalities so that they can be represented and processed in the same semantic space; Fusion representation learning module: Use graph neural network to perform fusion representation learning and generate comprehensive feature representation; Knowledge graph construction module: Integrate the fused feature representation into the knowledge graph to form a unified knowledge structure; The data preprocessing and standardization module includes: The preprocessing includes: removing noise, correcting spelling, filling missing values, and converting data in different formats into a structured format; The standardization includes: normalizing the numerical data to make its numerical range consistent; The feature extraction and vector representation module includes: Decompose the text into a sequence of words {w1,w2,...,w n }; Each word w i Mapped to vector v i ; Calculate sentence vectors by averaging Normalize the sentence vector The resulting uniform vector is denoted as v final ; Where n is the number of words in the sentence, v sentence is the vector representation of the sentence, μ is the mean of the vector, σ is the standard deviation of the vector, and v normalized is the normalized vector representation; The feature dimension reduction module includes: The similarity matrix between high-dimensional feature vectors is calculated using the following formula: Among them, p ij is the sample x i and x j The conditional probability between i Sample x i The distance to its nearest neighbor; The mapping position of samples is optimized by minimizing the KL divergence between the similarity between samples in high-dimensional space and the similarity between samples in low-dimensional space: Among them, q ij is the sample y i and j The conditional probability between Calculate the KL divergence between the conditional probability distribution P in the high-dimensional space and the conditional probability distribution Q in the low-dimensional space: The position of the sample in the low-dimensional space is used as the feature representation after dimensionality reduction; The modality alignment module comprises: Create an embedding space into which the feature vectors of each modality can be mapped: z=[z1,z2,...,z m ] Among them, z i Represents the i-th dimension vector in the shared space; Define the loss function to minimize the difference between different modalities: Among them, x i and i They represent the feature vectors of different modes respectively, and l(·) is the loss function used to measure the difference between two vectors; Through the gradient descent optimization algorithm, the loss function is minimized and the model parameters are updated so that the feature vectors of different modes are gradually aligned in the shared space: i * =argmin θ L(θ) Among them, θ represents the model parameters; The obtained model parameters θ * Used to map the eigenvectors of different modes into a shared space to achieve mode alignment: z i =f(x i ;θ * ) Among them, f(·) is the mapping function, which transforms the feature vector x i The vector z mapped to the shared space i .

Citation Information

Patent Citations

  • Multi-modal knowledge graph construction method

    CN112200317A

  • Power grid dispatching multi-mode knowledge graph construction method and system

    CN118035463A

Cited By

  • Multi-modal heterogeneous knowledge fusion construction and semantic enhancement retrieval system based on large model

    CN120705362A