Structural information enhanced multi-modal heterogeneous data fusion characterization method
This multimodal heterogeneous data fusion and representation method, which enhances structural information, solves the overfitting problem in multimodal data fusion and processing, improves data representation quality and model generalization ability, and is applicable to fields such as healthcare, security, and autonomous driving.
Patent Information
- Application Number
- CN202511057023.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-07-30
AI Technical Summary
The fusion and processing of multimodal heterogeneous data presents challenges. Traditional methods struggle to fully extract information from the data structure, leading to problems such as overfitting in large-scale data processing, which affects performance and generalization ability.
A multimodal heterogeneous data fusion representation method with enhanced structural information is adopted. Through preprocessing, feature extraction and transformation, graph structure enhancement technology, structural entropy regularization and soft allocation mechanism, a hypergraph structure is constructed to optimize data representation and clustering performance.
It improves the representation quality and discrimination ability of multimodal data, enhances the generalization ability of the model, improves adaptability, and is suitable for complex data scenarios.
Smart Images

Figure CN120996150A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, in particular to a multi-modal heterogeneous data fusion representation method with structural information enhancement. BACKGROUND
[0002] With the rapid development of information technology, data generation and acquisition have become increasingly easy, and data types have become increasingly diverse, including text, images, audio, video and other modalities. These multi-modal heterogeneous data play a key role in many fields such as medical care, security, autonomous driving, intelligent transportation, finance, etc. For example, in the medical field, doctors need to integrate patient medical records, X-ray films, CT images and other data to make accurate diagnoses. In security monitoring, the system needs to integrate video images and sound signals to identify potential security threats. In autonomous driving, vehicles need to process visual data from cameras, ranging data from radars, and other data from vehicle sensors to make safe driving decisions.
[0003] However, the fusion and processing of multi-modal heterogeneous data still face many challenges. Different modalities of data have different features and structures, and how to effectively fuse different modalities of data becomes the biggest problem. Traditional data representation methods are difficult to fully exploit the structural information in the data, so many large models cannot fully utilize these data, especially when the model is processing large-scale data, overfitting and other conditions are particularly prone to occur, thereby affecting the model performance and generalization ability. In addition, due to the complementarity and correlation between different modalities of data, this complementarity and correlation often needs to be further explored to effectively improve the quality and discriminative ability of data representation.
[0004] In recent years, in response to related problems, many multi-modal data fusion and representation learning methods have emerged. Among these methods, graph structure-based learning methods have received widespread attention due to their high efficiency in capturing data. However, these methods mostly rely on pre-defined graph structures, and their adaptability is relatively poor for unstructured data or complex data distribution. In the fusion process, how to fully utilize structural information to improve the discriminative ability and generalization ability of data representation is still a problem to be solved. Therefore, it is of great practical significance and broad application prospect to invent a technology that can efficiently fuse multi-modal heterogeneous data and improve data representation level. This technology can fundamentally improve the quality of data representation, enhance the discriminative ability of the model, and strengthen the generalization ability of the model, so that the model not only is accurate but also has stability when processing large-scale data. It can provide more powerful and efficient technical support for various complex data analysis and decision support tasks, and its application scenarios are widely involved in medical care, security, autonomous driving and other fields. SUMMARY
[0005] In order to solve the above technical problems, the application provides a multi-modal heterogeneous data fusion representation method with structural information enhancement.
[0006] In order to realize the above technical problems, the steps include: S1, obtaining multi-modal heterogeneous data through text, image, audio and video, performing feature extraction and conversion operation after performing preprocessing operation on the multi-modal heterogeneous data, and obtaining a multi-modal feature matrix in a unified feature space; Multi-modal heterogeneous data plays a key role in many fields such as medical treatment, security and protection, automatic driving and intelligent transportation; for example, in the medical field, doctors need to integrate patient medical records, X-ray films, CT images and patient language description data to make accurate diagnosis; in security monitoring, the system needs to fuse video images, sound signals and environmental audio to identify potential security threats; in an automatic driving vehicle, the vehicle needs to process visual data from a camera, ranging data from a radar and other data from a vehicle sensor to make a safe driving decision; In the application, the preprocessing operation includes cleaning, normalizing, aligning and missing value processing operation on the original input multi-modal heterogeneous data; In the application, the feature extraction and conversion operation is as follows: using a modal model to extract modal features, and mapping different modal features to the same hidden space through linear transformation; in the application, the modal model includes CNN for extracting image features and BERT for extracting text features.
[0007] S2, based on the multi-modal feature matrix, using a graph structure enhancement technology (GSL) to construct a graph structure with stability and interpretability, the steps are as follows: S2.1, calculating the inner product of the multi-modal feature matrix, constructing an initial similarity graph, and the expression is as follows: In the formula, denotes the similarity between the vertex and the vertex denotes the multi-modal feature matrix; the superscript denotes transposition; The application constructs an initial similarity graph by calculating the inner product of the data features, and based on the inner product operation, the similarity between different data points can be measured, which provides a basis for the construction of subsequent graph structure; for example, when processing a data set containing image and text data, the application constructs an initial similarity graph by calculating the inner product of the image features and the text features; this similarity graph based on the inner product can initially present the similarity between the data points, and provides a key reference for the optimization of the subsequent graph structure. S2.2, based on the initial similarity graph, an adjacency matrix is constructed, through sparse processing, the calculation complexity is reduced, the calculation efficiency is improved, and the local structure information between data points is highlighted; The expression for constructing the adjacency matrix based on the initial similarity graph is as follows: In the formula, denotes the adjacency matrix; denotes the activation function; The sparse processing method is to use the K-Nearest Neighbors (KNN) algorithm to retain the top k most similar neighbors of each node; The expression of sparse processing is as follows: In the formula, denotes the sparse similarity graph; denotes the similarity graph before sparse processing, which is the feature similarity graph in the present application ; denotes the vertex belongs to the top most similar neighbors of the vertex ; The sparse processing method can reduce the complexity of the graph, improve the calculation efficiency, and highlight the local structure information between data points; S2.3, post-processing operation is carried out on the adjacency matrix after sparse processing, including symmetry operation, activation function processing and row normalization processing, so as to guarantee the undirectedness of the graph and ensure the realization of normalization of the graph; Symmetry operation: in order to ensure that the graph is undirected, the present application realizes it by symmetrizing the adjacency matrix, and the expression is as follows: In the formula, denotes the adjacency matrix of the symmetry operation; Activation function processing: taking the adjacency matrix of the symmetry operation as input, applying the activation function (such as sigmoid) to enhance the stability of the graph structure, and the expression is as follows: In the formula, denotes the adjacency matrix after activation function processing; Row normalization processing: the row normalization processing is performed on the adjacency matrix after the activation function processing, so as to ensure that the weight sum of each node is 1, and the expression is as follows: In the formula, denotes the first xrow, the y elements of a column; represents the sum of all elements of the adjacency matrix after the activation function processing of the row x row; represents the adjacency matrix after row normalization processing; After the above graph structure enhancement step, the application can effectively improve the representation learning performance of multi-modal data, ensuring the accuracy and integrity of the model when processing complex multi-modal data. This method is particularly suitable for scenarios such as sentiment analysis, image classification, and speech recognition.
[0008] S3, based on the graph structure, a structure entropy regularized discriminative representation learning framework is constructed to improve the generalization ability and clustering performance of the model; wherein the structure entropy regularized discriminative representation learning framework includes an overfitting regularizer of information entropy and a class regularizer of structure entropy; The steps are as follows: S3.1, an overfitting regularizer of information entropy is constructed, including the purposes of punishing low entropy distribution, achieving fast convergence, and preventing overfitting, and the construction steps include: S3.1.1, punish low entropy distribution: punish low entropy distribution for the model output, introduce a low entropy distribution penalty, calculate the entropy of the conditional distribution for the SoftMax output of the neural network, and add it as a regularization term to the loss function, the expression is as follows: In the formula, represents the conditional probability distribution of the predicted class below the input of the model; represents the low entropy distribution penalty hyperparameter, used to control the strength of the entropy regularization term; represents the negative entropy term, used to punish low entropy distribution, wherein The expression of is as follows: In the formula, represents the predicted class of the a-th class; S3.1.2, fast convergence and prevention of overfitting: in order to achieve fast convergence and effectively prevent overfitting, the application designs a method of dynamically adjusting the strength of entropy regularization. In the initial stage of training, the entropy regularization strength is in a weak state. As the training process progresses, when the model begins to overfit, the strength of entropy regularization is gradually increased. This is achieved by setting a threshold, that is, when the entropy of the model output distribution is lower than the threshold, stronger entropy regularization is applied. The formula can be expressed as: In the formula, This represents the entropy threshold; strong regularization is triggered when the actual entropy falls below this value. S3.1.3, The deviation of the output distribution from the uniform distribution is penalized through label smoothing operation to smooth the distribution; the expression is as follows: In the formula, This represents the Kullback-Leibler divergence, used to measure the difference between two probability distributions; Indicates a uniform distribution; S3.2 Construct the category regularization calculation rule for structural entropy, the expression is as follows: In the formula, It is the main loss function; It is a regularization coefficient used to balance the relationship between the principal loss and the structural entropy regularization term; This represents structural entropy; in this way, the model can not only learn the main features of the data, but also maintain the balance between categories, thereby improving the ability to distinguish between different categories of data and the effect of handling imbalanced data. Following the steps outlined above, the structural entropy category regularization method can effectively guide the model to learn and obtain a better data representation, thereby improving the clustering performance of multimodal data. This method can enhance the model's ability to distinguish between different categories of data and improve the model's performance in handling imbalanced data, allowing the model to learn representations more accurately when faced with complex multimodal data.
[0009] S4. Design a structural information optimization framework and soft allocation mechanism to improve the quality of data representation and discriminative power; This invention conceives an innovative framework for optimizing structural information, based on a soft allocation mechanism to optimize graph structure learning. This allows the model to better capture structural information in the data. The soft allocation mechanism enables the model to dynamically adjust the connection weights between nodes during the learning process, thereby better matching the structural characteristics of the data. This dynamic adjustment capability allows the model to more flexibly handle complex relationships between different modalities of data, improving the quality and discriminative power of data representation. For example, when processing a multimodal dataset covering text, images, and audio data, the model can dynamically adjust the connection weights between text features, image features, and audio features through the soft allocation mechanism, more effectively capturing the structural information between them and improving the quality and discriminative power of data representation. The expression for the soft allocation mechanism is as follows: In the formula, and Represents vertices and vertex The feature vectors are derived from the row vectors of the multimodal feature matrix; Represents the inner product of eigenvectors; This represents the temperature parameter, used to control the weight distribution; Represents vertices The set of neighbors, among which, It is the vertex The index of the neighbor set; Represents vertices Neighboring vertices eigenvectors; superscript Indicates transpose; Represents a node and nodes Connection weights between them; Using this method, the soft allocation mechanism can dynamically adjust the connection weights between nodes, allowing the model to better adapt to the structural characteristics of the data.
[0010] S5. Based on the obtained graph structure, the discriminative representation learning framework of structural entropy regularization, and the structural information optimization framework and soft allocation mechanism, construct the hypergraph structure and calculate the structural entropy of the hypergraph structure. S5.1 Define the hypergraph structure; Define an undirected weighted hypergraph This undirected weighted hypergraph contains a set of vertices. Hyperedge set and weight matrix The specific definition is as follows: S5.1.1 Constructing the Vertex Set of the Hypergraph ; The set of vertices of a hypergraph The set of vertices in the graph structure corresponds to the data samples, i.e. The row vector; S5.1.2 Constructing the set of hyperedges of a hypergraph ; The set of superedges of a hypergraph Each hyperedge connects a set of related hypergraph vertices. The method for constructing hyperedges is as follows: For each vertex in the feature space of a hypergraph structure, the K-nearest neighbor algorithm is used to compute cross-modal similar neighbors. For example, in the medical field, a hyperedge includes { Images, medical records, electrocardiograms, and patient reports; S5.1.3 Constructing the weight matrix of the hypergraph ; Hypergraph weight matrix The weight of each hyperedge is stored using the following expression: In the formula, Indicates the super edge. ; Represents the two vertices of a superedge; Representing vertices in the feature space ; Represents vertices and Distance in feature space; The bandwidth parameter represents the Gaussian kernel function; S5.2 Define the cut edges and cut edge volumes of the hypergraph structure; Volume of the vertex set Expanding the calculation, which is defined as follows: , here The degree of a vertex, a metric that helps in understanding the "size" of a set of vertices within a hypergraph, is defined as the cut edges formed with the outside: In the formula, Represents the volume of the vertex set The set of cut edges; for The complement of the vertex set; Representing an edge At least one endpoint is in the set middle; Representing an edge At least one endpoint is in the set In the middle; these cut edges serve to connect the vertex sets. The roles of internal and external vertices; The expression for the cut volume is as follows: In the formula, The cut edge volume represents the degree of the hyperedge. This formula shows the strength of the connection between the vertex set and the rest of the hypergraph. The larger the cut edge volume, the stronger the connection between the vertex set and the rest of the hypergraph. Conversely, the larger the cut edge volume, the weaker the connection. S5.3 Generate the incidence matrix of the hypergraph structure, with the following expression: Correlation Matrix It can be used to represent the relationship between a vertex and a hyperedge, when the vertex... Belongs to superedge At that time, Otherwise, it is 0; S5.4. Based on the generated incidence matrix, construct a clique adjacency matrix to convert the higher-order relations of the hypergraph into ordinary graphs, as shown in the following expression: In the formula, Represents the clique-based adjacency matrix; The inverse matrix represents the hyperedge degree; the clique adjacency matrix provides a structured perspective for observing and analyzing hypergraphs. S5.5 Calculate the hypergraph entropy and cut edge entropy of the hypergraph structure; The expression for hypergraph entropy is as follows: In the formula, Represents the hypergraph entropy; Represents the hypergraph structure vertices in the hypergraph structure. The set of neighbors; Represents vertices in a hypergraph structure ; Represents vertices in a hypergraph structure and vertex soft links; The expression for the cut edge entropy is as follows: In the formula, This represents a category matrix, indicating the connections between categories. This indicates the total number of categories, such as disease categories in the medical field, where... An index representing the total number of categories; Hypergraph Entropy It can measure the complexity and information content of a hypergraph, such as cut edge entropy. This can reflect the uniformity of the cut edge distribution. These entropy values will be applied to the regularization model to improve its generalization ability, ensuring that the model performs well on the training data and maintains good performance on unseen data.
[0011] S6. Based on hypergraph structure entropy, guide multimodal unsupervised clustering; The steps include: S6.1, Multimodal data fusion; This invention fuses data from different modalities such as text, images, and audio, and maps them to the same feature space through feature extraction and transformation, providing a unified data representation for subsequent clustering analysis; S6.2, Structural entropy regularization-guided clustering; The clustering objective function can be expressed as: In the formula, This represents the basic clustering loss function (such as the squared error of K-means clustering). The regularization coefficient representing the balance structure entropy term and the basic loss; Represents vertices in a hypergraph structure The structural entropy is expressed as follows: ; After completing the step of optimizing structural entropy, the model can capture complex structural information in the data more efficiently, thereby further improving the accuracy of cluster analysis. S6.3 Clustering result optimization; This invention is based on the clustering results after structural entropy regularization, and further optimizes the cluster centers and cluster labels by iterative optimization until the minimum convergence condition is met. The iterative optimization process can be represented as: In the formula, Indicates the first During the nth iteration, the 1st The center vectors of each cluster; The set of data points belonging to the i-th cluster is expressed as follows: In the formula, Indicates the first During the nth iteration, it belongs to the... A set of data points in a cluster; Indicates the first Cluster centers at the next iteration; Indicates the number of clusters; This invention can not only achieve efficient fusion processing of multimodal heterogeneous data, but also complete high-quality characterization work, providing more efficient technical support for various complex tasks, completing tasks efficiently and safely, and has wide applications in fields such as medical care, security, and autonomous driving.
[0012] Beneficial effects of the present invention To achieve efficient fusion of multimodal data, this invention integrates text, images, audio, and other modalities more effectively and stably. The main purpose of this fusion method is to uncover the complementarity and correlation between different modalities, thereby improving the quality of the model's data representation and its discriminative ability. The application scenarios of this fusion method are very broad: in the medical field, combining patient medical records with medical image data can provide doctors with more comprehensive and complete information about the patient's condition; in the field of security monitoring, the fusion of video images and sound signals can more accurately identify potential security threats, thereby ensuring safety; in the field of autonomous vehicles, the integration of visual data and radar ranging data can provide accurate and useful information for safe driving decisions, thereby reducing the probability of accidents. To optimize the utilization of data structural information and address the problem of incomplete information utilization in traditional data representation methods, this invention optimizes data structural information based on structural entropy regularization. It studies the structural entropy of feature similarity maps, mainly including the calculation of local entropy and joint entropy, and incorporates the obtained structural entropy as a regularization term into the loss function. This enables the model to perform deep learning on the structural features of the data during training, thereby improving the discriminative ability of data representation.
[0013] To enhance model generalization ability and address the common problems of overfitting and class imbalance during model training, this invention proposes a structural entropy regularization method. The core of this invention lies in optimizing data structure information. By deeply exploring the inherent structural relationships within the data, it enhances the model's ability to process and analyze data from different categories. During training, the structural entropy regularization method dynamically adjusts the model's parameters, preventing the model from overlearning local features from old training data and reducing the probability of overfitting. Furthermore, this method balances the contributions of different classes of data during model training, improving the model's performance in recognizing a minority of classes.
[0014] To improve the processing capabilities of unstructured data, this invention can automatically extract structural information from the data. In real-world scenarios, a large amount of data lacks a clearly defined graph structure, and the construction process is extremely complex. Traditional learning methods that rely on graph structures are often limited by pre-defined graph structures and struggle to adapt well to unstructured or complex data distributions. Unlike traditional methods, the method proposed in this invention possesses the ability to automatically learn data structure information, improving the efficiency of processing unstructured data, effectively overcoming traditional limitations, and providing a new solution. Attached Figure Description
[0015] Figure 1 This is a flowchart of the steps of the present invention; Figure 2 This is an overall framework diagram of the present invention; Figure 3 This is the multimodal unsupervised clustering diagram of the present invention. Detailed Implementation
[0016] The present invention will be further described in detail below with reference to specific embodiments.
[0017] like Figure 1 and Figure 2 As shown, a multimodal heterogeneous data fusion and representation method with enhanced structural information includes the following steps: S1. Obtain multimodal heterogeneous data through text, images, audio and video. Perform preprocessing operations on the multimodal heterogeneous data and then perform feature extraction and transformation operations to obtain a multimodal feature matrix in a unified feature space. Multimodal heterogeneous data plays a crucial role in many fields such as healthcare, security, autonomous driving, and intelligent transportation. For example, in healthcare, doctors need to integrate various data such as patient medical records, X-rays, CT images, and patient verbal descriptions to make accurate diagnoses. In security monitoring, systems need to integrate video images, sound signals, and environmental audio to identify potential security threats. In autonomous vehicles, vehicles need to process visual data from cameras, ranging data from radar, and other data from vehicle sensors to make safe driving decisions. In this invention, the preprocessing operations include: cleaning, normalizing, aligning, and handling missing values of the original input multimodal heterogeneous data; In this invention, the feature extraction and transformation operation is performed as follows: using a modal model to extract features of each modality, and mapping different modal features to the same latent space through linear transformation; in this invention, the modal model includes: CNN to extract image features, and BERT to extract text features.
[0018] S2. Based on the multimodal feature matrix, construct a stable and interpretable graph structure using Graph Structure Enhancement (GSL) techniques. The steps are as follows: S2.1 Calculate the inner product of the multimodal feature matrices and construct the initial similarity graph, as shown in the following expression: In the formula, Represents vertices and vertex The similarity between them; Represents the multimodal feature matrix; superscript This indicates transpose; by calculating the inner product between feature vectors, a similarity graph reflecting the similarity between data points can be obtained. This similarity graph can initially present the similarity between data points and provide a key reference for subsequent graph structure optimization. For example, when processing a dataset that includes image and text data, the inner product between image features and text features can be calculated to construct the initial similarity graph. This invention constructs an initial similarity graph by calculating the inner product between data features. The inner product operation can measure the similarity between different data points, providing a foundation for the construction of subsequent graph structures. For example, when processing a dataset containing image and text data, this invention constructs an initial similarity graph by calculating the inner product of image features and text features. This inner product-based similarity graph can initially present the similarity between data points, providing a key reference for the optimization of subsequent graph structures. S2.2 Construct an adjacency matrix based on the initial similarity graph. By sparsifying the matrix, the computational complexity is reduced, the computational efficiency is improved, and the local structural information between data points is highlighted. The expression for constructing the adjacency matrix based on the initial similarity graph is as follows: In the formula, Represents the adjacency matrix; This represents an activation function used to introduce non-linear characteristics, such as the sigmoid activation function. The sparsity reduction method is to use the K-Nearest Neighbors (KNN) algorithm to retain the k most similar neighbors of each node; The expression for sparsification is as follows: In the formula, Represents a similar graph after sparsification; This represents a similarity graph before sparsification; in this invention, it is a feature similarity graph. ; Represents vertices Belongs to the distance from the vertex The former The most similar neighbor; Sparsity processing can reduce the complexity of the graph, improve computational efficiency, and highlight the local structural information between data points. For example, when processing a dataset containing 1000 data points, this invention can retain only the top 10 most similar neighbors of each node to construct a sparse similarity graph. This sparse similarity graph can reduce computational costs and more clearly present the local structural information between data points, providing a more accurate reference for subsequent graph structure optimization. S2.3. Post-processing operations on the expanded adjacency matrix after sparsification, including: symmetry operation, activation function processing and row normalization processing, to ensure the undirectedness of the graph and to ensure that the graph is normalized. Symmetry operation: To ensure the graph is undirected, this invention achieves this by symmetricizing the adjacency matrix, as shown in the following expression: In the formula, Represents the adjacency matrix for symmetrization operations; Activation function processing: Taking the adjacency matrix after symmetrization as input, an activation function (such as sigmoid) is applied to enhance the stability of the graph structure. The expression is as follows: In the formula, This represents the adjacency matrix after activation function processing; Row normalization: The adjacency matrix after activation function processing is row normalized to ensure that the sum of the weights of each node is 1. The expression is as follows: In the formula, This represents the first element in the adjacency matrix after activation function processing. x Okay, number y Column elements; The adjacency matrix after activation function processing is represented by the [missing information - likely a specific matrix or feature]. x Sum all elements in the row; This represents the adjacency matrix after row normalization. In addition to ensuring the undirectedness and normalization of the graph through post-processing operations, the graph structure discrimination ability is also improved through activation functions, so that the model can learn the intrinsic structural information of multimodal data more efficiently and make full use of it in actual operation. Through the above graph structure enhancement steps, the present invention can effectively improve the representation learning performance of multimodal data and ensure the accuracy and integrity of the model when processing complex multimodal data. This method is particularly suitable for scenarios such as sentiment analysis, image classification, and speech recognition.
[0019] S3. Based on graph structure, construct a discriminative representation learning framework with structural entropy regularization to improve the model's generalization ability and clustering performance; the discriminative representation learning framework with structural entropy regularization includes: overfitting regularization of information entropy and category regularization of structural entropy. In the field of multimodal data processing, neural network models often face the problem of overfitting. When overfitting occurs, the model performs well on the training data, but its generalization ability drops significantly on unseen data. To solve this problem, this invention employs various regularization techniques to improve model performance. By adjusting the probability distribution of the model's output, the generalization ability of the model is improved. This is based on maximizing the entropy of the model's output, which conforms to the maximum entropy principle, that is, under empirical constraints, the uniform probability distribution with the greatest uncertainty is selected. The steps are as follows: S3.1 Construct overfitting regularization rules for information entropy, including penalizing low-entropy distributions, achieving fast convergence, and preventing overfitting. The construction steps include: S3.1.1, Penalizing Low-Entropy Distribution: A penalty is imposed on the low-entropy distribution exhibited by the model output. This penalty is introduced by calculating the entropy of the conditional distribution for the SoftMax output of the neural network and adding it as a regularization term to the loss function. The expression is as follows: In the formula, Represents model input Next Prediction Category The conditional probability distribution; This represents a hyperparameter for penalizing low-entropy distributions, used to control the strength of the entropy regularization term; The term represents negative entropy, used to penalize low-entropy distributions, where... The expression is: In the formula, Indicates the predicted category of class a; S3.1.2 Fast Convergence and Overfitting Prevention: To achieve fast convergence and effectively prevent overfitting, this invention designs a method for dynamically adjusting the entropy regularization strength. At the beginning of training, the entropy regularization strength is relatively weak. As training progresses and the model begins to overfit, the entropy regularization strength is gradually increased. This is achieved by setting a threshold; when the entropy of the model's output distribution falls below this threshold, a stronger entropy regularization is applied. The formula can be expressed as: In the formula, This represents the entropy threshold; strong regularization is triggered when the actual entropy falls below this value. S3.1.3, The label smoothing operation penalizes the deviation of the output distribution from the uniform distribution to smooth the distribution, encouraging the distribution not to be too sharp; the expression is as follows: In the formula, This represents the Kullback-Leibler divergence, used to measure the difference between two probability distributions; Indicates a uniform distribution; S3.2 Construct the category regularization calculation rule for structural entropy, the expression is as follows: In the formula, It is the main loss function; It is a regularization coefficient used to balance the relationship between the principal loss and the structural entropy regularization term; This represents structural entropy; in this way, the model can not only learn the main features of the data, but also maintain the balance between categories, thereby improving the ability to distinguish between different categories of data and the effect of handling imbalanced data. Following the steps outlined above, the structural entropy category regularization method can effectively guide the model to learn and obtain a better data representation, thereby improving the clustering performance of multimodal data. This method can enhance the model's ability to distinguish between different categories of data and improve the model's performance in handling imbalanced data, allowing the model to learn representations more accurately when faced with complex multimodal data.
[0020] S4. Design a structural information optimization framework and soft allocation mechanism to improve the quality of data representation and discriminative power; This invention conceives an innovative framework for optimizing structural information, based on a soft allocation mechanism to optimize graph structure learning. This allows the model to better capture structural information in the data. The soft allocation mechanism enables the model to dynamically adjust the connection weights between nodes during the learning process, thereby better matching the structural characteristics of the data. This dynamic adjustment capability allows the model to more flexibly handle complex relationships between different modalities of data, improving the quality and discriminative power of data representation. For example, when processing a multimodal dataset covering text, images, and audio data, the model can dynamically adjust the connection weights between text features, image features, and audio features through the soft allocation mechanism, more effectively capturing the structural information between them and improving the quality and discriminative power of data representation. The expression for the soft allocation mechanism is as follows: In the formula, and Represents vertices and vertex The feature vectors are derived from the row vectors of the multimodal feature matrix; Represents the inner product of eigenvectors; This represents the temperature parameter, used to control the weight distribution; Represents vertices The set of neighbors, among which, It is the vertex The index of the neighbor set; Represents vertices Neighboring vertices eigenvectors; superscript Indicates transpose; Represents a node and nodes Connection weights between them; Using this method, the soft allocation mechanism can dynamically adjust the connection weights between nodes, allowing the model to better adapt to the structural characteristics of the data.
[0021] S5. Based on the obtained graph structure, the discriminative representation learning framework of structural entropy regularization, and the structural information optimization framework and soft allocation mechanism, construct the hypergraph structure and calculate the structural entropy of the hypergraph structure. S5.1 Define the hypergraph structure; Define an undirected weighted hypergraph This undirected weighted hypergraph contains a set of vertices. Hyperedge set and weight matrix The specific definition is as follows: S5.1.1 Constructing the Vertex Set of the Hypergraph ; The set of vertices of a hypergraph The set of vertices in the graph structure corresponds to the data samples, i.e. The row vector; S5.1.2 Constructing the set of hyperedges of a hypergraph ; The set of superedges of a hypergraph Each hyperedge connects a set of related hypergraph vertices. The method for constructing hyperedges is as follows: For each vertex in the feature space of a hypergraph structure, the K-nearest neighbor algorithm is used to compute cross-modal similar neighbors. For example, in the medical field, a hyperedge includes { Images, medical records, electrocardiograms, and patient reports; S5.1.3 Constructing the weight matrix of the hypergraph ; Hypergraph weight matrix The weight of each hyperedge is stored using the following expression: In the formula, Indicates the super edge. ; Represents the two vertices of a superedge; Representing vertices in the feature space ; Represents vertices and Distance in feature space; The bandwidth parameter represents the Gaussian kernel function; S5.2 Define the cut edges and cut edge volumes of the hypergraph structure; Volume of the vertex set Expanding the calculation, which is defined as follows: , here The degree of a vertex, a metric that helps in understanding the "size" of a set of vertices within a hypergraph, is defined as the cut edges formed with the outside: In the formula, Represents the volume of the vertex set The set of cut edges; for The complement of the vertex set; Representing an edge At least one endpoint is in the set middle; Representing an edge At least one endpoint is in the set In the middle; these cut edges serve to connect the vertex sets. The roles of internal and external vertices; The expression for the cut volume is as follows: In the formula, The cut edge volume represents the degree of the hyperedge. This formula shows the strength of the connection between the vertex set and the rest of the hypergraph. The larger the cut edge volume, the stronger the connection between the vertex set and the rest of the hypergraph. Conversely, the larger the cut edge volume, the weaker the connection. S5.3 Generate the incidence matrix of the hypergraph structure, with the following expression: Correlation Matrix It can be used to represent the relationship between a vertex and a hyperedge, when the vertex... Belongs to superedge At that time, Otherwise, it is 0; S5.4. Based on the generated incidence matrix, construct a clique adjacency matrix to convert the higher-order relations of the hypergraph into ordinary graphs, as shown in the following expression: In the formula, Represents the clique-based adjacency matrix; The inverse matrix represents the hyperedge degree; the clique adjacency matrix provides a structured perspective for observing and analyzing hypergraphs. S5.5 Calculate the hypergraph entropy and cut edge entropy of the hypergraph structure; The expression for hypergraph entropy is as follows: In the formula, Represents the hypergraph entropy; Represents the hypergraph structure vertices in the hypergraph structure. The set of neighbors; Represents vertices in a hypergraph structure ; Represents vertices in a hypergraph structure and vertex soft links; The expression for the cut edge entropy is as follows: In the formula, This represents a category matrix, indicating the connections between categories. This indicates the total number of categories, such as disease categories in the medical field, where... An index representing the total number of categories; Hypergraph Entropy It can measure the complexity and information content of a hypergraph, such as cut edge entropy. This can reflect the uniformity of the cut edge distribution. These entropy values will be applied to the regularization model to improve its generalization ability, ensuring that the model performs well on the training data and maintains good performance on unseen data.
[0022] S6, such as Figure 3 As shown, multimodal unsupervised clustering is guided based on hypergraph structure entropy; The steps include: S6.1, Multimodal data fusion; This invention fuses data from different modalities such as text, images, and audio, and maps them to the same feature space through feature extraction and transformation, providing a unified data representation for subsequent clustering analysis. For example, when processing a dataset that includes text, images, and audio data, this invention uses feature extraction and transformation to map them to the same feature space, providing a unified data representation for subsequent clustering analysis. This multimodal data fusion method can fully utilize the complementarity and correlation between different modalities of data, improving the accuracy and reliability of clustering analysis. S6.2, Structural entropy regularization-guided clustering; In the unsupervised clustering process, this invention introduces a structural entropy regularization term to guide the clustering process based on the optimized structural information of the data. Specifically, this invention calculates the structural entropy of multimodal data and adds it as a regularization term to the clustering objective function, encouraging a uniform state of cut edges between different categories in the clustering results, thus achieving more accurate clustering analysis. For example, when processing a multimodal dataset containing text, images, and audio data, this invention calculates the structural entropy of the multimodal data, allowing the model to better capture the structural information between text, images, and audio data, achieving more accurate clustering analysis. The clustering objective function can be expressed as: In the formula, This represents the basic clustering loss function (such as the squared error of K-means clustering). The regularization coefficient representing the balance structure entropy term and the basic loss; Represents vertices in a hypergraph structure The structural entropy is expressed as follows: ; After completing the step of optimizing structural entropy, the model can capture complex structural information in the data more efficiently, thereby further improving the accuracy of cluster analysis. S6.3 Clustering result optimization; This invention is based on the clustering results after structural entropy regularization. It further optimizes the cluster centers and cluster labels through iterative optimization until the minimum convergence condition is met. For example, when processing a dataset containing 1000 data points, the cluster centers and cluster labels are continuously adjusted through iterative optimization until the clustering results tend to stabilize. The iterative optimization process can be represented as: In the formula, Indicates the first During the nth iteration, the 1st The center vectors of each cluster; The set of data points belonging to the i-th cluster is expressed as follows: In the formula, Indicates the first During the nth iteration, it belongs to the... A set of data points in a cluster; Indicates the first Cluster centers at the next iteration; This indicates the number of clusters; in this way, the cluster centers and labels are effectively adjusted, thereby improving the accuracy and stability of the clustering results. Through the above technical solutions, this invention can not only achieve efficient fusion processing of multimodal heterogeneous data, but also complete high-quality characterization work, providing more efficient technical support for various complex tasks, completing tasks efficiently and safely, and is widely used in fields such as medical care, security, and autonomous driving.
[0023] The specific embodiments of the present invention have been described in detail above. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention.
Claims
1. A method for fusion and characterization of multimodal heterogeneous data with enhanced structural information, characterized in that, Includes the following steps: S1. Obtain multimodal heterogeneous data through text, images, audio and video. Perform preprocessing operations on the multimodal heterogeneous data and then perform feature extraction and transformation operations to obtain a multimodal feature matrix in a unified feature space. S2. Based on the multimodal feature matrix, a graph structure enhancement technique is used to construct a stable and interpretable graph structure; S3. Based on graph structure, construct a discriminative representation learning framework with structural entropy regularization to improve the model's generalization ability and clustering performance. The discriminative representation learning framework for structural entropy regularization includes: overfitting regularization of information entropy and category regularization of structural entropy; S4. Design a structural information optimization framework and soft allocation mechanism to improve the quality of data representation and discriminative power; S5. Based on the obtained graph structure, the discriminative representation learning framework of structural entropy regularization, and the structural information optimization framework and soft allocation mechanism, construct the hypergraph structure and calculate the structural entropy of the hypergraph structure. S6. Based on hypergraph structure entropy, guide multimodal unsupervised clustering to complete the multimodal heterogeneous data fusion and representation technology with enhanced structural information.
2. The multimodal heterogeneous data fusion and characterization method with enhanced structural information according to claim 1, characterized in that, The steps for constructing a stable and interpretable graph structure using graph structure enhancement techniques based on a multimodal feature matrix include: S2.1 Calculate the inner product of the multimodal feature matrices and construct the initial similarity graph; S2.2 Construct an adjacency matrix based on the initial similarity graph. By sparsifying the matrix, the computational complexity is reduced, the computational efficiency is improved, and the local structural information between data points is highlighted. S2.
3. Post-processing operations on the expanded adjacency matrix after sparsification, including: symmetry operation, activation function processing and row normalization processing, to ensure the undirectedness of the graph and to ensure that the graph is normalized. The symmetrization operation ensures that the graph is undirected, which is achieved by symmetrizing the adjacency matrix; The activation function takes the adjacency matrix of the symmetry operation as input and applies the activation function to enhance the stability of the graph structure. The row normalization process involves normalizing the adjacency matrix after activation function processing to ensure that the sum of the weights of each node is 1.
3. The multimodal heterogeneous data fusion and characterization method with enhanced structural information according to claim 1, characterized in that, The steps for constructing a discriminative representation learning framework based on graph structure and structural entropy regularization to improve the model's generalization ability and clustering performance are as follows: S3.1 Construct overfitting regularization rules for information entropy, including penalizing low-entropy distributions, achieving fast convergence, and preventing overfitting. S3.2 Construct the category regularization calculation rules for structural entropy.
4. The multimodal heterogeneous data fusion and characterization method with enhanced structural information according to claim 3, characterized in that, The overfitting regularization calculation rule for constructing information entropy includes penalizing low-entropy distributions, achieving fast convergence, and preventing overfitting. The steps include: S3.1.1, Penalizing Low-Entropy Distribution: Penalizing low-entropy distribution is implemented for the low-entropy distribution presented by the model output. For the SoftMax output of the neural network, the entropy of the conditional distribution is calculated and added as a regularization term to the loss function. S3.1.2 Fast Convergence and Overfitting Prevention: Set a threshold. When the entropy of the output distribution falls below this threshold, perform fast convergence and overfitting prevention operations. The expression is as follows: In the formula, Represents model input Next Prediction Category The conditional probability distribution; This represents a hyperparameter for penalizing low-entropy distributions, used to control the strength of the entropy regularization term; Indicates the entropy threshold; Represents the negative entropy term; S3.1.3, The deviation of the output distribution from the uniform distribution is penalized through label smoothing operation to smooth the distribution; the expression is as follows: In the formula, Represents model input Next Prediction Category The conditional probability distribution; This represents the Kullback-Leibler divergence, used to measure the difference between two probability distributions; It indicates a uniform distribution.
5. The method for structural information enhancement and multimodal heterogeneous data fusion characterization according to claim 1, characterized in that, The design structure information optimization framework and soft allocation mechanism are used to improve the quality of data representation and discriminative power. The expression of the soft allocation mechanism is as follows: In the formula, and Represents vertices and vertex The feature vectors are derived from the row vectors of the multimodal feature matrix; Represents the inner product of eigenvectors; This represents the temperature parameter, used to control the weight distribution; Represents vertices The set of neighbors, among which, It is the vertex The index of the neighbor set; Represents vertices Neighboring vertices eigenvectors; superscript Indicates transpose; Represents a node and nodes The connection weights between them.
6. The multimodal heterogeneous data fusion and characterization method with enhanced structural information according to claim 1, characterized in that, The steps for constructing a hypergraph structure and calculating its structural entropy, based on the obtained graph structure, the discriminative representation learning framework with structural entropy regularization, the structural information optimization framework, and the soft allocation mechanism, are as follows: S5.1 Define the hypergraph structure; S5.2 Define the cut edges and cut edge volumes of the hypergraph structure; S5.3, Generate the incidence matrix of the hypergraph structure; S5.4 Construct a clique adjacency matrix based on the generated incidence matrix; S5.5 Calculate the hypergraph entropy and cut edge entropy of the hypergraph structure. The expression for the hypergraph entropy is as follows: In the formula, Represents the hypergraph entropy; Represents the vertices in a hypergraph structure. The set of neighbors; Represents vertices in a hypergraph structure ; Represents vertices in a hypergraph structure and vertex Soft links.
7. The multimodal heterogeneous data fusion and characterization method with enhanced structural information according to claim 5, characterized in that, The hypergraph structure is defined as follows: S5.1 Define the hypergraph structure; Define an undirected weighted hypergraph This undirected weighted hypergraph contains a set of vertices. Hyperedge set and weight matrix The specific definition is as follows: S5.1.1 Constructing the Vertex Set of the Hypergraph ; The set of vertices of a hypergraph The set of vertices in the graph structure corresponds to the data samples, i.e. The row vector; S5.1.2 Constructing the set of hyperedges of a hypergraph ; The set of superedges of a hypergraph Each hyperedge connects a set of related hypergraph vertices. The method for constructing hyperedges is as follows: For each vertex in the feature space of the hypergraph structure, the K-nearest neighbor algorithm is used to calculate cross-modal similar neighbors; S5.1.3 Constructing the weight matrix of the hypergraph ; Hypergraph weight matrix The weight of each hyperedge is stored using the following expression: In the formula, Indicates the super edge. ; Represents the two vertices of a superedge; Representing vertices in the feature space and ; Represents vertices and Distance in feature space; This represents the bandwidth parameter of the Gaussian kernel function.
8. The multimodal heterogeneous data fusion and characterization method with enhanced structural information according to claim 1, characterized in that, The steps for guiding multimodal unsupervised clustering based on hypergraph structure entropy include: S6.1, Multimodal data fusion; S6.2, Structural entropy regularization-guided clustering, the expression is as follows: In the formula, Represents the basic clustering loss function; The regularization coefficient representing the balance structure entropy term and the basic loss; Represents vertices in a hypergraph structure structural entropy S6.3 Clustering result optimization; The clustering results are optimized by iteratively optimizing the cluster centers and cluster labels until the minimum convergence condition is met. The iterative optimization process can be represented as: In the formula, Indicates the first During the nth iteration, the 1st The center vectors of each cluster; This represents the set of data points belonging to the i-th cluster; Represents the vertices of a hypergraph structure eigenvectors.
Citation Information
Patent Citations
Multi-modal evolution feature automatic conformal representation method based on dynamic hypergraph network
CN113254729A
Vehicle and method for controlling the same
KR1020260078964A