An image processing method based on multi-source data fusion

Through sparse subspace representation, space-time alignment and high-order tensor decomposition, combined with dynamic weight allocation, the problem of insufficient correlation between multimodal data is solved, and efficient and accurate low-dimensional fusion representation is generated, which is suitable for image processing tasks.

CN119942284BActive Publication Date: 2025-07-18MIRROR VISION (ZHEJIANG) TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510030308.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-07-18
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

When processing multimodal data in high-dimensional feature, it has high computational complexity, a lot of redundant information, and insufficient correlation between modals, resulting in low efficiency and accuracy of the fusion model, and uneven contributions to the target task of different modalities, making it difficult for existing methods to dynamically adjust the modal weights.

Method used

Through sparse subspace representation modeling, space-time alignment and high-order tensor decomposition, combined with dynamic weight allocation, a subset of sparse features is extracted and space-time alignment is performed, high-order tensors are constructed for fusion, and the modal weights are dynamically adjusted to generate a low-dimensional compact fusion representation.

Benefits of technology

It realizes efficiently reducing multimodal data dimensions, improving computing efficiency, enhancing model generalization capabilities, accurately modeling high-order correlations between modes, dynamically utilizes key modal features, weakening the influence of redundant modalities, and generating high-quality low-dimensional embedded representations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942284B_ABST
    Figure CN119942284B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of image processing technology, and discloses an image processing method based on multi-source data fusion, including: Step 1, obtaining multi-modal image data, representing the data of each modality as a feature matrix, the feature matrix including the feature dimension and the number of samples of the modality, and performing standardization and normalization processing on the feature matrix. At the same time, noise reduction processing is performed on the feature matrix to remove abnormal points and noise signals in the data; Step 2, for the feature matrix obtained and preprocessed in Step 1, assuming that the modality data all share a low-dimensional subspace, the feature matrix is modeled by sparse subspace representation. Through sparse subspace representation modeling, a sparse regularization objective function is introduced to perform sparse feature extraction on the feature matrix, realizing the extraction of key features in high-dimensional multi-modal data and the elimination of redundant information, and obtaining the effects of reducing the curse of dimensionality, improving the calculation efficiency, and enhancing the generalization ability of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and specifically provides an image processing method based on multi-source data fusion. Background Art

[0002] With the rapid development of modern information technology, multi-modal data has been widely used in fields such as remote sensing, medical imaging, security monitoring, and intelligent driving. Multi-modal data comes from different sensors or data acquisition devices, such as optical images, radar images, infrared images, and lidar point clouds. Different modalities of data provide rich and complementary feature information. However, there are significant differences in the feature dimension, time scale, spatial distribution, and expression form of multi-modal data. How to effectively fuse multi-source data to generate a high-quality joint representation has become a key technical issue in the field of image processing.

[0003] Currently, the technical problems faced by multi-source data fusion methods in practical applications are as follows:

[0004] Multi-modal data usually has high-dimensional features. Directly processing high-dimensional data will lead to a significant increase in computational complexity, thereby introducing redundant information and reducing the efficiency and accuracy of the fusion model.

[0005] During the acquisition process of multi-modal data, differences in time and space dimensions may occur due to device characteristics or acquisition conditions. These spatio-temporal differences will prevent the alignment of features of different modalities and affect the data fusion effect.

[0006] Insufficient modeling ability of inter-modal correlation: Multi-modal data not only provides feature information within modalities but also reflects deeper semantic information through the correlation between modalities. Current traditional fusion methods are difficult to effectively capture the high-order correlation between modalities, resulting in insufficient expressiveness of the fusion model.

[0007] In practical scenarios, the contribution degrees of different modalities to the target task are usually unbalanced. The features of redundant modalities will interfere with the fusion result, and fusion methods that fail to dynamically adjust the modality weights cannot fully play the role of key modalities.

[0008] Therefore, those skilled in the art provide an image processing method based on multi-source data fusion to solve the above-mentioned problems. Summary of the Invention

[0009] Aiming at the deficiencies of the prior art, the present invention provides an image processing method based on multi-source data fusion to solve the problems raised in the above background art.

[0010] To achieve the above objectives, the present invention is realized through the following technical solutions: An image processing method based on multi-source data fusion, comprising:

[0011] Step 1: Obtain multi-modal image data, represent the data of each modality as a feature matrix. The feature matrix contains the feature dimension and the number of samples of the modality, and perform standardization and normalization processing on the feature matrix. At the same time, perform noise reduction processing on the feature matrix to remove outliers and noise signals in the data;

[0012] Step 2: For the feature matrix obtained and preprocessed in Step 1, assuming that the modality data shares a low-dimensional subspace, model the feature matrix through sparse subspace representation. By optimizing the objective function that includes the reconstruction error and the sparse regularization term, extract the sparse feature subset in the feature matrix. At the same time, screen the feature subset by evaluating the importance of the sparse features;

[0013] Step 3: Perform spatio-temporal alignment processing on the sparse feature subset extracted in Step 2. For the spatial and temporal differences existing in the multi-modal data, use the dynamic time warping method to synchronize the time dimension. At the same time, for the spatial transformation existing in different modalities, use the registration algorithm to map the feature subset to a unified coordinate system;

[0014] Step 4: Based on the aligned sparse feature subset in Step 3, construct a high-order tensor. Each dimension of the high-order tensor corresponds to the sparse representation of the modality features. Model the high-order tensor through a tensor decomposition model. The tensor decomposition model includes a core tensor and modality factor matrices. By constructing an optimization objective that includes the tensor decomposition reconstruction error and the sparse regularization term, jointly optimize the model parameters of the tensor decomposition to obtain the core tensor and modality factor matrices that can represent the joint features of the multi-modal data;

[0015] Step 5: Based on the core tensor and modality factor matrices obtained in Step 4, combined with the sparsity and correlation of the modality features, assign dynamic weights to each modality feature. The assignment of the dynamic weights is based on the contribution degree of the modality features to the target task, and fuse the modality features in a weighted manner to form a preliminary low-dimensional embedded representation;

[0016] Step 6: Combine the modality features fused in Step 5 with the core tensor and modality factor matrices obtained in Step 4 to generate the final low-dimensional compact fusion representation. The fusion representation has the feature structure of low dimension, sparsity and compactness, and the fusion representation can be directly used as input for downstream image processing tasks.

[0017] Preferably, the sparse subspace representation modeling in Step 2 includes the following optimization objective function:

[0018]

[0019] where, X i represents the feature matrix of modality i, and P i represents the projection matrix of modality i.

[0020] X i -P i U represents the projection matrix P i and the shared subspace U for the feature matrix X i approximation of

[0021] ||P i ||1 represents the sparse regularization term of the projection matrix P i , m represents the number of modalities,

[0022] # represents the sparse regularization parameter, represents the reconstruction error term of modality i.

[0023] Preferably, the projection matrix P in step 2 i is updated by iterative calculation through the following formula:

[0024]

[0025] where soft is the soft threshold function, is the learning rate, represents the gradient with respect to P i of

[0026] t is the number of iterations, # represents the sparse regularization parameter, represents the updated result of the sparse projection matrix of the i-th modality after the (t + 1)-th iteration,

[0027] represents the updated result of the sparse projection matrix of the i-th modality after the t-th iteration.

[0028] Preferably, the spatio-temporal alignment in step 3 adopts the following method:

[0029] Synchronization in the time dimension is achieved through the dynamic time warping method, which calculates the optimal alignment path of time series between different modalities;

[0030] Alignment in the space dimension adopts the geometric registration method based on affine transformation to map the feature subsets of different modalities into a unified spatial coordinate system, and the affine transformation matrix is obtained by minimizing the objective function.

[0031] Preferably, the geometric registration objective function of the affine transformation is:

[0032]

[0033] where T(y i ) is the affine transformation function that maps the input feature y i to the same space,

[0034] y' i is the registered target feature coordinate, || || 2 represents the Euclidean distance

[0035] n represents the total number of feature vectors in the modal data, y i represents the i-th feature vector in the original modal data

[0036] Preferably, the tensor decomposition model includes a core tensor and a modal factor matrix A m , satisfying the following representation

[0037]

[0038] where represents the core tensor, used to model the multi-modal joint features

[0039] A m represents the modal factor matrix represents the sparse feature subset of the multi-modal data, and m represents the number of modalities

[0040] Preferably, the optimization objective function of the tensor decomposition model is

[0041]

[0042] where represents the reconstruction error of the tensor decomposition

[0043] represents the sparse regularization term of the core tensor

[0044] B is the sparse regularization coefficient, used to control the trade-off between sparsity and reconstruction error

[0045] core tensor and modal factor matrix A m , represents the sparse feature subset of the multi-modal data, A i is the factor matrix of modality i, and m represents the number of modalities

[0046] Preferably, the dynamic weight allocation in step 5 is based on the importance of the modal features and is represented by the following formula

[0047]

[0048] where w i is the dynamic weight of modality i, Y i is the feature contribution score of modality i, m represents the number of modalities, and exp is the exponential function

[0049] Preferably, the low-dimensional compact fusion representation generated in step 6 is used for image processing tasks, including image classification, object detection, and image segmentation.

[0050] Preferably, the dynamic weight assignment in step 5 further combines a sparse regularization model to dynamically adjust the importance of modal features, and the adaptive assignment of dynamic weights is achieved through the following optimization objective function:

[0051]

[0052] where w i is the dynamic weight of modality i, is the feature tensor of modality i, is the core tensor, A m is the factor matrix of modality m;

[0053] is the tensor decomposition reconstruction error of modality i,

[0054] a is the sparse regularization coefficient, ||w||1 is the sparse regularization term of the dynamic weight vector, and m is the number of modalities.

[0055] The present invention provides an image processing method based on multi-source data fusion. It has the following beneficial effects:

[0056] 1. Through sparse subspace representation modeling, the present invention introduces a sparse regularization objective function to perform sparse feature extraction on the feature matrix, realizes the extraction of key features and the elimination of redundant information in high-dimensional multi-modal data, and obtains the effects of reducing the curse of dimensionality, improving computational efficiency, and enhancing the generalization ability of the model.

[0057] 2. By introducing spatio-temporal alignment methods, including dynamic time warping and geometric registration based on affine transformation, the present invention realizes the synchronization of multi-modal data in the time and space dimensions, obtains the effect of consistent expression of multi-modal features under a unified spatio-temporal framework, and provides a data basis for subsequent fusion and modeling.

[0058] 3. Through high-order tensor decomposition modeling, the present invention integrates the aligned multi-modal feature subsets into a high-order tensor, and combines the joint optimization of the core tensor and the modal factor matrix to realize the accurate modeling of high-order correlations between modalities, and obtains the effect of more expressive and consistent multi-modal joint features.

[0059] 4. Through the dynamic weight assignment mechanism, combined with the evaluation of the contribution of modal features to the target task, the present invention assigns weights according to the contribution ratio and performs weighted fusion on modal features, realizes the efficient utilization of key modalities and the weakening of the influence of redundant modalities, and obtains the effect of more accurate and robust low-dimensional embedded representation. Description of the Drawings

[0060] Figure 1 This is the flowchart of the present invention. Detailed implementation manners

[0061] To enable those skilled in the art to understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.

[0062] The following will describe the present invention in detail with reference to the accompanying drawings:

[0063] Embodiment:

[0064] Please refer to the attached Figure 1 , the embodiment of the present invention provides an image processing method based on multi-source data fusion, including:

[0065] Step 1: Obtain multi-modal image data, represent the data of each modality as a feature matrix, the feature matrix includes the feature dimension and the number of samples of the modality, and perform standardization and normalization processing on the feature matrix. At the same time, perform noise reduction processing on the feature matrix to remove outliers and noise signals in the data;

[0066] Step 2: For the feature matrix obtained and preprocessed in Step 1, assume that the modality data all share a low-dimensional subspace, represent the feature matrix through sparse subspace representation modeling, and extract the sparse feature subset in the feature matrix by optimizing the objective function including the reconstruction error and the sparse regularization term. At the same time, screen the feature subset by evaluating the importance of the sparse features;

[0067] Step 3: Perform spatio-temporal alignment processing on the sparse feature subset extracted in Step 2. For the spatial and temporal differences existing in the multi-modal data, use the dynamic time warping method to synchronize the time dimension. At the same time, for the spatial transformation existing in different modalities, use the registration algorithm to map the feature subset to a unified coordinate system;

[0068] Step 4: Based on the aligned sparse feature subset in Step 3, construct a high-order tensor. Each dimension of the high-order tensor corresponds to the sparse representation of the modality features. Model the high-order tensor through a tensor decomposition model. The tensor decomposition model includes a core tensor and modality factor matrices. By constructing an optimization objective including the tensor decomposition reconstruction error and the sparse regularization term, jointly optimize the model parameters of the tensor decomposition to obtain the core tensor and modality factor matrices that can characterize the joint features of the multi-modal data;

[0069] Step 5: Based on the core tensor and modal factor matrices obtained in Step 4, combined with the sparsity and correlation of modal features, assign dynamic weights to each modal feature. The assignment of dynamic weights is based on the contribution degree of the modal feature to the target task, and fuse the modal features through a weighted manner to form a preliminary low-dimensional embedded representation;

[0070] Step 6: Combine the fused modal features in Step 5 with the core tensor and modal factor matrices obtained in Step 4 to generate a final low-dimensional compact fused representation. The fused representation has the characteristic structure of low dimension, sparsity, and compactness, and can be directly used as input for downstream image processing tasks.

[0071] Benefits of Step 1: Through standardization and normalization processing, eliminate the distribution differences caused by different dimensions and numerical ranges between different modal data. Through noise reduction processing, remove outliers and noise signals, improve the quality and consistency of the data, and provide input data for subsequent modeling;

[0072] Benefits of Step 2: Sparse subspace representation modeling can extract discriminative key features in high-dimensional data, remove redundant features through sparse regularization, significantly reduce the data dimension, reduce computational complexity and storage overhead, and improve the effectiveness of data representation and the accuracy of downstream tasks by evaluating the importance of features and screening feature subsets;

[0073] Benefits of Step 3: Solve the out-of-sync problem of multimodal data in the time dimension through the dynamic time warping method, and solve the coordinate differences of different modalities in the spatial dimension through the geometric registration method, ensuring the consistency of multimodal data in time and space, and providing a feature expression basis for subsequent fusion and joint modeling;

[0074] Benefits of Step 4: Integrate the sparse feature subsets after spatio-temporal alignment into a high-order tensor, model the high-order correlations between modalities through the tensor decomposition method. The core tensor can characterize the complex semantic relationships between multimodal data, while the modal factor matrices are used to represent the low-dimensional embedded features unique to each modality. By jointly optimizing the tensor decomposition parameters, achieve the expression of semantic consistency between modalities and enhance the expressiveness of the data fusion model;

[0075] Benefits of Step 5: Through the dynamic weight assignment mechanism, based on the evaluation of the contribution degree of modal features to the target task, reasonably assign the weights of modalities, highlight the role of key modalities and weaken the interference of redundant modalities, generate a preliminary low-dimensional embedded representation through weighted fusion, improve the robustness and reliability of feature fusion, and lay a foundation for the generation of the final fused representation;

[0076] Benefits of Step 6: By combining the dynamic weight assignment results with the core tensor and factor matrices generated by the tensor decomposition model, the generated low-dimensional compact fusion representation is both sparse and has good expressive power. The fusion representation retains the most important feature information for the target task, while reducing redundancy and improving the efficiency and performance of downstream image processing tasks.

[0077] The sparse subspace representation modeling in Step 2 includes the following optimization objective function:

[0078]

[0079] where X i represents the feature matrix of modality i, P i represents the projection matrix of modality i,

[0080] X i -P i U represents the approximation of the projection matrix P i and the shared subspace U to the feature matrix X i ,

[0081] ||P i ||1 represents the sparse regularization term of the projection matrix P i , m represents the number of modalities,

[0082] # represents the sparse regularization parameter, represents the reconstruction error term of modality i.

[0083] By optimizing the objective function, the sparse subspace representation modeling can extract the effective key features in the modality feature matrix, and at the same time, through the sparse regularization term, it can constrain the projection matrix to reduce the interference of redundant features. The sparsity constraint ensures that the most discriminative features for downstream tasks are retained, while significantly reducing the feature dimension and computational complexity.

[0084] The reconstruction error term included in the optimization objective function can ensure the efficient approximation of the projection matrix and the shared subspace to the original feature matrix, retain the main information content in the feature matrix, ensure that the low-dimensional subspace can express the modality features, and provide data input for subsequent analysis and modeling.

[0085] By introducing the sparse regularization parameter, it is possible to assign weights to the important features in the projection matrix, weaken the influence of irrelevant and noisy features, improve the performance of the model during training, and enhance the generalization ability in different datasets and task scenarios.

[0086] In summary, by combining the reconstruction error term and the sparse regularization term, the optimization objective function achieves efficient feature selection, low-dimensional approximate representation, and unified feature representation in the sparse subspace representation modeling. The specific benefits include:

[0087] Reduce the feature dimension and decrease the computational complexity;

[0088] Extract key features and remove redundancy and noise;

[0089] Ensure the accuracy of reconstruction and improve the quality of feature representation;

[0090] Enhance the generalization ability of the model in different data scenarios;

[0091] Provide a consistent feature representation framework for subsequent spatio-temporal alignment and fusion steps.

[0092] Generally speaking, this optimized objective function provides a balance in terms of sparsity, reconstruction ability, and computational efficiency, solves the key problems in multi-modal high-dimensional feature representation, and is one of the core technical links in multi-source data fusion.

[0093] The projection matrix P in Step 2 i is updated through iterative calculation by the following formula:

[0094]

[0095] where soft is the soft threshold function, is the learning rate, represents the gradient with respect to P i and

[0096] t is the number of iterations, # represents the sparse regularization parameter, represents the updated result of the sparse projection matrix of the i-th modality after the (t + 1)-th iteration,

[0097] represents the updated result of the sparse projection matrix of the i-th modality after the t-th iteration.

[0098] Through iterative calculation, the projection matrix approaches the optimal solution, and each update is optimized along the direction of the gradient of the objective function to ensure the convergence of the sparse subspace representation model. As the iteration progresses, the value of the projection matrix tends to be stable, and a projection matrix that can effectively extract the key features of the modality is obtained.

[0099] The combination of sparse regularization and the soft threshold function enables the updated projection matrix to automatically weaken the influence of noise and redundant features, improve the robustness to high-noise data and multi-modal data, and enhance the stability and reliability of the sparse subspace representation model.

[0100] In summary, the update formula of the projection matrix provides an efficient and stable method for determining sparse features by introducing gradient descent, the soft threshold function, and sparse regularization. The main benefits include:

[0101] Ensure the convergence of the iterative process;

[0102] Implement sparse feature selection to retain the features that contribute most to the task;

[0103] Improve the optimization efficiency and reduce the computational complexity;

[0104] Enhance the robustness and adaptability of the model in high-dimensional and multi-modal data scenarios.

[0105] Play a core role in sparse subspace representation modeling and provide feature inputs for subsequent spatio-temporal alignment and fusion of multi-modal data.

[0106] The spatio-temporal alignment in Step 3 adopts the following method:

[0107] Synchronization in the time dimension is achieved through the dynamic time warping method, which calculates the optimal alignment path of time series between different modalities;

[0108] Alignment in the space dimension adopts a geometric registration method based on affine transformation to map the feature subsets of different modalities into a unified space coordinate system, and the affine transformation matrix is obtained by minimizing the objective function.

[0109] The objective function of the geometric registration of affine transformation is:

[0110]

[0111] where T(y i ) is the affine transformation function that maps the input feature y i to the same space,

[0112] y′ i is the target feature coordinate after registration, || || 2 represents the Euclidean distance,

[0113] n represents the total number of feature vectors in the modal data, and y i represents the i-th feature vector in the original modal data.

[0114] Benefits of the dynamic time warping method in the time dimension:

[0115] The dynamic time warping method can effectively handle problems such as differences in time series sampling rates and data lengths between different modalities. By calculating the optimal alignment path of the time series, it determines the most matching time points between modal data to achieve synchronization in the time dimension.

[0116] Dynamic time warping considers the global structure and local variations of the data, can capture complex dynamic relationships of time series, and makes different modal data highly consistent in the time dimension, providing a time alignment basis for subsequent data fusion.

[0117] Benefits of the geometric registration method in the spatial dimension:

[0118] Data from different modalities come from different acquisition devices or perspectives, resulting in differences in the spatial distribution of features. Geometric registration maps different modality feature subsets into a unified coordinate system through affine transformation, eliminating the spatial differences between modalities and making the feature representation more unified.

[0119] Through geometric operations such as rotation, scaling, and translation, affine transformation can accurately align the input modality features to the target coordinate system. At the same time, it minimizes the alignment error to ensure that the registered features can strictly match the target features.

[0120] Benefits of the objective function of the geometric registration of affine transformation:

[0121] By minimizing the global error, the registration process can balance the impact of local feature differences on the alignment accuracy, making the result insensitive to single-point errors or abnormal features, and improving the overall robustness of the alignment.

[0122] The affine transformation objective function directly optimizes the alignment error, making the features of different modalities have spatial deviations after being mapped to the same coordinate system, thereby ensuring the accuracy and consistency of subsequent feature fusion.

[0123] In summary, the spatio-temporal alignment method in step 3 realizes the unified expression of multi-modal data in the time and space dimensions through dynamic time warping and geometric registration, and has the following advantages:

[0124] The dynamic time warping method solves the problem of asynchronous multi-modal data in the time dimension and improves the accuracy of time feature alignment;

[0125] Geometric registration eliminates the spatial differences between modalities through affine transformation to ensure the consistency of features in the spatial dimension;

[0126] The affine transformation objective function minimizes the Euclidean distance, reduces the spatial alignment error to the lowest level, and improves the accuracy and robustness of registration.

[0127] Generally speaking, the spatio-temporal alignment process provides a unified and aligned spatio-temporal feature expression framework for multi-modal data, provides support for subsequent feature fusion and joint modeling, and at the same time significantly enhances the applicability and generalization ability of the model.

[0128] The tensor decomposition model includes a core tensor and a modal factor matrix A m , satisfying the following representation:

[0129]

[0130] Among them, represents the core tensor, which is used to model multi-modal joint features,

[0131] A m represents the modal factor matrix, represents the sparse feature subset of multimodal data, and m represents the number of modalities.

[0132] The optimization objective function of the tensor decomposition model is:

[0133]

[0134] where, represents the reconstruction error of the tensor decomposition;

[0135] represents the sparse regularization term of the core tensor;

[0136] B is the sparse regularization coefficient, which is used to control the trade-off between sparsity and reconstruction error,

[0137] core tensor and the modal factor matrix A m , represents the sparse feature subset of multimodal data, and A i is the factor matrix of modality i, and m represents the number of modalities.

[0138] The tensor decomposition model characterizes the joint features among multimodal data through the core tensor, and can capture the complex high-order correlations between different modal features. Compared with the traditional feature concatenation method, the tensor decomposition method can express the semantic relationship between modalities, making the joint representation of multimodal features accurate.

[0139] Reduce feature redundancy and improve feature compactness. The modal factor matrix extracts the low-dimensional embedded representation of each modality. At the same time, combined with the sparse regularization constraint, redundant and useless information can be removed, making the feature representation of each modality concise and compact, reducing the computational complexity, and providing feature input for subsequent fusion and applications.

[0140] Flexibly adapt to multimodal data scenarios. The tensor decomposition model can adapt to any number of modal data. The model can flexibly adjust through the core tensor and factor matrix to adapt to different types and numbers of multimodal data scenarios, and has good scalability.

[0141] In summary, through the tensor decomposition model and the optimization objective function, the following main benefits can be achieved:

[0142] Capture the high-order correlations of multimodal data and enhance the expression ability of multimodal joint features;

[0143] Remove redundant features and make the feature representation more compact and efficient;

[0144] Enhance the robustness of the model to noise and outliers through sparse regularization;

[0145] Ensure the accuracy of feature representation by combining the reconstruction error;

[0146] Adapt to the diversity and scalability of multimodal data;

[0147] Achieve a unified low-dimensional representation of multimodal features and solve the problem of inter-modal inconsistency.

[0148] Generally speaking, the tensor decomposition model plays a key role in the joint modeling of high-dimensional multimodal data, provides feature input for subsequent dynamic weight assignment and the generation of the final fusion representation, and significantly improves the efficiency and performance of the image processing pipeline.

[0149] The dynamic weight assignment in step 5 is based on the importance of modal features and is represented by the following formula:

[0150]

[0151] where w i is the dynamic weight of modality i, Y i is the feature contribution score of modality i, m represents the number of modalities, and exp is the exponential function.

[0152] The low-dimensional compact fusion representation generated in step 6 is used for image processing tasks, including image classification, object detection, and image segmentation.

[0153] The dynamic weight assignment in step 5 and the low-dimensional compact fusion representation in step 6 jointly support the efficient fusion of multi-source data, which is mainly reflected as follows:

[0154] The dynamic weight assignment mechanism adaptively adjusts the modal weights through the feature contribution degree, highlights the role of key modalities, weakens the interference of redundant modalities, and ensures the accuracy and robustness of the multimodal fusion features.

[0155] The low-dimensional compact fusion representation combines the dynamic weight and the tensor decomposition result to generate an efficient and sparse joint feature representation, significantly reducing the data dimension, improving the computing efficiency, and at the same time, meeting the requirements of semantic consistency and compactness of feature expression for downstream tasks.

[0156] Generally speaking, the dynamic weight assignment and the low-dimensional compact fusion representation are the core links of the multimodal data fusion method. Through the organic combination of the two, the present invention has achieved significant advantages in terms of the accuracy, efficiency, and adaptability of feature fusion, providing technical support for multimodal image processing tasks.

[0157] The dynamic weight assignment in step 5 further combines the sparse regularization model to dynamically adjust the importance of modal features, and realizes the adaptive assignment of dynamic weights through the following optimization objective function:

[0158]

[0159] Among them, w i is the dynamic weight of modality i, is the feature tensor of modality i, is the core tensor, A m is the factor matrix of modality m;

[0160] is the tensor decomposition reconstruction error of modality i,

[0161] a is the sparse regularization coefficient, ||w||1 is the sparse regularization term of the dynamic weight vector, and m is the number of modalities.

[0162] The dynamic weight assignment method combined with sparse regularization has the following remarkable benefits in multi-modal data fusion:

[0163] By adaptively adjusting the weights and sparse constraints, it strengthens the utilization of key modalities, and at the same time, weakens the influence of redundant modalities.

[0164] Combined with the optimization of tensor decomposition reconstruction error, it ensures that the generated fusion features have higher semantic consistency and discriminative power.

[0165] Sparse regularization reduces the interference of abnormal modalities and optimizes the computational efficiency at the same time, making the multi-modal data fusion more stable and efficient.

[0166] The dynamic weight assignment mechanism can flexibly cope with the diversity of different modal tasks and data scenarios, ensuring the wide applicability of the model.

[0167] Generally speaking, the dynamic weight assignment method combined with sparse regularization effectively solves the problems of weight assignment and redundant modality interference in multi-modal data fusion, provides an important guarantee for generating high-quality and low-dimensional compact fusion features, and is one of the key innovative links in multi-source data fusion technology.

[0168] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An image processing method based on multi-source data fusion, characterized in that, Including: Step 1: Obtain multi-modal image data, represent the data of each modality as a feature matrix. The feature matrix contains the feature dimension and the number of samples of the modality, and perform standardization and normalization processing on the feature matrix. At the same time, perform noise reduction processing on the feature matrix to remove outliers and noise signals in the data; Step 2: For the feature matrix obtained and preprocessed in Step 1, assume that the modality data shares a low-dimensional subspace. Model the feature matrix through sparse subspace representation, and extract the sparse feature subset in the feature matrix by optimizing the objective function including the reconstruction error and the sparse regularization term. At the same time, screen the feature subset by evaluating the importance of the sparse features; Step 3: Perform spatio-temporal alignment processing on the sparse feature subset extracted in Step 2. For the spatial and temporal differences existing in the multi-modal data, use the dynamic time warping method to synchronize the time dimension. At the same time, for the spatial transformation existing in different modalities, use the registration algorithm to map the feature subset to a unified coordinate system; Step 4: Based on the aligned sparse feature subset in Step 3, construct a high-order tensor. Each dimension of the high-order tensor corresponds to the sparse representation of the modality features. Model the high-order tensor through a tensor decomposition model. The tensor decomposition model includes a core tensor and modality factor matrices. By constructing an optimization objective including the tensor decomposition reconstruction error and the sparse regularization term, jointly optimize the model parameters of the tensor decomposition to obtain the core tensor and modality factor matrices that can characterize the joint features of the multi-modal data; Step 5: Based on the core tensor and modality factor matrices obtained in Step 4, combine the sparsity and correlation of the modality features, assign dynamic weights to each modality feature. The assignment of the dynamic weights is based on the contribution degree of the modality features to the target task, and fuse the modality features in a weighted manner to form a preliminary low-dimensional embedded representation; Step 6: Combine the fused modality features in Step 5 with the core tensor and modality factor matrices obtained in Step 4 to generate a final low-dimensional compact fusion representation. The fusion representation has a feature structure with low dimension, sparsity, and compactness, and the fusion representation can be directly used as input for downstream image processing tasks.

2. The image processing method based on multi-source data fusion according to claim 1, wherein The sparse subspace representation modeling in Step 2 includes the following objective function: Among them, X i represents the feature matrix of mode i, and P i represents the projection matrix of mode i. X i - P i U represents the projection matrix P i and the shared subspace U for the feature matrix X i approximation of ||P i ||1 represents the projection matrix P i with a sparse regularization term, m represents the number of modalities, # represents the sparse regularization parameter, represents the reconstruction error term of modality i.

3. The image processing method based on multi-source data fusion according to claim 2, characterized in that The projection matrix P in step 2 i is updated through iterative calculation using the following formula: where soft is the soft threshold function, is the learning rate, represents the gradient with respect to P i and t is the number of iterations, and # represents the sparse regularization parameter. represents the updated result of the sparse projection matrix of the i-th modality after the (t + 1)-th iteration. Denote the updated result of the sparse projection matrix of the $i$-th modality after the $t$-th iteration.

4. A method for image processing based on multi-source data fusion according to claim 1, characterized in that, The spatio-temporal alignment in Step 3 adopts the following method: The synchronization in the time dimension is achieved through the dynamic time warping method, and the dynamic time warping method calculates the optimal alignment path of the time series between different modalities; The alignment in the spatial dimension adopts the geometric registration method based on affine transformation to map the feature subsets of different modalities to a unified spatial coordinate system, and the affine transformation matrix is obtained by minimizing the objective function.

5. The image processing method based on multi-source data fusion according to claim 4, characterized in that The geometric registration objective function of the affine transformation is: Among them, T(y i ) is an affine transformation function that maps the input feature y i to the same space. y′ i are the coordinates of the registered target features, || || 2 represents the Euclidean distance n represents the total number of eigenvectors in the modal data, and y i represents the i-th eigenvector in the original modal data.

6. The image processing method based on multi-source data fusion according to claim 1, wherein The tensor decomposition model includes a core tensor and a modal factor matrix A m , satisfying the following representation: Among them, represents the core tensor for modeling multi-modal joint features, A m represents the modal factor matrix, represents the sparse feature subset of multi-modal data, and m represents the number of modalities.

7. The image processing method based on multi-source data fusion according to claim 6, characterized in that The optimization objective function of the tensor decomposition model is: Among them, represents the reconstruction error of tensor decomposition; Represents the sparse regularization term of the core tensor; B is the sparse regularization coefficient, which is used to control the trade-off between sparsity and reconstruction error, Core tensor and modal factor matrix A m , represents a sparse feature subset of multimodal data, and A i is the factor matrix of modality i, and m represents the number of modalities.

8. An image processing method based on multi-source data fusion according to claim 1, characterized in that, The dynamic weight assignment in Step 5 is based on the importance of the modality features and is represented by the following formula: Among them, w i is the dynamic weight of mode i, Y i is the feature contribution score of mode i, m represents the number of modes, and exp is the exponential function.

9. A method for image processing based on multi-source data fusion according to claim 1, characterized in that The generated low-dimensional compact fusion representation in Step 6 is used for image processing tasks, including image classification, object detection, and image segmentation.

10. An image processing method based on multi-source data fusion according to claim 1, characterized in that The dynamic weight assignment in the above step 5 further combines a sparse regularization model to dynamically adjust the importance of modal features, and the adaptive assignment of dynamic weights is achieved through the following optimization objective function: Among them, w i is the dynamic weight of mode i, is the feature tensor of mode i, is the core tensor, A m is the factor matrix of mode m; is the tensor decomposition reconstruction error for modality i, a is the sparse regularization coefficient, ||w||1 is the sparse regularization term of the dynamic weight vector, and m is the number of modalities.

Citation Information

Patent Citations

  • Medical image fusion method based on shared multi-dimensional component tensor dictionary learning

    CN117934305A

  • High-precision surveying and mapping method based on multi-source data fusion

    CN119203034A