Image processing method based on multi-source data fusion

Through technical means such as sparse subspace representation, space-time alignment, tensor decomposition and dynamic weight allocation, the problem of insufficient correlation between high-dimensional data processing in multi-source data fusion is solved, and efficient and accurate multi-modal data fusion is achieved.

CN119942284AActive Publication Date: 2025-05-06MIRROR VISION (ZHEJIANG) TECHNOLOGY CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510030308.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-06
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

The existing multi-source data fusion method has high computational complexity, many redundant information, and insufficient correlation between modals when processing high-dimensional multi-modal data, resulting in poor fusion effect.

Method used

Through technical means such as sparse subspace representation modeling, space-time alignment, high-order tensor decomposition, and dynamic weight allocation, sparse features of multimodal data are extracted, space-time alignment is performed, high-order correlations between modes are modeled, and modal weights are dynamically adjusted for fusion.

Benefits of technology

It realizes effective dimensionality reduction of high-dimensional multimodal data, improves computing efficiency, enhances model generalization capabilities, captures high-order correlations between modes, generates more expressive multimodal joint features, and improves the accuracy and robustness of fusion results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942284A_ABST
    Figure CN119942284A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, and discloses an image processing method based on multi-source data fusion, which comprises the following steps: step 1, acquiring multi-modal image data, representing data of each modal as a feature matrix which comprises feature dimensions and sample quantity of the modals, carrying out standardization and normalization processing on the feature matrix, and meanwhile, carrying out data fusion on the feature dimensions and the sample quantity of the modals; performing noise reduction processing on the feature matrix, and removing abnormal points and noise signals in the data; 2, for the feature matrix obtained and preprocessed in the step 1, assuming that modal data share a low-dimensional subspace, and modeling the feature matrix through sparse subspace representation; through sparse subspace representation modeling, a sparse regularization objective function is introduced to carry out sparse feature extraction on a feature matrix, extraction of key features in high-dimensional multi-modal data and elimination of redundant information are realized, and the effects of reducing dimension disasters, improving calculation efficiency and enhancing model generalization ability are obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to an image processing method based on multi-source data fusion. Background Art

[0002] With the rapid development of modern information technology, multimodal data has been widely used in remote sensing, medical imaging, security monitoring and intelligent driving. Multimodal data comes from different sensors or data acquisition devices, such as optical images, radar images, infrared images and lidar point clouds. Data of different modalities provide rich and complementary feature information. However, multimodal data has significant differences in feature dimension, time scale, spatial distribution and expression form. How to effectively fuse multi-source data to generate high-quality joint representation has become a key technical issue in the field of image processing.

[0003] At present, the technical difficulties faced by multi-source data fusion methods in practical applications are as follows:

[0004] Multimodal data usually has high-dimensional characteristics. Directly processing high-dimensional data will lead to a significant increase in computational complexity, which will in turn introduce redundant information and reduce the efficiency and accuracy of the fusion model.

[0005] During the acquisition process, multimodal data may cause differences in time and space dimensions due to different equipment characteristics or acquisition conditions. The time and space differences will make it impossible to align the features of different modalities, affecting the data fusion effect.

[0006] Insufficient modeling capability for inter-modal correlation: Multimodal data not only provides feature information within the modality, but also reflects deeper semantic information through the correlation between modalities. Current traditional fusion methods are difficult to effectively capture high-order correlations between modalities, resulting in insufficient expressiveness of the fusion model.

[0007] In actual scenarios, the contribution of different modalities to the target task is usually unbalanced. The characteristics of redundant modalities will interfere with the fusion results, and fusion methods that fail to dynamically adjust the modal weights cannot give full play to the role of key modalities.

[0008] Therefore, those skilled in the art provide an image processing method based on multi-source data fusion to solve the above-mentioned problems. Summary of the invention

[0009] In view of the deficiencies in the prior art, the present invention provides an image processing method based on multi-source data fusion to solve the problems raised in the above background technology.

[0010] To achieve the above objectives, the present invention is implemented through the following technical solutions: an image processing method based on multi-source data fusion, comprising:

[0011] Step 1: Acquire multimodal image data, represent the data of each modality as a feature matrix, wherein the feature matrix includes the feature dimension and sample quantity of the modality, and standardize and normalize the feature matrix. At the same time, perform noise reduction on the feature matrix to remove abnormal points and noise signals in the data;

[0012] Step 2: For the feature matrix obtained and preprocessed in step 1, assume that all modal data share a low-dimensional subspace, model the feature matrix through a sparse subspace representation, and extract a sparse feature subset in the feature matrix by optimizing an objective function including reconstruction error and a sparse regularization term. At the same time, the feature subset is screened by evaluating the importance of the sparse features.

[0013] Step 3: Perform spatiotemporal alignment on the sparse feature subsets extracted in step 2. In view of the spatial and temporal differences in multimodal data, a dynamic time warping method is used to synchronize the time dimension. At the same time, in view of the spatial transformation of different modalities, a registration algorithm is used to map the feature subsets to a unified coordinate system.

[0014] Step 4: Based on the sparse feature subsets aligned in step 3, a high-order tensor is constructed, wherein each dimension of the high-order tensor corresponds to a sparse representation of the modal feature, and the high-order tensor is modeled by a tensor decomposition method. The tensor decomposition model includes a core tensor and a modal factor matrix. By constructing an optimization objective including a tensor decomposition reconstruction error and a sparse regularization term, the model parameters of the tensor decomposition are jointly optimized to obtain a core tensor and a modal factor matrix that can characterize the joint features of multimodal data;

[0015] Step 5: Based on the core tensor and modal factor matrix obtained in step 4, combined with the sparsity and correlation of the modal features, dynamic weights are assigned to each modal feature. The dynamic weights are assigned based on the contribution of the modal features to the target task. The modal features are fused in a weighted manner to form a preliminary low-dimensional embedding representation.

[0016] Step 6: Combine the modal features fused in step 5 with the core tensor and modal factor matrix obtained in step 4 to generate the final low-dimensional compact fusion representation. The fusion representation has a low-dimensional, sparse and compact feature structure. The fusion representation can be directly used as input for downstream image processing tasks.

[0017] Preferably, the sparse subspace representation modeling in step 2 includes the following optimization objective function:

[0018]

[0019] Among them, X i represents the characteristic matrix of mode i, P i represents the projection matrix of mode i,

[0020] X i -P i U represents the projection matrix P i and the shared subspace U for the feature matrix X i The approximation of

[0021] ∥P i ∥1 represents the projection matrix P i The sparse regularization term, m represents the number of modes,

[0022] # represents the sparse regularization parameter, represents the reconstruction error term of mode i.

[0023] Preferably, the projection matrix P in step 2 is i The update is iteratively calculated using the following formula:

[0024]

[0025] Among them, soft is the soft threshold function, learning rate, Indicates about P i The gradient of

[0026] t is the number of iterations, # represents the sparse regularization parameter, represents the update result of the sparse projection matrix of the i-th mode after the t+1th iteration,

[0027] Represents the updated result of the sparse projection matrix of the i-th mode after the t-th iteration.

[0028] Preferably, the spatiotemporal alignment in step 3 adopts the following method:

[0029] Synchronization in the time dimension is achieved through the dynamic time warping method, which calculates the optimal alignment path of time series between different modalities;

[0030] The alignment in the spatial dimension adopts a geometric registration method based on affine transformation to map feature subsets of different modalities into a unified spatial coordinate system. The affine transformation matrix is ​​obtained by minimizing the objective function.

[0031] Preferably, the geometric registration objective function of the affine transformation is:

[0032]

[0033] Among them, T(y i ) is an affine transformation function, which transforms the input feature y i Mapped to the same space,

[0034] y′ i is the target feature coordinate after registration, || || 2 represents the Euclidean distance,

[0035] n represents the total number of eigenvectors in the modal data, y i represents the i-th eigenvector in the original modal data.

[0036] Preferably, the tensor decomposition model includes a core tensor and the modal factor matrix A m , satisfying the following representation:

[0037]

[0038] in, Represents the core tensor used to model multimodal joint features.

[0039] A m represents the modal factor matrix, represents a sparse feature subset of multimodal data, and m represents the number of modalities.

[0040] Preferably, the optimization objective function of the tensor decomposition model is:

[0041]

[0042] in, represents the reconstruction error of tensor decomposition;

[0043] represents the sparse regularization term of the core tensor;

[0044] B is the sparse regularization coefficient, which is used to control the trade-off between sparsity and reconstruction error.

[0045] Core Tensor and the modal factor matrix A m , Represents a sparse feature subset of multimodal data, A i is the factor matrix of mode i, and m represents the number of modes.

[0046] Preferably, the dynamic weight allocation in step 5 is based on the importance of the modal features and is expressed by the following formula:

[0047]

[0048] Among them, w i is the dynamic weight of mode i, Y i is the feature contribution score of mode i, m represents the number of modes, and exp is an exponential function.

[0049] Preferably, the low-dimensional compact fusion representation generated in step 6 is used for image processing tasks, including image classification, object detection and image segmentation.

[0050] Preferably, the dynamic weight allocation in step 5 is further combined with a sparse regularization model to dynamically adjust the importance of modal features, and the adaptive allocation of dynamic weights is achieved by optimizing the following objective function:

[0051]

[0052] Among them, w i is the dynamic weight of mode i, is the characteristic tensor of mode i, is the core tensor, A m is the factor matrix of mode m;

[0053] is the tensor decomposition reconstruction error of mode i,

[0054] a is the sparse regularization coefficient, ∥w∥1 is the sparse regularization term of the dynamic weight vector, and m is the number of modes.

[0055] The present invention provides an image processing method based on multi-source data fusion. It has the following beneficial effects:

[0056] 1. The present invention uses sparse subspace representation modeling and introduces a sparse regularization objective function to perform sparse feature extraction on the feature matrix, thereby realizing the extraction of key features and elimination of redundant information in high-dimensional multimodal data, thereby reducing the curse of dimensionality, improving computational efficiency, and enhancing the generalization ability of the model.

[0057] 2. The present invention introduces a spatiotemporal alignment method, including dynamic time warping and geometric alignment based on affine transformation, to achieve synchronization of multimodal data in time and space dimensions, obtain the effect of consistent expression of multimodal features under a unified spatiotemporal framework, and provide a data basis for subsequent fusion and modeling.

[0058] 3. The present invention integrates the aligned multimodal feature subsets into high-order tensors through high-order tensor decomposition modeling, and combines the joint optimization of the core tensor and the modal factor matrix to achieve accurate modeling of high-order correlations between modalities, thereby obtaining a more expressive and consistent multimodal joint feature effect.

[0059] 4. The present invention adopts a dynamic weight allocation mechanism, combines the evaluation of the contribution of modal features to the target task, allocates weights according to the contribution ratio and performs weighted fusion on the modal features, so as to achieve efficient utilization of key modalities and weaken the influence of redundant modalities, and obtain a more accurate and robust low-dimensional embedding representation effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION

[0061] In order to make the technical personnel in the technical field understand the scheme of the present invention, the technical scheme in the embodiment of the present invention will be clearly and completely described below in combination with the drawings in the embodiment of the present invention. Obviously, the described embodiment is a partial embodiment of the present invention, not a complete embodiment. Based on the embodiment of the present invention, other embodiments obtained by ordinary technicians in the field without creative work should fall within the scope of protection of the present invention.

[0062] The present invention is described in detail below in conjunction with the accompanying drawings:

[0063] Example:

[0064] Please see attached Figure 1 The embodiment of the present invention provides an image processing method based on multi-source data fusion, comprising:

[0065] Step 1: Obtain multimodal image data, represent the data of each modality as a feature matrix, which contains the feature dimension and sample quantity of the modality, and standardize and normalize the feature matrix. At the same time, perform noise reduction on the feature matrix to remove abnormal points and noise signals in the data;

[0066] Step 2: For the feature matrix obtained and preprocessed in step 1, assume that all modal data share a low-dimensional subspace, model the feature matrix through a sparse subspace representation, and extract a sparse feature subset in the feature matrix by optimizing an objective function including reconstruction error and a sparse regularization term. At the same time, the feature subset is screened by evaluating the importance of the sparse features.

[0067] Step 3: Perform spatiotemporal alignment on the sparse feature subsets extracted in step 2. In view of the spatial and temporal differences in multimodal data, a dynamic time warping method is used to synchronize the time dimension. At the same time, in view of the spatial transformation of different modalities, a registration algorithm is used to map the feature subsets to a unified coordinate system.

[0068] Step 4: Based on the sparse feature subsets aligned in step 3, a high-order tensor is constructed. Each dimension of the high-order tensor corresponds to a sparse representation of the modal feature. The high-order tensor is modeled by a tensor decomposition method. The tensor decomposition model includes a core tensor and a modal factor matrix. By constructing an optimization objective including a tensor decomposition reconstruction error and a sparse regularization term, the model parameters of the tensor decomposition are jointly optimized to obtain a core tensor and a modal factor matrix that can characterize the joint features of multimodal data.

[0069] Step 5: Based on the core tensor and modal factor matrix obtained in step 4, combined with the sparsity and correlation of the modal features, dynamic weights are assigned to each modal feature. The dynamic weights are assigned based on the contribution of the modal features to the target task. The modal features are fused in a weighted manner to form a preliminary low-dimensional embedding representation.

[0070] Step 6: Combine the modal features fused in step 5 with the core tensor and modal factor matrix obtained in step 4 to generate the final low-dimensional compact fusion representation. The fusion representation has a low-dimensional, sparse and compact feature structure. The fusion representation can be directly used as input for downstream image processing tasks.

[0071] Benefits of step 1: Standardization and normalization can eliminate the distribution differences between different modal data due to different dimensions and numerical ranges. Noise reduction can remove outliers and noise signals, improve data quality and consistency, and provide input data for subsequent modeling.

[0072] Benefits of step 2: Sparse subspace representation modeling can extract key discriminative features from high-dimensional data, remove redundant features through sparse regularization, significantly reduce data dimensions, reduce computational complexity and storage overhead, and improve the effectiveness of data representation and the accuracy of downstream tasks by evaluating the importance of features and screening feature subsets;

[0073] Benefits of step 3: The dynamic time warping method is used to solve the synchronization problem of multimodal data in the time dimension, and the geometric registration method is used to solve the coordinate differences of different modalities in the spatial dimension, ensuring the consistency of multimodal data in time and space, and providing a feature expression basis for subsequent fusion and joint modeling;

[0074] Benefits of step 4: The sparse feature subsets after time-space alignment are integrated into high-order tensors. The high-order correlation between modalities is modeled through tensor decomposition methods. The core tensor can represent the complex semantic relationship between multimodal data, and the modal factor matrix is ​​used to represent the low-dimensional embedding features unique to each modality. By jointly optimizing the tensor decomposition parameters, the semantic consistency between modalities can be expressed, which improves the expressiveness of the data fusion model.

[0075] Benefits of step 5: Through the dynamic weight allocation mechanism, based on the evaluation of the contribution of modal features to the target task, the weights of the modalities are reasonably allocated, the role of key modalities is highlighted and the interference of redundant modalities is weakened. A preliminary low-dimensional embedding representation is generated through weighted fusion, which improves the robustness and reliability of feature fusion and lays the foundation for the final fusion representation generation;

[0076] Benefits of step 6: Combining the core tensors and factor matrices generated by the dynamic weight allocation results and the tensor decomposition model, the generated low-dimensional compact fusion representation is both sparse and has good expressive power. The fusion representation retains the most important feature information for the target task while reducing redundancy and improving the efficiency and performance of downstream image processing tasks.

[0077] The sparse subspace representation modeling in step 2 includes the following optimization objective function:

[0078]

[0079] Among them, X i represents the characteristic matrix of mode i, P i represents the projection matrix of mode i,

[0080] X i -P i U represents the projection matrix P i and the shared subspace U for the feature matrix X i The approximation of

[0081] ∥P i ∥1 represents the projection matrix P i The sparse regularization term, m represents the number of modes,

[0082] # represents the sparse regularization parameter, represents the reconstruction error term of mode i.

[0083] By optimizing the objective function, sparse subspace representation modeling can extract effective key features from the modal feature matrix, while constraining the projection matrix through sparse regularization terms to reduce the interference of redundant features. Sparsity constraints ensure that the most discriminative features for downstream tasks are retained, while significantly reducing feature dimensions and computational complexity.

[0084] The reconstruction error term included in the optimization objective function can ensure that the projection matrix and shared subspace are efficient approximations of the original feature matrix, retain the main information content in the feature matrix, ensure that the low-dimensional subspace can express the modal characteristics, and provide data input for subsequent analysis and modeling.

[0085] By introducing sparse regularization parameters, we can give weights to important features in the projection matrix, weaken the impact of irrelevant and noisy features, improve the performance of the model during training, and enhance the generalization ability in different data sets and task scenarios.

[0086] In summary, the optimization objective function combines the reconstruction error term and the sparse regularization term to achieve efficient feature selection, low-dimensional approximate expression, and unified feature representation in sparse subspace representation modeling. The specific benefits include:

[0087] Reduce feature dimensions and computational complexity;

[0088] Extract key features and remove redundancy and noise;

[0089] Ensure the accuracy of reconstruction and improve the quality of feature representation;

[0090] Enhance the generalization ability of the model in different data scenarios;

[0091] Provide a consistent feature representation framework for subsequent spatiotemporal alignment and fusion steps.

[0092] In general, this optimization objective function provides a balance between sparsity, reconstruction capability, and computational efficiency, solves the key problem in multimodal high-dimensional feature representation, and is one of the core technical links in multi-source data fusion.

[0093] The projection matrix P in step 2 i The update is iteratively calculated using the following formula:

[0094]

[0095] Among them, soft is the soft threshold function, learning rate, Indicates about P i The gradient of

[0096] t is the number of iterations, # represents the sparse regularization parameter, represents the update result of the sparse projection matrix of the i-th mode after the t+1th iteration,

[0097] Represents the updated result of the sparse projection matrix of the i-th mode after the t-th iteration.

[0098] Through iterative calculation, the projection matrix approaches the optimal solution, and each update is optimized along the direction of the objective function gradient to ensure the convergence of the sparse subspace representation model. As the iteration proceeds, the value of the projection matrix tends to be stable, and a projection matrix that can effectively extract the key features of the modal is obtained.

[0099] The combination of sparse regularization and soft threshold function enables the updated projection matrix to automatically weaken the influence of noise and redundant features, improve the robustness to high-noise data and multimodal data, and enhance the stability and reliability of the sparse subspace representation model.

[0100] In summary, the projection matrix update formula provides an efficient and stable method for determining sparse features by introducing gradient descent, soft threshold function and sparse regularization. The main benefits include:

[0101] Ensure the convergence of the iterative process;

[0102] Implement sparse feature selection to retain the features that contribute most to the task;

[0103] Improve optimization efficiency and reduce computational complexity;

[0104] Enhance the robustness and adaptability of the model in high-dimensional and multimodal data scenarios.

[0105] It plays a core role in sparse subspace representation modeling and provides feature input for subsequent spatiotemporal alignment and fusion of multimodal data.

[0106] The spatiotemporal alignment in step 3 is performed using the following method:

[0107] Synchronization in the time dimension is achieved through the dynamic time warping method, which calculates the optimal alignment path of time series between different modalities;

[0108] The alignment in the spatial dimension adopts a geometric registration method based on affine transformation to map feature subsets of different modalities into a unified spatial coordinate system. The affine transformation matrix is ​​obtained by minimizing the objective function.

[0109] The objective function of geometric registration of affine transformation is:

[0110]

[0111] Among them, T(y i ) is an affine transformation function, which transforms the input feature y i Mapped to the same space,

[0112] y′ i is the target feature coordinate after registration, || || 2 represents the Euclidean distance,

[0113] n represents the total number of eigenvectors in the modal data, y i represents the i-th eigenvector in the original modal data.

[0114] Benefits of Dynamic Time Warping in the time dimension:

[0115] The dynamic time warping method can effectively deal with the problems of different sampling rates and data lengths of time series between different modalities. By calculating the optimal alignment path of the time series, the most matching time point between the modal data is determined to achieve synchronization in the time dimension.

[0116] Dynamic time warping takes into account the global structure and local changes of the data, can capture complex dynamic relationships of time series, make data of different modalities highly consistent in the time dimension, and provide a time alignment basis for subsequent data fusion.

[0117] Benefits of geometric registration methods in spatial dimensions:

[0118] Different modal data come from different acquisition devices or perspectives, resulting in differences in feature spatial distribution. Geometric registration maps different modal feature subsets to a unified coordinate system through affine transformation, eliminating spatial differences between modalities and making feature expression more unified.

[0119] Affine transformation can accurately align the input modal features to the target coordinate system through geometric operations such as rotation, scaling, and translation. At the same time, it minimizes the alignment error and ensures that the registered features can strictly match the target features.

[0120] Benefits of the affine transformation geometric registration objective function:

[0121] By minimizing the global error, the registration process can balance the impact of local feature differences on alignment accuracy, making the result insensitive to single-point errors or abnormal features, while improving the overall robustness of the alignment.

[0122] The affine transformation objective function directly optimizes the alignment error so that the features of different modalities have spatial deviations after being mapped to the same coordinate system, thereby ensuring the accuracy and consistency of subsequent feature fusion.

[0123] In summary, the spatiotemporal alignment method in step 3 achieves unified expression of multimodal data in time and space dimensions through dynamic time warping and geometric registration, which has the following advantages:

[0124] Dynamic time warping solves the problem of asynchrony of multimodal data in the time dimension and improves the accuracy of temporal feature alignment;

[0125] Geometric registration eliminates the spatial differences between modalities through affine transformation to ensure the consistency of features in spatial dimensions;

[0126] The affine transformation objective function minimizes the spatial alignment error and improves the accuracy and robustness of the registration by minimizing the Euclidean distance.

[0127] In general, the spatiotemporal alignment process provides a unified and aligned spatiotemporal feature expression framework for multimodal data, provides support for subsequent feature fusion and joint modeling, and significantly enhances the applicability and generalization ability of the model.

[0128] The tensor decomposition model includes the core tensor and the modal factor matrix A m , satisfying the following representation:

[0129]

[0130] in, Represents the core tensor used to model multimodal joint features.

[0131] A m represents the modal factor matrix, represents a sparse feature subset of multimodal data, and m represents the number of modalities.

[0132] The optimization objective function of the tensor decomposition model is:

[0133]

[0134] in, represents the reconstruction error of tensor decomposition;

[0135] represents the sparse regularization term of the core tensor;

[0136] B is the sparse regularization coefficient, which is used to control the trade-off between sparsity and reconstruction error.

[0137] Core Tensor and the modal factor matrix A m , Represents a sparse feature subset of multimodal data, A i is the factor matrix of mode i, and m represents the number of modes.

[0138] The tensor decomposition model represents the joint features between multimodal data through the core tensor, and can capture the complex high-order correlations between different modal features. Compared with the traditional feature concatenation method, the tensor decomposition method can express the semantic relationship between modalities and make the joint representation of multimodal features accurate.

[0139] Reduce feature redundancy and improve feature compactness. The modal factor matrix extracts the low-dimensional embedding representation of each modality. At the same time, combined with sparse regularization constraints, it can eliminate redundant and useless information, making the feature representation of each modality concise and compact, reducing computational complexity, and providing feature input for subsequent fusion and application.

[0140] The tensor decomposition model can flexibly adapt to multimodal data scenarios and can adapt to any number of modal data. The model can adapt to multimodal data scenarios of different types and quantities through flexible adjustment of core tensors and factor matrices, and has good scalability.

[0141] In summary, the following main benefits can be achieved through the tensor decomposition model and optimization of the objective function:

[0142] Capture high-order correlations of multimodal data and improve the expressiveness of multimodal joint features;

[0143] Eliminate redundant features to make feature representation more compact and efficient;

[0144] Enhance the model's robustness to noise and outliers through sparse regularization;

[0145] Combined with reconstruction error to ensure the accuracy of feature representation;

[0146] Adapt to the diversity and scalability of multimodal data;

[0147] Achieve a unified low-dimensional representation of multimodal features and solve the inconsistency problem between modalities.

[0148] In general, the tensor decomposition model plays a key role in the joint modeling of high-dimensional multimodal data, providing feature input for the subsequent dynamic weight allocation and the generation of the final fusion representation, and significantly improving the efficiency and performance of the image processing process.

[0149] The dynamic weight assignment in step 5 is based on the importance of the modal features and is expressed by the following formula:

[0150]

[0151] Among them, w i is the dynamic weight of mode i, Y i is the feature contribution score of mode i, m represents the number of modes, and exp is an exponential function.

[0152] The low-dimensional compact fused representation generated in step 6 is used for image processing tasks, including image classification, object detection, and image segmentation.

[0153] The dynamic weight allocation in step 5 and the low-dimensional compact fusion representation in step 6 jointly provide support for the efficient fusion of multi-source data, which is mainly reflected in the following aspects:

[0154] The dynamic weight allocation mechanism adaptively adjusts the modal weights through feature contribution, highlights the role of key modalities, weakens the interference of redundant modalities, and ensures the accuracy and robustness of multimodal fusion features.

[0155] The low-dimensional compact fusion representation combines dynamic weights and tensor decomposition results to generate an efficient and sparse joint feature representation, which significantly reduces data dimensions and improves computational efficiency. At the same time, it meets the requirements of downstream tasks for semantic consistency and compactness of feature expression.

[0156] In general, dynamic weight allocation and low-dimensional compact fusion representation are the core links of the multimodal data fusion method. Through the organic combination of the two, the present invention has achieved significant advantages in the accuracy, efficiency and adaptability of feature fusion, providing technical support for multimodal image processing tasks.

[0157] The dynamic weight allocation in step 5 is further combined with the sparse regularization model to dynamically adjust the importance of modal features, and the adaptive allocation of dynamic weights is achieved through the following optimization objective function:

[0158]

[0159] Among them, w i is the dynamic weight of mode i, is the characteristic tensor of mode i, is the core tensor, A m is the factor matrix of mode m;

[0160] is the tensor decomposition reconstruction error of mode i,

[0161] a is the sparse regularization coefficient, ∥w∥1 is the sparse regularization term of the dynamic weight vector, and m is the number of modes.

[0162] The dynamic weight allocation method combined with sparse regularization has the following significant benefits in multimodal data fusion:

[0163] Through adaptive weight adjustment and sparse constraints, the utilization of key modes is strengthened, while the impact of redundant modes is weakened.

[0164] Combined with tensor decomposition and reconstruction error optimization, it ensures that the generated fusion features have higher semantic consistency and discriminability.

[0165] Sparse regularization reduces the interference of abnormal modes and optimizes computational efficiency, making multimodal data fusion more stable and efficient.

[0166] The dynamic weight allocation mechanism can flexibly cope with the diversity of different modal tasks and data scenarios, ensuring the wide applicability of the model.

[0167] In general, the dynamic weight allocation method combined with sparse regularization effectively solves the problems of weight allocation and redundant modal interference in multimodal data fusion, provides an important guarantee for generating high-quality, low-dimensional and compact fusion features, and is one of the key innovative links in multi-source data fusion technology.

[0168] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An image processing method based on multi-source data fusion, characterized in that: include: Step 1: Acquire multimodal image data, represent the data of each modality as a feature matrix, wherein the feature matrix includes the feature dimension and sample quantity of the modality, and standardize and normalize the feature matrix. At the same time, perform noise reduction on the feature matrix to remove abnormal points and noise signals in the data; Step 2: For the feature matrix obtained and preprocessed in step 1, assume that all modal data share a low-dimensional subspace, model the feature matrix through a sparse subspace representation, and extract a sparse feature subset in the feature matrix by optimizing an objective function including reconstruction error and a sparse regularization term. At the same time, the feature subset is screened by evaluating the importance of the sparse features. Step 3: Perform spatiotemporal alignment on the sparse feature subsets extracted in step 2. In view of the spatial and temporal differences in multimodal data, a dynamic time warping method is used to synchronize the time dimension. At the same time, in view of the spatial transformation of different modalities, a registration algorithm is used to map the feature subsets to a unified coordinate system. Step 4: Based on the sparse feature subsets aligned in step 3, a high-order tensor is constructed, wherein each dimension of the high-order tensor corresponds to a sparse representation of the modal feature, and the high-order tensor is modeled by a tensor decomposition method. The tensor decomposition model includes a core tensor and a modal factor matrix. By constructing an optimization objective including a tensor decomposition reconstruction error and a sparse regularization term, the model parameters of the tensor decomposition are jointly optimized to obtain a core tensor and a modal factor matrix that can characterize the joint features of multimodal data; Step 5: Based on the core tensor and modal factor matrix obtained in step 4, combined with the sparsity and correlation of the modal features, dynamic weights are assigned to each modal feature. The dynamic weights are assigned based on the contribution of the modal features to the target task. The modal features are fused in a weighted manner to form a preliminary low-dimensional embedding representation. Step 6: Combine the modal features fused in step 5 with the core tensor and modal factor matrix obtained in step 4 to generate the final low-dimensional compact fusion representation. The fusion representation has a low-dimensional, sparse and compact feature structure. The fusion representation can be directly used as input for downstream image processing tasks.

2. The image processing method based on multi-source data fusion according to claim 1, characterized in that: The sparse subspace representation modeling in step 2 includes the following optimization objective function: Among them, X i represents the characteristic matrix of mode i, P i represents the projection matrix of mode i, X i -P i U represents the projection matrix P i and the shared subspace U for the feature matrix X i Approximation of ∥P i ∥1 represents the projection matrix P i The sparse regularization term, m represents the number of modes, # represents the sparse regularization parameter, represents the reconstruction error term of mode i.

3. The image processing method based on multi-source data fusion according to claim 2, characterized in that: The projection matrix P in step 2 i The update is iteratively calculated using the following formula: Among them, soft is the soft threshold function, learning rate, Indicates about P i The gradient of t is the number of iterations, # represents the sparse regularization parameter, represents the update result of the sparse projection matrix of the i-th mode after the t+1th iteration, Represents the updated result of the sparse projection matrix of the i-th mode after the t-th iteration.

4. The image processing method based on multi-source data fusion according to claim 1, characterized in that: The spatiotemporal alignment in step 3 is performed using the following method: Synchronization in the time dimension is achieved through the dynamic time warping method, which calculates the optimal alignment path of time series between different modalities; The alignment in the spatial dimension adopts a geometric registration method based on affine transformation to map feature subsets of different modalities into a unified spatial coordinate system. The affine transformation matrix is ​​obtained by minimizing the objective function.

5. The image processing method based on multi-source data fusion according to claim 4, characterized in that: The geometric registration objective function of the affine transformation is: Among them, T(y i ) is an affine transformation function, which transforms the input feature y i Mapped to the same space, is the target feature coordinate after registration, ∥∥ 2 represents the Euclidean distance, n represents the total number of eigenvectors in the modal data, y i represents the i-th eigenvector in the original modal data.

6. The image processing method based on multi-source data fusion according to claim 1, characterized in that: The tensor decomposition model includes the core tensor and the modal factor matrix A m , satisfying the following representation: in, Represents the core tensor used to model multimodal joint features. A m represents the modal factor matrix, represents a sparse feature subset of multimodal data, and m represents the number of modalities.

7. The image processing method based on multi-source data fusion according to claim 6, characterized in that: The optimization objective function of the tensor decomposition model is: in, represents the reconstruction error of tensor decomposition; represents the sparse regularization term of the core tensor; B is the sparse regularization coefficient, which is used to control the trade-off between sparsity and reconstruction error. Core Tensor and the modal factor matrix A m , Represents a sparse feature subset of multimodal data, A i is the factor matrix of mode i, and m represents the number of modes.

8. The image processing method based on multi-source data fusion according to claim 1, characterized in that: The dynamic weight allocation in step 5 is based on the importance of the modal features and is expressed by the following formula: Among them, w i is the dynamic weight of mode i, Y i is the feature contribution score of mode i, m represents the number of modes, and exp is an exponential function.

9. The image processing method based on multi-source data fusion according to claim 1, characterized in that: The low-dimensional compact fusion representation generated in step 6 is used for image processing tasks, including image classification, object detection and image segmentation.

10. The image processing method based on multi-source data fusion according to claim 1, characterized in that: The dynamic weight allocation in step 5 is further combined with the sparse regularization model to dynamically adjust the importance of modal features, and the adaptive allocation of dynamic weights is achieved through the following optimization objective function: Among them, w i is the dynamic weight of mode i, is the characteristic tensor of mode i, is the core tensor, A m is the factor matrix of mode m; is the tensor decomposition reconstruction error of mode i, a is the sparse regularization coefficient, ∥w∥1 is the sparse regularization term of the dynamic weight vector, and m is the number of modes.

Citation Information

Patent Citations

  • NMF and low-rank tensor-based incomplete multi-modal media data clustering method

    CN114461961A

  • Multi-modal medical image classification method based on tensor decomposition subspace fusion

    CN117237690A

  • Medical image fusion method based on shared multi-dimensional component tensor dictionary learning

    CN117934305A

  • Target fusion method and system based on radar track projection

    CN118736369A

  • High-precision surveying and mapping method based on multi-source data fusion

    CN119203034A