Multi-modal sensor data adaptive alignment method and system based on deep learning

By constructing a deep learning multi-path feature extraction network and a dynamic adaptive alignment model, and utilizing a vector alignment point convolutional network and a small positional offset alignment technique, the problem of decreased alignment accuracy of multimodal sensor data in dynamic scenarios is solved, and high-precision alignment and fusion of multimodal data is achieved.

CN120822196AActive Publication Date: 2025-10-21SICHUAN COOLBY COMM EQUIP CO LTD

Patent Information

Application Number
CN202511339741.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2025-10-21
Estimated Expiration
2045-09-19

AI Technical Summary

Technical Problem

In existing technologies, multimodal sensor data may be misaligned in terms of time, space, or semantic dimensions, leading to distorted fusion results and system decision-making errors, especially in dynamic scenarios where alignment accuracy decreases.

Method used

By constructing a deep learning-based multi-path feature extraction network and a dynamic adaptive alignment model, and employing a vector alignment point convolutional network and a small positional offset alignment technique, fine-grained alignment is achieved by dynamically adjusting the alignment offset of multimodal features in the spatial position dimension.

Benefits of technology

It significantly improves the accuracy and environmental adaptability of multimodal feature alignment, is suitable for multimodal data processing in dynamic scenarios, reduces alignment errors, and improves the accuracy of data fusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120822196A_ABST
    Figure CN120822196A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-modal sensor data adaptive alignment method and system based on deep learning, and belongs to the technical field of multi-modal data processing, and the method comprises the following steps: S1, obtaining respective spatial point location data from a plurality of heterogeneous sensors, and obtaining a multi-modal sensor data set; s2, constructing a multi-path feature extraction network model to extract spatial convolution feature representations of different modal sensor data; s3, calculating the cross-modal correlation between different modal features by adopting a position vector feature sequence dimension regularity mode, and dynamically adjusting the alignment offset of the multi-modal features in the spatial position dimension based on the cross-modal correlation; s4, performing continuous position domain transformation on the multi-modal data by utilizing a micro position offset alignment technology so as to realize fine-grained alignment of the multi-modal data; the method has the beneficial effects that the alignment offset of the multi-modal feature in the spatial position dimension is dynamically adjusted, so that the alignment operation can adapt to the dynamic change of data in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of multimodal data processing, and in particular relates to a method and system for adaptive alignment of multimodal sensor data based on deep learning. Background Art

[0002] With the rapid development of the Internet of Things (IoT) and intelligent sensing technology, single sensors are no longer sufficient to meet the information collection needs of complex scenarios. For example, autonomous driving requires the coordinated operation of cameras (visual modality), lidar (point cloud modality), millimeter-wave radar (range-velocity modality), and inertial measurement units (IMUs, motion state modality). Industrial equipment monitoring requires the integration of vibration sensors (mechanical vibration modality), temperature sensors (thermal modality), and acoustic sensors (acoustic modality) for fault warning.

[0003] The fusion analysis of multimodal sensor data is the core link to improve the decision-making ability of intelligent systems. Among them, data alignment is the premise and key to multimodal sensor data fusion. If different modal data are misaligned in the time dimension (timing misalignment), spatial dimension (position deviation) or semantic dimension (feature mismatch), it will directly lead to distortion of the fusion results and even cause system decision-making errors. Summary of the Invention

[0004] The present invention provides a multimodal sensor data adaptive alignment method and system based on deep learning, which is used to solve the technical problem in the prior art that multimodal sensor data cannot be aligned. By calculating cross-modal correlation in a dimensional regular manner of position vector feature sequence, the static limitation of fixed offset alignment is broken through. Based on the cross-modal correlation calculated in real time, the alignment offset of multimodal features in the spatial position dimension can be dynamically adjusted. The dynamic adjustment mechanism enables the alignment operation to adapt to the dynamic changes of data in real time, avoiding the problem of decreased alignment accuracy of static alignment in scenarios with dynamic data changes, significantly improving the accuracy and environmental adaptability of multimodal feature alignment, and being particularly suitable for multimodal data processing in dynamic scenarios.

[0005] In order to achieve the above object, the present invention is implemented by the following technical solutions:

[0006] A multimodal sensor data adaptive alignment method based on deep learning is characterized by comprising the following steps:

[0007] Step S1. Acquire spatial point data from multiple heterogeneous sensors to obtain a multimodal sensor dataset;

[0008] Step S2. Construct a multi-path feature extraction network model to extract spatial convolutional feature representations of sensor data of different modalities. Each path of the multi-path feature extraction network model uses a vector-aligned point convolutional network to adapt to the positional variation characteristics of different sensors and output the optimal position offset between the modal data.

[0009] Step S3. Construct a dynamic adaptive alignment model, input the extracted spatial convolution features into the dynamic adaptive alignment model, calculate the cross-modal correlation between different modal features using the position vector feature sequence dimension regularization method, and dynamically adjust the alignment offset of the multimodal features in the spatial position dimension based on the cross-modal correlation;

[0010] Step S4: Using the micro-position offset alignment technology, the multimodal data is transformed in the continuous position domain to achieve fine-grained alignment of the multimodal data.

[0011] Optionally, in step S1, when obtaining the spatial point data of multiple heterogeneous sensors, the points of each heterogeneous sensor are fused and obtained as :

[0012] ;

[0013] in, for The final output spatial point vector after moment fusion; For the Position sensors in Dynamic adjustment coefficient of the moment; is the total number of sensors involved in fusion; It is the sum of the reliability of each sensor’s point vector; It is A data source or sensor at a time The original vector of is the normalization factor, and the point vector representing each sensor is the spatial point vector.

[0014] Optionally, in step S2, the working steps of the multi-path feature extraction network are:

[0015] Receive sensor data of different modalities through the input layer;

[0016] The path feature extraction method is that each path contains: vector alignment point convolution layer, batch normalization layer, ReLU activation function and maximum pooling layer, and feature extraction is performed on the data in each layer;

[0017] The output layer outputs the fused feature representation and the optimal position offset between each modality.

[0018] Optionally, the spatial convolution feature of sensor data of different modalities is expressed as follows: After inputting data of different modalities, the point convolution calculation formula is:

[0019] ;

[0020] in, After the convolution operation, the output feature is at position The value at is the modal number of the convolution kernel; is the convolution kernel In modal , channel index The adjustment parameters at , Separate Modal and Mode, the mode range is [1, ]; is the channel index of the input feature, is the total number of channel indexes;

[0021] is the input feature In position , channel index The value at

[0022] Is the position vector, representing the feature position Corresponding to the physical location information of the sensor; is the position vector The processing function;

[0023] For each neighborhood location , calculate the convolution kernel weight ×Input eigenvalue , then all Add the results and multiply by the position function , so that the final result It also contains the spatial neighborhood information and sensor location information of the input features.

[0024] Optionally, the alignment loss of the optimal position offset is calculated as:

[0025] ;

[0026] in, For the Modal and Modal feature alignment loss, which is used to measure the degree of spatial alignment between two modal features; and Respectively Hedi modality; and is the spatial dimension of the feature, expressed as ; For the Modal features in spatial position The eigenvector at ;

[0027] For the Mode relative to the The spatial offset of the mode, is the spatial offset in the horizontal direction; is the spatial offset in the vertical direction;

[0028] It is the square of the Euclidean norm or the L2 norm, which is used to calculate the distance between two feature vectors. The smaller the distance, the more similar the two features are.

[0029] No. Modal characteristics With the Modal shift characteristics The L2 norm squared is ;

[0030] Specifically, the optimal offset The calculation is:

[0031] ;

[0032] in, 、 and are the offsets of the x-axis, y-axis, and z-axis respectively. The argmin operation means finding the variable value that minimizes the objective function. , let the loss Minimum.

[0033] Optionally, in step S3, the position vector feature sequence dimension is regularized in such a way that the ordered set of position vectors sampled in time series is represented by ,in, is the sequence length;

[0034] Specifically, the dimension regularization method of the position vector feature sequence is:

[0035] ;

[0036] in, For the The mean of the nearest neighbor distances of the convolution kernel modes; For the The point set of the convolution kernel mode; for The number of points included; for The Points, for The Points, for Neighboring points The set of nearest neighbors of ; for point To the nearest neighbor distance;

[0037] It's the right point Neighbor search, find the points closest to it , forming a set ; Traverse the collection Each point in , accumulate;

[0038] By calculating the sum of the average distances of all points to their neighbors, the core is to quantify the density distribution characteristics of the points, which is to calculate the average distance of the local neighbors of a single point to reflect the density around the point. In order to make the point result independent of the number of neighbors;

[0039] If the point result is small, the points are dense; if the point result is large, the points are sparse. It is used for scene recognition. It is to make the characteristics of the points comparable.

[0040] Optionally, in step S3, the alignment offset in the spatial position dimension is applicable to point object alignment. The core is to sum the absolute offsets in each dimension and combine them with spatial correlation correction to avoid a single-dimensional offset masking the overall misalignment:

[0041] ;

[0042] in, Is the basic space alignment offset, is the final calculation result, represents the object and The overall degree of misalignment in space; is the spatial correlation between the benchmark and the target, To align the datum object, Align objects to the target; Dimension Alignment distance; Dimension The geometric property correction coefficient of (i.e., the correction coefficient of the point to be matched); Target In dimension Characteristic parameters on ; To align the base object In dimension Characteristic parameters on ;

[0043] Dimension offset, single dimension Target object With the benchmark object The absolute offset of , the absolute value is added to eliminate the influence of direction and determine the size of the deviation;

[0044] To calculate the correction offset for each dimension, the key dimensions are highlighted, spatial correction is used to adapt the object type, and absolute offset is used to quantify the single-dimensional deviation;

[0045] is the summation term, It is the summation to calculate the correction offset of each dimension;

[0046] Point object alignment involves summing the corrected offsets in each dimension and combining them with spatial correlation corrections to obtain the alignment offset in the spatial position dimension, making the alignment offset more consistent with the actual impact.

[0047] Optionally, in step S4, the slight position offset alignment technology is based on the differences in spatial feature representation of multimodal data, and predicts the position offset caused by these differences through model learning. Then, the data is transformed according to the predicted offset so that data of different modalities can be better matched in the position domain.

[0048] Optionally, the multimodal data is transformed into a continuous position domain. The specific steps are as follows:

[0049] Step a: Feature extraction and offset prediction: extract features from multimodal data and then predict the position offset between features;

[0050] Step b: Interpolation sampling and transformation: According to the predicted offset, the original features are interpolated and sampled to achieve fine-grained spatial alignment;

[0051] Step c: progressive alignment: A multi-stage, progressive approach is used for alignment. The offset result of the previous stage affects the offset prediction of the next stage, making the alignment more refined and stable, avoiding the loss of high-frequency details and the accumulation of cascading errors.

[0052] A multimodal sensor data adaptive alignment system based on deep learning is characterized by including:

[0053] The data acquisition module is used to acquire spatial point data collected by multiple heterogeneous sensors, and integrate the spatial point data into a multimodal sensor dataset after preprocessing. The preprocessing includes data denoising, outlier removal, and format standardization.

[0054] The multi-path feature extraction module is connected to the data acquisition module and has a built-in multi-path feature extraction network model. Each path of the multi-path feature extraction network model is configured with a vector-aligned point convolutional network. The vector-aligned point convolutional network can adapt to the position change characteristics of corresponding heterogeneous sensors. The multi-path feature extraction module uses the multi-path feature extraction network model to extract features from sensor data of different modes in the multimodal sensor dataset, obtain the spatial convolution feature representation of the sensor data of each modality, and output the optimal position offset between the data of each modality.

[0055] The dynamic adaptive alignment module is connected to the multi-path feature extraction module to construct a dynamic adaptive alignment model. The dynamic adaptive alignment model receives the spatial convolution features output by the multi-path feature extraction module, uses the position vector feature sequence dimensionality regularization method to unify the dimensions of the spatial convolution features of different modalities, and then calculates the cross-modal correlation between the features of different modalities based on the dimensionality unified features. The alignment offset of the multimodal features in the spatial position dimension is dynamically adjusted according to the cross-modal correlation;

[0056] The fine-grained alignment execution module is connected to the dynamic adaptive alignment module and adopts the micro-position offset alignment technology. According to the alignment offset adjusted by the dynamic adaptive alignment module, the multimodal data in the multimodal sensor dataset is transformed in the continuous position domain to achieve fine-grained alignment of the multimodal data.

[0057] Beneficial effects of the present invention:

[0058] 1. The present invention constructs a multi-path feature extraction network model and adopts a vector-aligned point convolutional network in each path to perform customized feature extraction based on the position change characteristics of different sensors. Compared with the existing method of extracting all modal features with a unified convolution kernel, it can accurately adapt to the spatial position characteristic differences of different sensors. The unique spatial convolution features of each modal data are captured through the vector-aligned point convolutional network, and the optimal position offset between the modal data is output at the same time, which provides an accurate initial offset reference for subsequent alignment operations and effectively reduces the alignment error caused by insufficient adaptability of feature extraction.

[0059] 2. The position vector feature sequence dimension regularization method of the present invention calculates cross-modal correlation, breaking through the static limitations of fixed offset alignment. In practical applications, it can dynamically adjust the alignment offset of multimodal features in the spatial position dimension based on the cross-modal correlation calculated in real time. This dynamic adjustment mechanism enables the alignment operation to adapt to the dynamic changes of data in real time, avoiding the problem of decreased alignment accuracy of static alignment in scenarios with dynamic data changes, significantly improving the accuracy and environmental adaptability of multimodal feature alignment, and is particularly suitable for multimodal data processing in dynamic scenarios.

[0060] 3. The present invention uses micro-position offset alignment technology to perform step-by-step progressive alignment, adopting a multi-stage, step-by-step approach to alignment. The offset result of the previous stage will affect the offset prediction of the next stage, making the alignment more precise and stable, avoiding the loss of high-frequency details and the accumulation of error cascades, which is more advantageous than the unidirectional deformation convolution strategy. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0062] Figure 1 Schematic diagram of the system structure of the present invention;

[0063] Figure 2 Schematic diagram of the workflow of the present invention. DETAILED DESCRIPTION

[0064] The embodiments of the present application are described in detail below with reference to the accompanying drawings.

[0065] Example 1;

[0066] like Figure 1 As shown, this embodiment provides a multimodal sensor data adaptive alignment system based on deep learning, including:

[0067] The data acquisition module is used to acquire spatial point data collected by multiple heterogeneous sensors, and integrate the spatial point data into a multimodal sensor dataset after preprocessing. The preprocessing includes data denoising, outlier removal, and format standardization.

[0068] A multi-path feature extraction module is connected to the data acquisition module and has a built-in multi-path feature extraction network model. Each path of the multi-path feature extraction network model is configured with a vector-aligned point convolutional network. The vector-aligned point convolutional network can adapt to the position change characteristics of the corresponding heterogeneous sensors. The multi-path feature extraction module uses the multi-path feature extraction network model to extract features from the sensor data of different modes in the multimodal sensor data set, obtain the spatial convolution feature representation of the sensor data of each modality, and output the optimal position offset between the data of each modality.

[0069] The dynamic adaptive alignment module is connected to the multi-path feature extraction module to construct a dynamic adaptive alignment model. The dynamic adaptive alignment model receives the spatial convolution features output by the multi-path feature extraction module, uses the position vector feature sequence dimensionality regularization method to unify the dimensions of the spatial convolution features of different modalities, and then calculates the cross-modal correlation between the features of different modalities based on the dimensionality unified features. The alignment offset of the multimodal features in the spatial position dimension is dynamically adjusted according to the cross-modal correlation;

[0070] The fine-grained alignment execution module is connected to the dynamic adaptive alignment module and adopts the micro-position offset alignment technology. According to the alignment offset adjusted by the dynamic adaptive alignment module, the multimodal data in the multimodal sensor dataset is transformed in the continuous position domain to achieve fine-grained alignment of the multimodal data.

[0071] Example 2;

[0072] Based on Example 1, Figure 2 As shown, this embodiment provides a multimodal sensor data adaptive alignment method based on deep learning, including the following steps:

[0073] Step S1. Acquire spatial point data from multiple heterogeneous sensors to obtain a multimodal sensor dataset;

[0074] By acquiring spatial point data from multiple heterogeneous sensors and forming a multimodal sensor dataset, the limitations of single-modality independent acquisition and weak data correlation in the sensor data acquisition process are broken. The original spatial point information of different heterogeneous sensors (such as visual sensors, lidars, and inertial measurement units) in the same spatial scene can be fully retained, avoiding the loss of key spatial features due to differences in sampling strategies during the data acquisition stage. This lays a high-quality data foundation for subsequent feature extraction and alignment operations, allowing subsequent processing links to be carried out based on complete and accurate original data, reducing the interference of data missing or distortion on the final alignment effect.

[0075] Step S2. Construct a multi-path feature extraction network model to extract spatial convolutional feature representations of sensor data of different modalities. Each path of the multi-path feature extraction network model uses a vector-aligned point convolutional network to adapt to the positional variation characteristics of different sensors and output the optimal position offset between the modal data.

[0076] By constructing a multi-path feature extraction network model, a vector-aligned point-wise convolutional network is employed on each path to perform customized feature extraction based on the positional variation characteristics of different sensors (e.g., installation offsets for some sensors, slight jitter in sensor position in dynamic scenarios, etc.). Compared to existing methods that extract features from all modalities using a unified convolution kernel, this invention accurately adapts to the differences in spatial positional characteristics of different sensors. Using a vector-aligned point-wise convolutional network, it captures the unique spatial convolution features of each modal data, while simultaneously outputting the optimal position offset between the modal data. This provides a precise initial offset reference for subsequent alignment operations, effectively reducing alignment errors caused by insufficient feature extraction adaptability.

[0077] Step S3. Construct a dynamic adaptive alignment model, input the extracted spatial convolution features into the dynamic adaptive alignment model, calculate the cross-modal correlation between different modal features using the position vector feature sequence dimension regularization method, and dynamically adjust the alignment offset of the multimodal features in the spatial position dimension based on the cross-modal correlation;

[0078] By constructing a dynamic adaptive alignment model, cross-modal correlations are calculated in a dimensionally regularized manner using position vector feature sequences, breaking through the static limitations of fixed offset alignment. In practical applications, multimodal sensor data may experience dynamic changes in spatial position relationships due to environmental interference (e.g., electromagnetic interference) or device motion (e.g., sensor movement). The model can dynamically adjust the alignment offset of multimodal features in the spatial position dimension based on cross-modal correlations calculated in real time. This dynamic adjustment mechanism enables alignment operations to adapt to dynamic changes in data in real time, avoiding the problem of decreased alignment accuracy in static alignment methods in scenarios where data is dynamically changing. This significantly improves the accuracy and environmental adaptability of multimodal feature alignment, making it particularly suitable for multimodal data processing in dynamic scenarios (e.g., slight jitter in sensor position in dynamic scenarios).

[0079] Step S4: Using the micro-position offset alignment technology, the multimodal data is transformed in the continuous position domain to achieve fine-grained alignment of the multimodal data.

[0080] The adopted micro-position offset alignment technology optimizes the micro-position deviations that may remain in the previous steps, and realizes the fine-grained alignment of multimodal data in the continuous position domain. In high-precision data application scenarios (such as precision measurement), even small position deviations may lead to poor data fusion and increased errors in subsequent analysis. The micro-position offset alignment technology can accurately correct small position offsets by transforming and adjusting the continuous position domain, so that the multimodal data can be highly matched in spatial position. The implementation of fine-grained alignment not only improves the consistency and reliability of multimodal data, but also provides high-quality data support for subsequent data fusion, feature fusion and decision analysis (such as target recognition and environmental perception based on multimodal data), effectively improving the application value of multimodal sensor data and expanding its application scope in high-precision fields.

[0081] Example 3;

[0082] Based on Example 2, in step S1, multiple heterogeneous sensors include but are not limited to lidar, visual camera, millimeter-wave radar and inertial measurement unit (IMU), and the acquired spatial point data is the three-dimensional coordinate information of each sensor. The spatial position information (such as their respective spatial coordinates) of each sensor is acquired from sensors of different types and principles (i.e., lidar, visual camera, millimeter-wave radar and inertial measurement unit (IMU)) and summarized into a collection of multi-source spatial position data.

[0083] Specifically, when acquiring the spatial point data of multiple heterogeneous sensors, the set of heterogeneous sensors is , at the moment The location of each sensor is ;in, Indicates the The sensors in The spatial point vector output at the moment is Describes the behavior of a sensor (or data source) in a heterogeneous sensor at time Spatial location information; is a vector The coordinate expansion form of 、 and They correspond to the coordinate components of the X-axis, Y-axis, and Z-axis in the spatial rectangular coordinate system, It is transpose, which is to write the original row vector into column vector form.

[0084] After fusion, the points of each heterogeneous sensor are obtained as follows: :

[0085] ;

[0086] in, for The final output spatial point vector after moment fusion; For the The sensors at the Dynamic adjustment coefficient of the moment, In fact, it is The location of each sensor; is the total number of sensors involved in fusion; The reliability of each sensor's point vector is summed up (for high-precision sensors, The reliability of the point vector of the sensor with large and high precision is more significant). It is A data source or sensor at a time The original vector (for example, the vector information of position or features collected by different sensors); is the normalization factor, representing that the point vector of each sensor is a spatial point vector (otherwise the summation will deviate from the actual physical meaning due to the absolute value of the dynamic adjustment coefficient being too large or too small);

[0087] The result is that by dynamically adjusting The fusion result always prioritizes the most reliable sensor in the current scene, obtains high-precision absolute coordinates, and dynamically adjusts As sensor status (accuracy and consistency) changes in real time, the system automatically adapts to the reliability of sensors in different scenarios and fuses the spatial position vectors of each sensor. By averaging multiple original vectors, a comprehensive fused vector is generated. This fusion result integrates multi-source information while reflecting the credibility and importance of each source data. This optimizes the consistency of multi-source data, improves multi-sensor fusion positioning, and integrates multimodal features.

[0088] No. The sensors at the Dynamic adjustment coefficient of time Specifically:

[0089] ;

[0090] in, For sensors Real-time accuracy variance; For sensors Point consistency measurement with other sensors; Experimental calibration coefficient to balance the real-time accuracy variance and point consistency measurement; Represents the sensor status 、 The reliability of each sensor's position vector is adaptively adjusted to the real-time changes. The experimental calibration coefficients are dynamically adjusted through real-time accuracy variance and position consistency measurements, facilitating the adjustment of each component. Adaptive adjustment prioritizes real-time accuracy variance and position consistency measurements, achieving adaptive optimization.

[0091] Example 4;

[0092] Based on Example 2, in step S2, the working steps of the multi-path feature extraction network are:

[0093] Receive sensor data of different modalities (such as visual, infrared, and radar signals) through the input layer;

[0094] The path feature extraction method is that each path contains: vector aligned point convolution layer (VAPC layer), batch normalization layer, ReLU activation function and maximum pooling layer, and feature extraction is performed on the data in each layer;

[0095] The output layer outputs the fused feature representation and the optimal position offset between each modality.

[0096] In step S2, for the multi-path feature extraction network model, assume that the system has Sensor modality, The raw data of the modal is represented as:

[0097] ;

[0098] in, Indicates the Sensor data of each modality; It is a symbol in mathematics, indicating The data type belongs to the space attribute; is the field of real numbers, indicating The elements of are real numbers;

[0099] To describe the dimension of the tensor, it is an abstraction of the spatial structure of sensor data. is the height or first spatial dimension of the data (e.g. the number of vertical pixels in an image, the number of rows after projection of a point cloud), The width or second spatial dimension of the data (e.g., the number of horizontal pixels in an image or the number of columns after projection of a point cloud); is the number of channels (for example, RGB channels for images, XYZ coordinate channels for point clouds, and distance-velocity channels for radar data). 、 and Different (for example, images are 2D + channels, radar is 1D range-Doppler + channels), the formula uses a unified symbol system to describe This modal data is a three-dimensional tensor in the real number field, which facilitates the subsequent multi-path feature extraction network to perform unified or differentiated processing on different modalities.

[0100] The spatial convolution feature representation of sensor data of different modalities is: input data of different modalities After that, the point convolution calculation formula is:

[0101] ;

[0102] in, After the convolution operation, the output feature is at position The value at is the modal number of the convolution kernel; is the convolution kernel In modal , channel index The adjustment parameters at , Separate Modal and Mode, the mode range is [1, ]; is the channel index of the input feature, is the total number of channel indices.

[0103] is the input feature In position , channel index The value at (when convolution, As the center, take the neighborhood of the convolution kernel size on the input feature).

[0104] Is the position vector, representing the feature position Corresponding to the physical location information of the sensor; is the position vector Processing functions (such as nonlinear mapping, functions that encode position features, so that the convolution results can be integrated into the position information).

[0105] Spatial neighborhood sampling is the output location , in the input feature Take one from the top The neighborhood of , the coordinate range is:

[0106] Row extraction ( From 1 to ) → corresponding input row index ;

[0107] Columns extracted as ( From 1 to ) → corresponding input column index ; Channel index is extracted as From 1 to (Traverse all input channels).

[0108] For each neighborhood location , calculate the convolution kernel weight ×Input eigenvalue , then all Then multiply by the position function , so that the final result It also contains the spatial neighborhood information of the input features (i.e., the convolution kernel) and the sensor location information (i.e., ).

[0109] The convolution result integrates the physical location information of the sensor. In a multi-sensor scenario, different sensors (camera, radar) are installed in different locations (i.e. encoding position differences), by The role of convolution is to adaptively adjust the processing of sensor data at different positions, combine the sensor position and convolution operation, and make feature extraction adapt to the scene of multiple sensor positions changing.

[0110] Specifically, for the vector alignment point convolutional network, the first The VAPC layer of the modal path is used to calculate the input data Perform feature extraction, and the VAPC layer output features are:

[0111] ;

[0112] in, Indicates the The modal feature extraction path (corresponding to the processing branch of a certain modal sensor data) is in the first Features of the output of the layer (network level, such as convolutional layer, fully connected layer, etc.); Indicates that it belongs to the real number field ,Right now Features are real numbers;

[0113] is the height or first spatial dimension of the feature (e.g., the number of vertical pixels of the convolutional feature, or the length dimension of the sequence feature); is the width or second spatial dimension of the feature (e.g., the number of horizontal pixels of the convolutional feature, or another dimension of the sequence feature); For the The modal path The number of channels of the layer; describes the dimension or shape of the feature), which is the construction of the feature space structure; the essence of the output feature of the VAPC layer is to convert the first Path No. The layer is characterized by a three-dimensional field of real numbers, the shape of which is given by definition.

[0114] Furthermore, the position-aware convolution operation of the vector-aligned point convolutional network is:

[0115] ;

[0116] in, For the The modal path The number of channels of the layer, For the Layer features in spatial position , channel index The eigenvalue at ; is the scaling factor, which is used to control the output amplitude to avoid the value being too large or too small; is the size of the convolution kernel, that is, the number of sampling points;

[0117] For the The convolution kernel size of the layer, The spatial index inside the convolution kernel (range is 1- ), are the output and input channel indexes respectively ( is the output channel index, is the input channel index);

[0118] For the Layer features at position ,aisle The eigenvalue at It is the spatial sampling position of the input features in the convolution operation;

[0119] is the position vector The processing function, is the feature position The corresponding physical location code (for example: the coordinates of the sensor in 3D space or nonlinear transformation);

[0120] It is a nonlinear function (such as MLP and trigonometric function) used to map physical location to position coordinates so that feature transformation can be integrated into position information. For the Layer, channel index The corresponding bias term.

[0121] The position of the current layer feature is , in the previous layer of features Above, according to the convolution kernel size , sample its neighborhood position , , through the convolution kernel Index the input channel The feature map is mapped to the output channel ,pass Modulate features with physical position information and add channel-specific bias , get the current layer features, and then sum the triples, through Controls the overall amplitude.

[0122] It is a convolution formula that integrates physical location information, allowing feature transformation to simultaneously consider spatial neighborhood relationships and sensor physical locations, and adapt to feature extraction scenarios of multimodal sensor data.

[0123] The alignment loss calculation formula for the optimal position offset is:

[0124] ;

[0125] in, For the Modal and Modal feature alignment loss, which is used to measure the degree of spatial alignment between two modal features; and Respectively Hedi modality; and is the spatial dimension of the feature, expressed as ;

[0126] For the Modal features in spatial position The feature vector at (depending on the channel processing method), if the feature is multi-channel (such as: channel), then is the channel feature splicing at that position;

[0127] For the Mode relative to the The spatial offset of the mode, is the spatial offset in the horizontal direction; The vertical spatial offset.

[0128] The square of the Euclidean norm or the L2 norm is used to calculate the distance between two feature vectors. The smaller the distance, the more similar the two features are. Modal characteristics With the Modal shift characteristics The L2 norm squared is .

[0129] The essence of calculating the optimal position offset is to calculate the pixel-by-pixel L2 distance between the two features after the offset and add them up. Characteristics of modality , applying a spatial offset , get the offset position In scenarios where features are spatially misaligned due to different sensor installation positions, the features of the two modalities are aligned as much as possible through offset; for each spatial position of the feature , calculate the Modal characteristics With the Modal shift characteristics The L2 distance is , L2 distance measures the difference between two eigenvectors. To calculate the function of L2 distance, all spatial positions (Traverse ) is added to get the final alignment loss .

[0130] Optimal offset The calculation is:

[0131] ;

[0132] in, 、 and are the offsets of the x-axis, y-axis, and z-axis respectively. The argmin operation means finding the variable value that minimizes the objective function. , let the loss Min. Alignment loss Measured the Modal characteristics and Modal feature applied offset The smaller the difference, the more aligned the two modal features are under this offset, and the higher the spatial position matching degree.

[0133] Example 5;

[0134] Based on Example 2, in step S3, the dynamic adaptive alignment model is, given the input features , the aligned point convolution obtains an offset at each sampling position of the standard convolution kernel , so that the sampling point position becomes , is the fixed sampling offset of the standard convolution, and the standard convolution kernel is Sampling points at , then the output feature At the base position of the output feature The value at is:

[0135] ;

[0136] in, is the modal number of the convolution kernel; is the size of the convolution kernel or the number of sampling points; The output feature map is at position The position value at For the Convolution of modal sampling points; Represents the spatial mapping of features as a function of input features, where the input features are determined by position; is the input feature position of the sampling point after spatial mapping;

[0137] The input features of the deformed sampling points , convolution with the modal sampling point Sum and get the output .

[0138] The dimensional regularization method of the position vector feature sequence is an ordered set of position vectors sampled in time series, denoted as ,in, is the sequence length (number of sampling points), which is also the dimension to be regularized. The feature dimension of each position vector is fixed, such as 2D is 2-dimensional. The regularization goal is to unify the sequence length. .

[0139] The dimension regularization method for the position vector feature sequence is:

[0140] ;

[0141] in, For the The mean of the nearest neighbor distances of the convolution kernel modes; For the The point set of the convolution kernel mode (i.e., the spatial region divided from the original point cloud); for The number of points included; for The Points, for The Points, for Neighboring points The set of nearest neighbors of ; for point To the nearest neighbor distance;

[0142] It's the right point Neighbor search, find the points closest to it , forming a set ; Traverse the collection Each point in , and accumulate.

[0143] By calculating the sum of the average distances of all points' neighbors, the core is to quantify the density distribution characteristics of the points. It is to calculate the average distance of the local neighbors of a single point to reflect the density around the point (the smaller the distance, the denser the points). In order to make the point result independent of the number of neighbors.

[0144] If the point result is small, the points are dense (such as near neighbors); if the point result is large, the points are sparse (such as far neighbors), which is used for scene recognition. It is to make the characteristics of the points comparable.

[0145] Alignment offset in the spatial position dimension is specifically applicable to point object alignment (e.g., point matching). The core is to sum the absolute offsets in each dimension and combine them with spatial correlation correction to prevent a single dimension offset from masking the overall misalignment:

[0146] ;

[0147] in, Is the basic space alignment offset, is the final calculation result, represents the object and The overall degree of misalignment in space; The spatial correlation between the benchmark and the target (such as the distance between two points, the basic positional relationship between quantitative objects), Alignment datum objects (e.g., datum points, which are references to the theoretical alignment state); Align the target object (e.g., the point to be matched is the subject of the actual misalignment state); Dimension Alignment distance; Dimension The geometric property correction coefficient of (i.e., the correction coefficient of the point to be matched); Target In dimension Characteristic parameters on ; To align the base object In dimension Feature parameters on (such as the X coordinate of a point);

[0148] Dimension offset, single dimension Target object With the benchmark object The absolute offset is added to eliminate the influence of direction and only consider the size of the deviation.

[0149] To calculate the correction offset for each dimension, the key dimensions are highlighted, spatial correction is used to adapt the object type, and absolute offset is used to quantify the single-dimensional deviation;

[0150] is the summation term, It is the summation to calculate the correction offset of each dimension;

[0151] Point object alignment involves summing the corrected offsets in each dimension and combining them with spatial correlation corrections to obtain an alignment offset in the spatial position dimension. This allows the alignment offset to better reflect the actual impact (smaller offsets have less impact on distant objects, while larger offsets have a greater impact on closer objects).

[0152] By introducing (Spatial Correlation) solves the difference between small offsets of distant objects and small offsets of close objects. For example, when two points are 10km apart, a 1m offset on the X axis is acceptable; but when they are 1m apart, a 1m offset on the X axis is a serious misalignment. This method reduces the offset between two points and is suitable for spatial position alignment (such as polar coordinate alignment and longitude and latitude alignment of targets). It converts linear dimension offset into spatial dimension offset calculation to avoid errors in adjacent distances.

[0153] Example 6;

[0154] Based on Example 2, in step S4, the slight position offset alignment technique uses model learning to predict the position offset caused by differences in spatial feature representation of multimodal data. The data is then transformed based on the predicted offset to better align the data from different modalities in the position domain. For example, in remote sensing image fusion, there is a difference in spatial resolution between high-resolution panchromatic images (PAN) and low-resolution multispectral images (MS). This technique can learn the offset of each pixel in the PAN image relative to the corresponding position in the MS image, allowing the PAN image to be adjusted to achieve fine alignment with the MS image.

[0155] The specific steps for transforming the continuous position domain of multimodal data are as follows:

[0156] Step a: Feature extraction and offset prediction: Extract features from multimodal data and then predict the position offset between features. For example, the Modality-Aware Feature Alignment Network (MFAN) aligns the features of the PAN and MS and then predicts the spatial shift of each position.

[0157] Step b: Interpolation sampling and transformation: Based on the predicted offset, the original features are interpolated and sampled to achieve fine-grained spatial alignment. For example, MFAN uses a learnable interpolation function (LIF) to perform bilinear interpolation through pixel-level transformation offsets to finely align the high-frequency spatial structure (PAN) with the low-frequency spectral information (MS).

[0158] Step c: progressive alignment: A multi-stage, progressive approach is used for alignment. The offset result of the previous stage affects the offset prediction of the next stage, making the alignment more precise and stable, avoiding the loss of high-frequency details and the accumulation of cascade errors. This is more advantageous than the unidirectional deformable convolution strategy.

[0159] Specifically, multi-stage progressive adjustment and step-by-step alignment involve each stage performing further processing based on the results of the previous stage. For example, the Bidirectional Stepwise Feature Alignment (BSFA) strategy proposed in BSAFusion uses forward and backward alignment layers to predict the deformation field between features layer by layer. This multi-level step-by-step alignment operation does not complete the alignment all at once, but instead allows features to be adjusted incrementally during the alignment process.

[0160] The migration results of the previous stage will serve as an important basis for the migration prediction of the next stage. For example, when facing the problem of remote sensing resolution mismatch, the migration obtained in the previous stage is used to interpolate and sample the original features, and then the migration prediction of the next stage is carried out based on this, so that the alignment can be continuously refined.

[0161] Because alignment is performed step by step, each adjustment is relatively small, preventing the significant changes to image features that occur with single-step alignment. This allows for better preservation of high-frequency details in the image. For example, the bidirectional stepwise feature alignment strategy uses forward and backward passes to gradually predict the deformation field. Each predicted deformation field is a smaller adjustment based on the current feature state, thus avoiding the loss of high-frequency details caused by a large-scale, one-time deformation.

[0162] Each stage is based on the previous stage's relatively accurate performance. Even if there are errors in the previous stage, they can be corrected through continuous adjustments in subsequent stages, unlike the unidirectional deformation convolution strategy, where errors accumulate as the deformation progresses. For example, in BSAFusion's bidirectional stepwise feature alignment, the stepwise predictions in the forward and reverse directions complement and correct each other, effectively suppressing the cascading accumulation of errors.

[0163] Applied to remote sensing image fusion, it can accurately align and fuse high-resolution PAN images with low-resolution MS images to generate images with both high spatial and spectral resolution, thereby improving the quality and application value of remote sensing images.

[0164] When applied to visual multimodal alignment and processing data, the slight position offset alignment technology can achieve fine-grained alignment of spatial positions, enhance the understanding of the deep correlation between visual positions, and improve the performance of the task.

[0165] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope of the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A multimodal sensor data adaptive alignment method based on deep learning, characterized by: The steps include: Step S1. Acquire spatial point data from multiple heterogeneous sensors to obtain a multimodal sensor dataset; Step S2. Construct a multi-path feature extraction network model to extract spatial convolutional feature representations of sensor data of different modalities. Each path of the multi-path feature extraction network model uses a vector-aligned point convolutional network to adapt to the positional variation characteristics of different sensors and output the optimal position offset between the modal data. Step S3. Construct a dynamic adaptive alignment model, input the extracted spatial convolution features into the dynamic adaptive alignment model, calculate the cross-modal correlation between different modal features using the position vector feature sequence dimension regularization method, and dynamically adjust the alignment offset of the multimodal features in the spatial position dimension based on the cross-modal correlation; Step S4: Using the micro-position offset alignment technology, the multimodal data is transformed in the continuous position domain to achieve fine-grained alignment of the multimodal data.

2. The method for adaptive alignment of multimodal sensor data based on deep learning according to claim 1, characterized in that: In step S1, when obtaining the spatial point data of each of the plurality of heterogeneous sensors, the points of each heterogeneous sensor are fused and obtained as : ; in, for The final output spatial point vector after moment fusion; For the The sensors at the Dynamic adjustment coefficient of the moment; is the total number of sensors involved in fusion; It is the sum of the reliability of each sensor’s point vector; It is A data source or sensor at a time The original vector of is the normalization factor, and the point vector representing each sensor is the spatial point vector.

3. The method for adaptive alignment of multimodal sensor data based on deep learning according to claim 1, characterized in that: In step S2, the working steps of the multi-path feature extraction network are: Receive sensor data of different modalities through the input layer; The path feature extraction method is that each path contains: vector alignment point convolution layer, batch normalization layer, ReLU activation function and maximum pooling layer, and feature extraction is performed on the data in each layer; The output layer outputs the fused feature representation and the optimal position offset between each modality.

4. The method for adaptive alignment of multimodal sensor data based on deep learning according to claim 3, characterized in that: The spatial convolution feature of the sensor data of different modalities is expressed as follows: after inputting the data of different modalities, the point convolution calculation formula is obtained: ; in, After the convolution operation, the output feature is at position The value at is the modal number of the convolution kernel; is the convolution kernel In modal , channel index The adjustment parameters at , Separate Modal and Mode, the mode range is [1, ]; is the channel index of the input feature, is the total number of channel indexes; is the input feature In position , channel index The value at Is the position vector, representing the feature position Corresponding to the physical location information of the sensor; is the position vector The processing function; For each neighborhood location , calculate the convolution kernel weight ×Input eigenvalue , then all Add the results and multiply by the position function , so that the final result It also contains the spatial neighborhood information and sensor location information of the input features.

5. The method for adaptive alignment of multimodal sensor data based on deep learning according to claim 3, characterized in that: The calculation formula for the alignment loss of the optimal position offset is: ; in, For the Modal and Modal feature alignment loss, which is used to measure the degree of spatial alignment between two modal features; and Respectively Hedi modality; and is the spatial dimension of the feature, expressed as ; For the Modal features in spatial position The eigenvector at ; For the Mode relative to the The spatial offset of the mode, is the spatial offset in the horizontal direction; is the spatial offset in the vertical direction; It is the square of the Euclidean norm or the L2 norm, which is used to calculate the distance between two feature vectors. The smaller the distance, the more similar the two features are. No. Modal characteristics With the Modal shift characteristics The L2 norm squared is ; Specifically, the optimal offset The calculation is: ; in, 、 and are the offsets of the x-axis, y-axis, and z-axis respectively. The argmin operation means finding the variable value that minimizes the objective function. , let the loss Minimum.

6. The method for adaptive alignment of multimodal sensor data based on deep learning according to claim 1, characterized in that: In step S3, the position vector feature sequence dimension regularization method is an ordered set of position vectors sampled in time series, which is recorded as ,in, is the sequence length; Specifically, the dimension regularization method of the position vector feature sequence is: ; in, For the The mean of the nearest neighbor distances of the convolution kernel modes; For the The point set of the convolution kernel mode; for The number of points included; for The Points, for The Points, for Neighboring points The set of nearest neighbors of ; for point To the nearest neighbor distance; It's the right point Neighbor search, find the points closest to it , forming a set ; Traverse the collection Each point in , accumulate; By calculating the sum of the average distances of all points to their neighbors, the core is to quantify the density distribution characteristics of the points, which is to calculate the average distance of the local neighbors of a single point to reflect the density around the point. In order to make the point result independent of the number of neighbors; If the point result is small, the points are dense; if the point result is large, the points are sparse. It is used for scene recognition. It is to make the characteristics of the points comparable.

7. The method for adaptive alignment of multimodal sensor data based on deep learning according to claim 1, characterized in that: In step S3, the alignment offset in the spatial position dimension is applicable to point object alignment. The core is to sum the absolute offsets of each dimension and combine it with spatial correlation correction to avoid a single dimension offset masking the overall misalignment: ; in, Is the basic space alignment offset, is the final calculation result, represents the object and The overall degree of misalignment in space; is the spatial correlation between the benchmark and the target, To align the datum object, Align objects to the target; Dimension Alignment distance; Dimension Geometric property correction factor; Target In dimension Characteristic parameters on ; To align the base object In dimension Characteristic parameters on ; Dimension offset, single dimension Target object With the benchmark object The absolute offset of , the absolute value is added to eliminate the influence of direction and determine the size of the deviation; To calculate the correction offset for each dimension, the key dimensions are highlighted, spatial correction is used to adapt the object type, and absolute offset is used to quantify the single-dimensional deviation; is the summation term, It is the summation to calculate the correction offset of each dimension; Point object alignment involves summing the corrected offsets in each dimension and combining them with spatial correlation corrections to obtain the alignment offset in the spatial position dimension, making the alignment offset more consistent with the actual impact.

8. The method for adaptive alignment of multimodal sensor data based on deep learning according to claim 1, characterized in that: In step S4, the slight position offset alignment technology is based on the differences in spatial feature representation of multimodal data, and predicts the position offset caused by these differences through model learning. Then, the data is transformed according to the predicted offset so that the data of different modalities can be better matched in the position domain.

9. The method for adaptive alignment of multimodal sensor data based on deep learning according to claim 8, characterized in that: The multimodal data is transformed into a continuous position domain, and the specific steps are as follows: Step a: Feature extraction and offset prediction: extract features from multimodal data and then predict the position offset between features; Step b: Interpolation sampling and transformation: According to the predicted offset, the original features are interpolated and sampled to achieve fine-grained spatial alignment; Step c: progressive alignment: A multi-stage, progressive approach is used for alignment. The offset result of the previous stage affects the offset prediction of the next stage, making the alignment more refined and stable, avoiding the loss of high-frequency details and the accumulation of cascading errors.

10. A multimodal sensor data adaptive alignment system based on deep learning, which adopts a multimodal sensor data adaptive alignment method based on deep learning according to any one of claims 1 to 9, characterized in that: include: The data acquisition module is used to acquire spatial point data collected by multiple heterogeneous sensors, and integrate the spatial point data into a multimodal sensor dataset after preprocessing. The preprocessing includes data denoising, outlier removal, and format standardization. A multi-path feature extraction module is connected to the data acquisition module and has a built-in multi-path feature extraction network model. Each path of the multi-path feature extraction network model is configured with a vector-aligned point convolutional network. The vector-aligned point convolutional network can adapt to the position change characteristics of the corresponding heterogeneous sensors. The multi-path feature extraction module uses the multi-path feature extraction network model to extract features from the sensor data of different modes in the multimodal sensor data set, obtain the spatial convolution feature representation of the sensor data of each modality, and output the optimal position offset between the data of each modality. The dynamic adaptive alignment module is connected to the multi-path feature extraction module to construct a dynamic adaptive alignment model. The dynamic adaptive alignment model receives the spatial convolution features output by the multi-path feature extraction module, uses the position vector feature sequence dimensionality regularization method to unify the dimensions of the spatial convolution features of different modalities, and then calculates the cross-modal correlation between the features of different modalities based on the dimensionality unified features. The alignment offset of the multimodal features in the spatial position dimension is dynamically adjusted according to the cross-modal correlation; The fine-grained alignment execution module is connected to the dynamic adaptive alignment module and adopts the micro-position offset alignment technology. According to the alignment offset adjusted by the dynamic adaptive alignment module, the multimodal data in the multimodal sensor dataset is transformed in the continuous position domain to achieve fine-grained alignment of the multimodal data.

Citation Information

Patent Citations

  • Fine-grained action recognition method based on cross-modal knowledge alignment

    CN118196888A

  • High-precision multi-mode sensor space alignment method in road monitoring

    CN120164173A

  • Historical and cultural oriented space classification and regeneration method and system

    CN120448874A

Cited By

  • Progressive fine tuning method and system for multi-modal pre-training model

    CN121010981A