A Deep Learning-Based Adaptive Alignment Method and System for Multimodal Sensor Data
By using a deep learning-based adaptive alignment method for multimodal sensor data, the spatial position offset of multimodal features is dynamically adjusted, solving the problem of multimodal sensor data misalignment and achieving high-precision data alignment and fusion.
Patent Information
- Application Number
- CN202511339741.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-19
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-09-19
AI Technical Summary
Multimodal sensor data are misaligned in time, space and semantic dimensions, resulting in distorted fusion results. Existing technologies struggle to achieve accurate data alignment.
A deep learning-based adaptive alignment method for multimodal sensor data is adopted. By constructing a multi-path feature extraction network and a dynamic adaptive alignment model, cross-modal correlation is calculated, the alignment offset of multimodal features in the spatial position dimension is dynamically adjusted, and fine-grained alignment is achieved by combining micro-position offset alignment technology.
It significantly improves the accuracy and environmental adaptability of multimodal feature alignment, is suitable for multimodal data processing in dynamic scenarios, reduces alignment errors, and improves the accuracy and consistency of data fusion.
Smart Images

Figure CN120822196B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of multimodal data processing technology, and specifically relates to a method and system for adaptive alignment of multimodal sensor data based on deep learning. Background Technology
[0002] With the rapid development of the Internet of Things (IoT) and intelligent sensing technologies, single sensors are no longer sufficient to meet the information collection needs of complex scenarios. For example, in autonomous driving scenarios, it is necessary to rely on the coordinated work of cameras (visual modality), lidar (point cloud modality), millimeter-wave radar (range-velocity modality), and inertial measurement units (IMU, motion state modality); in industrial equipment monitoring, it is necessary to combine vibration sensors (mechanical vibration modality), temperature sensors (thermal modality), and acoustic sensors (acoustic modality) to achieve fault early warning.
[0003] The fusion analysis of multimodal sensor data is a core link in improving the decision-making ability of intelligent systems. Among them, data alignment is the premise and key to multimodal sensor data fusion. If different modal data are misaligned in the time dimension (temporal misalignment), spatial dimension (positional deviation), or semantic dimension (feature mismatch), it will directly lead to the distortion of fusion results and even cause system decision-making errors. Summary of the Invention
[0004] This invention provides a deep learning-based adaptive alignment method and system for multimodal sensor data, addressing the technical problem of inability to align multimodal sensor data in existing technologies. By calculating cross-modal correlation through position vector feature sequence dimensional regularization, it overcomes the static limitations of fixed offset alignment. Based on real-time calculated cross-modal correlation, it dynamically adjusts the alignment offset of multimodal features in the spatial position dimension. This dynamic adjustment mechanism enables the alignment operation to adapt to dynamic changes in data in real time, avoiding the accuracy degradation problem of static alignment in scenarios with dynamic data changes. It significantly improves the accuracy and environmental adaptability of multimodal feature alignment, making it particularly suitable for multimodal data processing in dynamic scenarios.
[0005] To achieve the above objectives, the present invention is implemented through the following technical solution:
[0006] A deep learning-based adaptive alignment method for multimodal sensor data is characterized by the following steps:
[0007] Step S1. Obtain spatial location data from multiple heterogeneous sensors to obtain a multimodal sensor dataset;
[0008] Step S2. Construct a multi-path feature extraction network model to extract spatial convolutional feature representations of sensor data from different modalities; wherein, each path of the multi-path feature extraction network model adopts a vector-aligned point convolutional network to adapt to the positional variation characteristics of different sensors and output the optimal positional offset between each modal data.
[0009] Step S3. Construct a dynamic adaptive alignment model. Input the extracted spatial convolutional features into the dynamic adaptive alignment model. Calculate the cross-modal correlation between different modal features using the position vector feature sequence dimension normalization method. Dynamically adjust the alignment offset of multimodal features in the spatial position dimension based on the cross-modal correlation.
[0010] Step S4. Utilize micro-position offset alignment technology to perform continuous position domain transformation on multimodal data to achieve fine-grained alignment of multimodal data.
[0011] Optionally, in step S1, when acquiring the spatial location data of each of the multiple heterogeneous sensors, the locations of each heterogeneous sensor are fused and obtained as follows: :
[0012] ;
[0013] in, for The spatial point vector of the final output after time-mapping; For the first A position sensor in Dynamic adjustment coefficient at any given time; The total number of sensors participating in the fusion; It involves summing the reliability of the point vector for each sensor; It is the first A data source or sensor at any time The original vector; The normalization factor represents the spatial point vector of each sensor.
[0014] Optionally, in step S2, the working steps of the multi-path feature extraction network are as follows:
[0015] The input layer receives sensor data in different modes.
[0016] The path feature extraction method is that each path includes: vector-aligned point convolutional layer, batch normalization layer, ReLU activation function and max pooling layer, and feature extraction is performed on the data in each layer;
[0017] The output layer outputs the fused feature representation and the optimal position offset between each mode.
[0018] Optionally, the spatial convolution features of sensor data from different modalities can be represented as follows: After inputting data from different modalities, the point convolution calculation formula is:
[0019] ;
[0020] in, After the convolution operation, the output features are located at... The value at that location; The number of modes of the convolution kernel; convolution kernel In modality Channel Index Adjustment parameters at the location; , The respective Modality and the first Modal range, the modal range is [1, ]; It is the channel index of the input feature. This represents the total number of channel indexes.
[0021] Input features In position Channel Index The value at that location;
[0022] This is a position vector, representing the feature location. Corresponding to the physical location information of the sensor; For position vectors The processing function;
[0023] For each neighborhood location Calculate the convolution kernel weights × Input feature value Then put all Add the results together, then multiply by the position function. Let the final result It also includes spatial neighborhood information of the input features and sensor location information.
[0024] Optionally, the alignment loss for the optimal position offset is calculated using the following formula:
[0025] ;
[0026] in, For the first Modality and the first Alignment loss of modal features is used to measure the degree of spatial alignment between two modal features; and The first and the Modality; and The spatial dimension of the feature is represented as ; For the first Modal features in spatial location The eigenvector at that location;
[0027] For the first Mode relative to the first Spatial offset of the mode, This represents the spatial offset in the horizontal direction. This represents the spatial offset in the vertical direction.
[0028] The square of the Euclidean norm or the square of the L2 norm is used to calculate the distance between two eigenvectors. The smaller the distance, the more similar the two features are.
[0029] No. Modal features With the Features after mode shift The square of the L2 norm is ;
[0030] Specifically, the optimal offset The calculation is as follows:
[0031] ;
[0032] in, , and These are the offsets along the x, y, and z axes, respectively. The argmin operation finds the variable values that minimize the objective function, and finds a set of... Let the loss Minimum.
[0033] Optionally, in step S3, the position vector feature sequence dimension normalization method is an ordered set composed of temporally sampled position vectors, denoted as... ,in, The sequence length;
[0034] Specifically, the dimension normalization method for the position vector feature sequence is as follows:
[0035] ;
[0036] in, For the first The mean of the nearest neighbor distances of each convolution kernel mode; For the first The set of points representing the modalities of each convolutional kernel; for The number of points included; for The first in One point, for The first in One point, for Nearest neighbor points The nearest neighbor set; For point To the nearest point The distance;
[0037] Point Nearest neighbor search: find the points that are closest to it. , forming a set ; Traverse the set Each point in Accumulate;
[0038] By calculating the sum of the average nearest distances of all points, the core concept is to quantify the density distribution characteristics of the points. This involves calculating the average local nearest distance of a single point to reflect the density of its surroundings. To make the result independent of the number of nearest neighbors;
[0039] For small point values, the points are densely packed; for large point values, the points are sparsely packed. This is used for scene recognition. This makes the characteristics of the points comparable.
[0040] Optionally, in step S3, the alignment offset in the spatial position dimension is applicable to point object alignment. The core is to sum the absolute offsets in each dimension and combine them with spatial correlation correction to avoid a single-dimensional offset from masking the overall misalignment.
[0041] ;
[0042] in, It is the base space alignment offset, the final calculation result, representing the object. and The numerical value of the overall degree of misalignment in space; The spatial correlation between the baseline and the target To align with the reference object, Align the target object; For dimension Alignment distance; For dimension The geometric attribute correction coefficient (i.e., the correction coefficient of the point to be matched); For the target object In dimensions Feature parameters on; To align the reference object In dimensions Feature parameters on;
[0043] For dimension offset, single dimension Above, target object With reference object The absolute offset is used to eliminate the influence of direction and determine the magnitude of the deviation.
[0044] To calculate the corrected offset for each dimension, highlighting key dimensions, spatial correction is used to adapt to the object type, and absolute offset is used to quantify the single-dimensional deviation.
[0045] For the summation term, It sums up the corrected offset for each dimension;
[0046] Point object alignment involves summing the correction offsets in each dimension and combining them with spatial correlation correction to obtain the alignment offset in the spatial position dimension, making the alignment offset more closely reflect the actual impact.
[0047] Optionally, in step S4, the small position offset alignment technique is based on the differences in spatial feature representation of multimodal data. The model learns to predict the position offset caused by these differences, and then transforms the data according to the predicted offset so that the data of different modalities can be better matched in the position domain.
[0048] Optionally, the multimodal data can be transformed in the continuous location domain. The specific steps are as follows:
[0049] Step a: Feature extraction and offset prediction; extract features from multimodal data and then predict the positional offsets between features;
[0050] Step b: Interpolation sampling and transformation; based on the predicted offset, the original features are interpolated to achieve fine-grained spatial alignment;
[0051] Step c: Progressive alignment; Alignment is performed in a multi-stage, progressive manner. The offset result of the previous stage will affect the offset prediction of the next stage, making the alignment more refined and stable, and avoiding the loss of high-frequency details and the accumulation of cascading errors.
[0052] A deep learning-based multimodal sensor data adaptive alignment system, characterized in that it includes:
[0053] The data acquisition module is used to acquire spatial point data collected by multiple heterogeneous sensors, and integrate the spatial point data after preprocessing to form a multimodal sensor dataset; the preprocessing includes data denoising, outlier removal and format standardization.
[0054] The multi-path feature extraction module, connected to the data acquisition module, has a built-in multi-path feature extraction network model. Each path of the multi-path feature extraction network model is configured with a vector-aligned point convolutional network. The vector-aligned point convolutional network can adapt to the positional change characteristics of the corresponding heterogeneous sensors. The multi-path feature extraction module extracts features from the sensor data of different modalities in the multi-modal sensor dataset through the multi-path feature extraction network model, obtains the spatial convolutional feature representation of each modal sensor data, and outputs the optimal positional offset between each modal data.
[0055] The dynamic adaptive alignment module is connected to the multi-path feature extraction module and constructs a dynamic adaptive alignment model. The dynamic adaptive alignment model receives the spatial convolution features output by the multi-path feature extraction module, uses the position vector feature sequence dimension normalization method to unify the dimensions of the spatial convolution features of different modalities, calculates the cross-modal correlation between the features of different modalities based on the dimension-unified features, and dynamically adjusts the alignment offset of the multimodal features in the spatial position dimension according to the cross-modal correlation.
[0056] The fine-grained alignment execution module, connected to the dynamic adaptive alignment module, employs a micro-position offset alignment technique. Based on the alignment offset adjusted by the dynamic adaptive alignment module, it performs continuous position domain transformation on the multimodal data in the multimodal sensor dataset to achieve fine-grained alignment of the multimodal data.
[0057] The beneficial effects of this invention are:
[0058] 1. This invention constructs a multi-path feature extraction network model, employing a vector-aligned point convolutional network for each path. It performs customized feature extraction based on the positional variation characteristics of different sensors. Compared to the existing method of extracting all modal features using a unified convolutional kernel, this invention can accurately adapt to the spatial positional characteristics differences of different sensors. By capturing the unique spatial convolutional features of each modal data through the vector-aligned point convolutional network, it outputs the optimal positional offset between each modal data, providing a precise initial offset reference for subsequent alignment operations and effectively reducing alignment errors caused by insufficient feature extraction adaptability.
[0059] 2. The position vector feature sequence dimension regularization method of the present invention calculates cross-modal correlation, which breaks through the static limitation of fixed offset alignment. In practical applications, it can dynamically adjust the alignment offset of multimodal features in the spatial position dimension based on real-time calculated cross-modal correlation. This dynamic adjustment mechanism enables the alignment operation to adapt to the dynamic changes of data in real time, avoiding the problem of decreased alignment accuracy of static alignment in scenarios with dynamic data changes. It significantly improves the accuracy and environmental adaptability of multimodal feature alignment, and is especially suitable for multimodal data processing in dynamic scenarios.
[0060] 3. This invention utilizes a micro-position offset alignment technique to perform progressive alignment. It adopts a multi-stage, progressive approach to alignment, where the offset result of the previous stage affects the offset prediction of the next stage, making the alignment more refined and stable. This avoids the loss of high-frequency details and the cascading accumulation of errors, and has advantages over unidirectional deformable convolution strategies. Attached Figure Description
[0061] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0062] Figure 1 This is a schematic diagram of the system structure of the present invention;
[0063] Figure 2 This is a schematic diagram of the workflow of the present invention. Detailed Implementation
[0064] The embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0065] Example 1;
[0066] like Figure 1 As shown, this embodiment provides a deep learning-based multimodal sensor data adaptive alignment system, including:
[0067] The data acquisition module is used to acquire spatial point data collected by multiple heterogeneous sensors, and integrate the spatial point data after preprocessing to form a multimodal sensor dataset; the preprocessing includes data denoising, outlier removal and format standardization.
[0068] The multi-path feature extraction module, connected to the data acquisition module, has a built-in multi-path feature extraction network model. Each path of the multi-path feature extraction network model is configured with a vector-aligned point convolutional network. The vector-aligned point convolutional network can adapt to the positional change characteristics of the corresponding heterogeneous sensors. The multi-path feature extraction module extracts features from the sensor data of different modalities in the multimodal sensor dataset through the multi-path feature extraction network model to obtain the spatial convolutional feature representation of each modal sensor data, and outputs the optimal positional offset between each modal data.
[0069] The dynamic adaptive alignment module is connected to the multi-path feature extraction module and constructs a dynamic adaptive alignment model. The dynamic adaptive alignment model receives the spatial convolution features output by the multi-path feature extraction module, uses the position vector feature sequence dimension normalization method to unify the dimensions of the spatial convolution features of different modalities, calculates the cross-modal correlation between the features of different modalities based on the dimension-unified features, and dynamically adjusts the alignment offset of the multimodal features in the spatial position dimension according to the cross-modal correlation.
[0070] The fine-grained alignment execution module, connected to the dynamic adaptive alignment module, employs a micro-position offset alignment technique. Based on the alignment offset adjusted by the dynamic adaptive alignment module, it performs continuous position domain transformation on the multimodal data in the multimodal sensor dataset to achieve fine-grained alignment of the multimodal data.
[0071] Example 2;
[0072] Based on Example 1, such as Figure 2 As shown, this embodiment provides a deep learning-based adaptive alignment method for multimodal sensor data, including the following steps:
[0073] Step S1. Obtain spatial location data from multiple heterogeneous sensors to obtain a multimodal sensor dataset;
[0074] By acquiring spatial location data from multiple heterogeneous sensors and forming a multimodal sensor dataset, the limitations of independent acquisition of a single mode and weak data correlation in the sensor data acquisition process are overcome. It can completely preserve the original spatial location information of different heterogeneous sensors (such as vision sensors, LiDAR, and inertial measurement units) in the same spatial scene, avoiding the loss of key spatial features due to differences in sampling strategies during the data acquisition stage. This lays a high-quality data foundation for subsequent feature extraction and alignment operations, enabling subsequent processing steps to be carried out based on complete and accurate original data, reducing the interference of missing or distorted data on the final alignment effect.
[0075] Step S2. Construct a multi-path feature extraction network model to extract spatial convolutional feature representations of sensor data from different modalities; wherein, each path of the multi-path feature extraction network model adopts a vector-aligned point convolutional network to adapt to the positional variation characteristics of different sensors and output the optimal positional offset between each modal data.
[0076] By constructing a multi-path feature extraction network model, a vector-aligned point-based convolutional network is used for each path to perform customized feature extraction based on the positional variation characteristics of different sensors (such as installation offset of some sensors, slight positional jitter of sensors in dynamic scenes, etc.). Compared with the existing method of extracting all modal features with a unified convolutional kernel, this invention can accurately adapt to the differences in spatial positional characteristics of different sensors. By capturing the unique spatial convolutional features of each modal data through the vector-aligned point-based convolutional network, it outputs the optimal positional offset between each modal data, providing a precise initial offset reference for subsequent alignment operations, and effectively reducing alignment errors caused by insufficient feature extraction adaptability.
[0077] Step S3. Construct a dynamic adaptive alignment model. Input the extracted spatial convolutional features into the dynamic adaptive alignment model. Calculate the cross-modal correlation between different modal features using the position vector feature sequence dimension normalization method. Dynamically adjust the alignment offset of multimodal features in the spatial position dimension based on the cross-modal correlation.
[0078] By constructing a dynamic adaptive alignment model, cross-modal correlations are calculated using a dimension-normalized method for position vector feature sequences, overcoming the static limitations of fixed-offset alignment. In practical applications, the spatial positional relationships of multimodal sensor data may dynamically change due to environmental interference (e.g., electromagnetic interference) and device movement (e.g., sensor movement). The model can dynamically adjust the alignment offset of multimodal features in the spatial positional dimension based on real-time calculated cross-modal correlations. This dynamic adjustment mechanism enables the alignment operation to adapt to the dynamic changes in data in real time, avoiding the accuracy degradation problem of static alignment methods in scenarios with dynamic data changes. It significantly improves the accuracy and environmental adaptability of multimodal feature alignment, and is particularly suitable for multimodal data processing in dynamic scenarios (e.g., where sensor positions fluctuate slightly in dynamic scenarios).
[0079] Step S4. Utilize micro-position offset alignment technology to perform continuous position domain transformation on multimodal data to achieve fine-grained alignment of multimodal data.
[0080] The employed micro-position offset alignment technique optimizes for any residual micro-positional deviations from previous steps, achieving fine-grained alignment of multimodal data across a continuous positional domain. In high-precision data applications (such as precision measurement), even minute positional deviations can lead to poor data fusion and increased errors in subsequent analysis. This micro-position offset alignment technique, through transformation and adjustment of the continuous positional domain, precisely corrects these micro-positional deviations, enabling a high degree of spatial matching between multimodal data. This fine-grained alignment not only improves the consistency and reliability of multimodal data but also provides high-quality data support for subsequent data fusion, feature fusion, and decision analysis (such as target recognition and environmental perception based on multimodal data), effectively enhancing the application value of multimodal sensor data and expanding its application scope in high-precision fields.
[0081] Example 3;
[0082] Based on Embodiment 2, in step S1, multiple heterogeneous sensors include, but are not limited to, lidar, vision camera, millimeter-wave radar and inertial measurement unit (IMU). The acquired spatial position data is the three-dimensional coordinate information of each sensor. The spatial position information (e.g., their respective spatial coordinates) of each sensor is acquired from sensors of different types and with different principles (i.e. lidar, vision camera, millimeter-wave radar and inertial measurement unit (IMU)) and summarized into a set of multi-source spatial position data.
[0083] Specifically, when acquiring spatial location data from multiple heterogeneous sensors, the set of each heterogeneous sensor is as follows: At any moment The locations of each sensor are as follows ;in, Indicates the first One sensor in The spatial point vector output at any given time, i.e. Describes the state of a sensor (or data source) in a heterogeneous sensor system at time [time]. Spatial location information; It is a vector The coordinate expansion form, , and These correspond to the coordinate components of the X, Y, and Z axes in a Cartesian coordinate system, respectively. It is a transpose, which is to write the original row vector in the form of a column vector.
[0084] After fusion, the locations of each heterogeneous sensor are obtained as follows: :
[0085] ;
[0086] in, for The spatial point vector of the final output after time-mapping; For the first Sensors at each location The dynamic adjustment coefficient at any given time. Actually, it's the first one. The location of each sensor; The total number of sensors participating in the fusion; This involves summing the reliability of the point vector for each sensor (for high-precision sensors, ...). The reliability of the point vector of a large, high-precision sensor has a more significant impact. It is the first A data source or sensor at any time The original vector (e.g., vector information of location or features collected by different sensors); As a normalization factor, it represents the spatial point vector of each sensor (otherwise, the summation will deviate from the actual physical meaning due to the absolute value of the dynamic adjustment coefficient being too large or too small).
[0087] The result is through dynamic adjustment The fusion results always prioritize the most reliable sensors in the current scenario to obtain high-precision absolute coordinates, and are dynamically adjusted. As sensor status (accuracy and consistency) changes in real time, the system automatically adapts to the reliability of sensors in different scenarios and fuses the spatial point vectors of various sensors. By averaging multiple original vectors, a comprehensive fusion vector is obtained, which integrates multi-source information and reflects the differences in the credibility and importance of each source data. This optimizes the consistency of multi-source data, improves multi-sensor fusion positioning, and the fusion of multi-modal features.
[0088] No. Sensors at each location Dynamic adjustment coefficient at time Specifically:
[0089] ;
[0090] in, For sensors The real-time accuracy variance; For sensors Measurement of point consistency with other sensors; Experimental calibration coefficients to balance real-time accuracy variance and location consistency measures; Represents the state of the sensor , The system adapts to the real-time changes in the location vector of each sensor, adjusting its reliability accordingly. By dynamically adjusting the experimental calibration coefficients using real-time accuracy variance and location consistency metrics, it facilitates the adjustment of each component. Adaptive adjustment prioritizes real-time accuracy variance and location consistency metrics, achieving adaptive optimization.
[0091] Example 4;
[0092] Based on Example 2, the working steps of the multi-path feature extraction network in step S2 are as follows:
[0093] The input layer receives sensor data of different modalities (such as visual, infrared, and radar signals).
[0094] The path feature extraction method is that each path contains: a vector-aligned point convolutional layer (VAPC layer), a batch normalization layer, a ReLU activation function, and a max pooling layer, and features are extracted from the data in each layer;
[0095] The output layer outputs the fused feature representation and the optimal position offset between each mode.
[0096] In step S2, for the multi-path feature extraction network model, let the system have... Type 1 sensor mode, the first The raw data for the modality is represented as follows:
[0097] ;
[0098] in, Indicates the first Sensor data for each modality; In mathematics, "belongs" is a symbol that represents... The data type belongs to the space allocation; For the real number field, represent The elements are real numbers;
[0099] The dimension of a tensor describes an abstraction of the spatial structure of sensor data. The height of the data or the first spatial dimension (e.g., the number of vertical pixels in an image, the number of rows after point cloud projection). For the width of the data or the second spatial dimension (e.g., the number of horizontal pixels in an image, the number of columns after point cloud projection); This refers to the number of channels (e.g., the RGB three channels of an image, the XYZ coordinate channels of a point cloud, and the range-velocity channels of radar data). Different sensor modalities... , and The differences (e.g., image processing is 2D + channels, radar is 1D range - Doppler + channels) are described by a unified notation system in the formula. Each modality data is a three-dimensional tensor over the real number domain, which facilitates the subsequent multi-path feature extraction network to perform unified or differentiated processing on different modalities.
[0100] The spatial convolutional features of sensor data from different modalities are represented as follows: input data from different modalities Then, the formula for calculating point convolution is:
[0101] ;
[0102] in, After the convolution operation, the output features are located at... The value at that location; The number of modes of the convolution kernel; convolution kernel In modality Channel Index Adjustment parameters at the location; , The respective Modality and the first Modal range, the modal range is [1, ]; It is the channel index of the input feature. This represents the total number of channel indexes.
[0103] Input features In position Channel Index The value at (during convolution, in) Centered on the input features, a neighborhood of the kernel size is selected.
[0104] This is a position vector, representing the feature location. Corresponding to the physical location information of the sensor; For position vectors Processing functions (such as nonlinear mapping, functions that encode location features, and functions that incorporate location information into the convolution results).
[0105] Spatial neighborhood sampling is the output location In input features Take one The neighborhood of , with coordinates ranging from:
[0106] Row extraction as ( From 1 to → Corresponding input row index ;
[0107] Column extraction as ( From 1 to → Corresponding input column index Channel index extraction is From 1 to (Iterate through all input channels).
[0108] For each neighborhood location Calculate the convolution kernel weights × Input feature value Then put all Add the results together. Then multiply by the position function. Let the final result It simultaneously includes spatial neighborhood information of the input features (i.e., the convolution kernel) and sensor location information (i.e., ... ).
[0109] The convolution result incorporates the physical location information of the sensors. In multi-sensor scenarios, different sensors (cameras, radars) are installed in different locations (i.e., (Encoding position differences), through The role of convolution is to adaptively adjust the processing of sensor data at different locations. By combining sensor location with convolution operations, feature extraction can be adapted to scenarios with changing sensor locations.
[0110] Specifically, for vector-aligned point-level convolutional networks, the first... The VAPC layer of the modal path handles the input data. Feature extraction is performed, and the VAPC layer outputs the following features:
[0111] ;
[0112] in, Indicates the first Modal feature extraction path (the processing branch corresponding to data from a specific modal sensor) in the first... Features output by layers (network layers, such as convolutional layers, fully connected layers, etc.); Indicates belonging to the real number field ,Right now The characteristic is that it is a real number;
[0113] The height or first spatial dimension of the feature (e.g., the number of vertical pixels in the feature after convolution, the length dimension of the sequence feature). This refers to the width of the feature or a second spatial dimension (e.g., the number of horizontal pixels in the feature after convolution, or another dimension of the sequence feature). For the first The first modal path The number of channels in the layer; (Describing the dimension or shape of features) is the construction of the feature space structure; the essence of the VAPC layer outputting features is to transform the first... Path number The characteristic of a layer is a three-dimensional real number field, the shape of which is determined by... definition.
[0114] Furthermore, the position-aware convolution operation of the vector-aligned point convolutional network is as follows:
[0115] ;
[0116] in, For the first The first modal path The number of channels in the layer For the first Layer features in spatial location Channel Index Eigenvalues at; This is the scaling factor, used to control the output amplitude and avoid the value being too large or too small; This refers to the size of the convolution kernel, i.e., the number of sampling points;
[0117] For the first kernel size of the layer Spatial index inside the convolution kernel (range 1- ), These are the output and input channel indices, respectively. It is the output channel index. (This is the input channel index).
[0118] For the first Layer features at location ,aisle eigenvalues at that location It refers to the spatial sampling location of the input features during the convolution operation;
[0119] For position vectors The processing function, It is a characteristic location The corresponding physical location code (e.g., the sensor's coordinates in 3D space) (or nonlinear transformation);
[0120] These are nonlinear functions (such as MLPs and trigonometric functions) used to map physical locations to location coordinates, incorporating location information into feature transformations. For the first Layer and Channel Index The corresponding bias term.
[0121] The position of the current layer feature is In the previous layer of features Above, according to the kernel size Sample its neighborhood location , Through convolution kernel Index the input channels Feature mapping to output channel ,pass Features are modulated using physical location information, plus a channel-specific offset. To obtain the features of the current layer, the terms are summed three times, and then... Control the overall amplitude.
[0122] It is a convolution formula that integrates physical location information, allowing feature transformation to consider both spatial neighborhood relationships and sensor physical location, thus adapting to feature extraction scenarios of multimodal sensor data.
[0123] The formula for calculating the alignment loss of the optimal position offset is:
[0124] ;
[0125] in, For the first Modality and the first Alignment loss of modal features is used to measure the degree of spatial alignment between two modal features; and The first and the Modality; and The spatial dimension of the feature is represented as ;
[0126] For the first Modal features in spatial location The feature vector at a given location (depending on the channel processing method), if the feature is multi-channel (e.g., the first...). (channel), then It is a splicing of channel features at this location;
[0127] For the first Mode relative to the first Spatial offset of the mode, This represents the spatial offset in the horizontal direction. This represents the spatial offset in the vertical direction.
[0128] The square of the Euclidean norm or the square of the L2 norm is used to calculate the distance between two eigenvectors. The smaller the distance, the more similar the two features are. Modal features With the Features after mode shift The square of the L2 norm is .
[0129] The calculation of the optimal position offset essentially involves calculating the pixel-by-pixel L2 distance between two features at the offset positions and summing them up. Specifically, this is done for the first... Modal characteristics Apply spatial offset To obtain the offset position In scenarios where spatial misalignment of features occurs due to different sensor installation positions, offsetting is used to align the features of the two modalities as much as possible; for each spatial position of the feature... Calculate the first Modal features With the Features after mode shift The L2 distance is The L2 distance measures the difference between two feature vectors. The function for calculating the L2 distance is to consider all spatial locations... (traversal) The L2 distances of the two sides are summed to obtain the final alignment loss. .
[0130] Optimal offset The calculation is as follows:
[0131] ;
[0132] in, , and These are the offsets along the x, y, and z axes, respectively. The argmin operation finds the variable values that minimize the objective function, and finds a set of... Let the loss Minimum alignment loss The first Modal characteristics and the first Modal feature offset applied The smaller the difference, the more aligned the two modal features are under that offset, and the higher the spatial matching degree.
[0133] Example 5;
[0134] Based on Example 2, in step S3, the dynamic adaptive alignment model is, given input features... Aligned point convolution produces an offset at each sampling position of the standard convolution kernel. This makes the sampling point position change , Given a fixed sampling offset for standard convolution, let the standard convolution kernel have... There are 1 sampling point, located at 1 Then output features At the reference position of the output feature The value at that location is:
[0135] ;
[0136] in, The number of modes of the convolution kernel; This refers to the size of the convolution kernel or the number of sampling points; To output feature map at location The position value at that location; For the first Convolution of modal sampling points; The spatial mapping of the input features is represented in functional form, and the input features are determined by their location; The input feature positions of the sampling points after spatial mapping;
[0137] Input features of the deformed sampling points Convolution with modal sampling points Summing gives the output. .
[0138] The dimension normalization method of the position vector feature sequence is an ordered set composed of temporally sampled position vectors, denoted as... ,in, The sequence length (number of sampling points) is also the dimension to be normalized. The feature dimension of each position vector is fixed, such as 2D being 2-dimensional. The normalization goal is to unify the sequence length. .
[0139] The dimension normalization method for the position vector feature sequence is as follows:
[0140] ;
[0141] in, For the first The mean of the nearest neighbor distances of each convolution kernel mode; For the first The set of points of each convolutional kernel mode (i.e., the spatial region divided from the original point cloud). for The number of points included; for The first in One point, for The first in One point, for Nearest neighbor points The nearest neighbor set; For point To the nearest point The distance;
[0142] Point Nearest neighbor search: find the points that are closest to it. , forming a set ; Traverse the set Each point in , and then accumulate them.
[0143] The core of this method is to quantify the density distribution characteristics of points by calculating the sum of the average distances of the nearest neighbors of all points. It calculates the average distance of the local nearest neighbors of a single point to reflect the density of the points around that point (the smaller the distance, the denser the points). To make the result of a point independent of the number of its neighbors.
[0144] For small point values, points are densely packed (e.g., nearest neighbors); for large point values, points are sparsely packed (e.g., distant neighbors). This is used for scene recognition. This makes the characteristics of the points comparable.
[0145] The alignment offset in the spatial dimension, specifically applicable to point object alignment (e.g., point matching), is essentially the sum of the absolute offsets in each dimension, combined with spatial correlation correction, to prevent a single-dimensional offset from masking the overall misalignment.
[0146] ;
[0147] in, It is the base space alignment offset, the final calculation result, representing the object. and The numerical value of the overall degree of misalignment in space; The spatial correlation between the baseline and the target (e.g., the distance between two points, quantifying the basic positional relationship between objects). To align with a reference object (e.g., a reference point, which is a reference for the theoretical alignment state); Align the target object (e.g., the point to be matched, which is the main body of the actual misalignment state); For dimension Alignment distance; For dimension The geometric attribute correction coefficient (i.e., the correction coefficient of the point to be matched); For the target object In dimensions Feature parameters on; To align the reference object In dimensions Feature parameters on the surface (e.g., the X coordinate of a point);
[0148] For dimension offset, single dimension Above, target object With reference object The absolute offset is added to eliminate the influence of direction, so only the magnitude of the deviation is considered.
[0149] To calculate the corrected offset for each dimension, highlighting key dimensions, spatial correction is used to adapt to the object type, and absolute offset is used to quantify the single-dimensional deviation.
[0150] For the summation term, It sums up the corrected offset for each dimension;
[0151] Point object alignment involves summing the correction offsets in each dimension and combining them with spatial correlation correction to obtain the alignment offset in the spatial position dimension, making the alignment offset more in line with the actual impact (small offsets have a small impact on distant objects, while large offsets have a large impact on close objects).
[0152] By introducing (Spatial correlation) resolves the difference between small offsets of distant objects and small offsets of nearby objects. For example, when two points are 10km apart, a 1m offset on the X-axis is acceptable; however, when they are 1m apart, a 1m offset on the X-axis is a serious misalignment. It reduces the offset between two points, making it suitable for spatial alignment (such as polar coordinate alignment or latitude and longitude alignment of targets). It converts linear dimensional offset into spatial dimensional offset calculation, avoiding errors in proximity distance.
[0153] Example 6;
[0154] Based on Example 2, in step S4, the micro-positional offset alignment technique is based on the differences in spatial feature representations of multimodal data. It uses model learning to predict the positional offsets caused by these differences. Then, the data is transformed according to the predicted offsets to better match the data from different modalities in the positional domain. For example, in remote sensing image fusion, there is a spatial resolution difference between a high-resolution panchromatic image (PAN) and a low-resolution multispectral image (MS). This technique can learn the offset of each pixel in the PAN image relative to its corresponding position in the MS image, thereby adjusting the PAN image to achieve fine alignment with the MS image.
[0155] The specific steps for transforming multimodal data in the continuous location domain are as follows:
[0156] Step a: Feature extraction and offset prediction; extract features from multimodal data and then predict the positional offsets between features. For example, the Modal Aware Feature Alignment Network (MFAN) aligns the features of PAN and MS and then predicts the spatial transformation offset of each position.
[0157] Step b: Interpolation sampling and transformation; based on the predicted offset, the original features are interpolated to achieve fine-grained spatial alignment. For example, MFAN uses a learnable interpolation function (LIF) to perform bilinear interpolation by transforming the offset at the pixel level, thus finely aligning the high-frequency spatial structure (PAN) with the low-frequency spectral information (MS).
[0158] Step c: Progressive alignment; Alignment is performed in a multi-stage, progressive manner. The offset result of the previous stage will affect the offset prediction of the next stage, making the alignment more refined and stable, avoiding the loss of high-frequency details and the accumulation of cascading errors. This is more advantageous than the unidirectional deformable convolution strategy.
[0159] Specifically, multi-stage progressive adjustment and step-by-step alignment involve further processing at each stage based on the results of the previous stage. For example, the bidirectional progressive feature alignment (BSFA) strategy proposed in BSA Fusion predicts the deformation field between features layer by layer through forward and backward alignment layers. This multi-level progressive alignment operation does not complete the alignment all at once, but allows the features to be progressively adjusted during the alignment process.
[0160] The offset results from the previous stage will serve as an important basis for the offset prediction in the next stage. For example, when dealing with the problem of remote sensing resolution mismatch, the offset obtained from the previous stage is used to interpolate and sample the original features, and then the offset prediction in the next stage is performed based on this, so that the alignment can be continuously refined.
[0161] Because the alignment is performed incrementally, each adjustment is relatively small, avoiding excessive changes to image features as in single-step alignment, thus better preserving high-frequency details in the image. For example, the bidirectional incremental feature alignment strategy predicts the deformation field incrementally in both forward and backward directions. Each predicted deformation field is a small adjustment based on the current feature state, avoiding the loss of high-frequency details caused by a large-scale deformation at once.
[0162] Each stage builds upon the accuracy of the previous stage. Even if there are errors in the previous stage, they can be corrected through continuous adjustments in subsequent stages, unlike unidirectional deformable convolution strategies where errors accumulate with each deformation. For example, in BSAsuion's bidirectional progressive feature alignment, progressive predictions in both the forward and backward directions complement and correct each other, effectively suppressing the cascading accumulation of errors.
[0163] When applied to remote sensing image fusion, it can accurately align and fuse high-resolution PAN images with low-resolution MS images to generate images with both high spatial and spectral resolution, thereby improving the quality and application value of remote sensing images.
[0164] When applied to visual multimodal alignment, fine-grained spatial alignment can be achieved through micro-positional offset alignment technology during data processing, enhancing the understanding of deep relationships between visual positions and improving task performance.
[0165] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope described in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A deep learning-based adaptive alignment method for multimodal sensor data, characterized in that, Includes the following steps: Step S1. Acquire spatial point data from multiple heterogeneous sensors to obtain a multimodal sensor dataset; wherein, the multiple heterogeneous sensors include LiDAR, vision camera, millimeter-wave radar and inertial measurement unit; Step S2. Construct a multi-path feature extraction network model to extract spatial convolutional feature representations of sensor data from different modalities; wherein, each path of the multi-path feature extraction network model adopts a vector-aligned point convolutional network to adapt to the positional variation characteristics of different sensors and output the optimal positional offset between each modal data. In step S2, the working steps of the multi-path feature extraction network are as follows: The input layer receives sensor data in different modes. The path feature extraction method is that each path includes: vector-aligned point convolutional layer, batch normalization layer, ReLU activation function and max pooling layer, and feature extraction is performed on the data in each layer; The output layer outputs the fused feature representation and the optimal position offset between each mode; The spatial convolution features of the sensor data from different modalities are represented as follows: After inputting data from different modalities, the point convolution calculation formula is: ; in, After the convolution operation, the output features are located at... The value at that location; The number of modes of the convolution kernel; convolution kernel In modality Channel Index Adjustment parameters at the location; , The respective Modality and the first Modal range, the modal range is [1, ]; It is the channel index of the input feature. Total number of channel indexes; Input features In position Channel Index The value at that location; This is a position vector, representing the feature location. Corresponding to the physical location information of the sensor; For position vectors The processing function; For each neighborhood location Calculate the convolution kernel weights × Input feature value Then put all Add the results together, then multiply by the position function. Let the final result It also includes spatial neighborhood information of the input features and sensor location information; Step S3. Construct a dynamic adaptive alignment model. Input the extracted spatial convolutional features into the dynamic adaptive alignment model. Calculate the cross-modal correlation between different modal features using the position vector feature sequence dimension normalization method. Dynamically adjust the alignment offset of multimodal features in the spatial position dimension based on the cross-modal correlation. In step S3, the alignment offset in the spatial position dimension is applicable to point object alignment. The core is to sum the absolute offsets in each dimension and combine it with spatial correlation correction to avoid a single-dimensional offset from masking the overall misalignment. ; in, It is the base space alignment offset, the final calculation result, representing the object. and The numerical value of the overall degree of misalignment in space; The spatial correlation between the baseline and the target To align with the reference object, Align the target object; For dimension Alignment distance; For dimension Geometric property correction coefficients; For target object In dimensions Feature parameters on; To align the reference object In dimensions Feature parameters on; For dimension offset, single dimension Above, target object With reference object The absolute offset is used to eliminate the influence of direction and determine the magnitude of the deviation. To calculate the corrected offset for each dimension, highlighting key dimensions, spatial correction is used to adapt to the object type, and absolute offset is used to quantify the single-dimensional deviation. For the summation term, It sums up the corrected offset for each dimension; Point object alignment involves summing the correction offsets in each dimension and combining them with spatial correlation correction to obtain the alignment offset in the spatial position dimension, making the alignment offset more closely reflect the actual impact. Step S4. Utilize micro-position offset alignment technology to perform continuous position domain transformation on multimodal data to achieve fine-grained alignment of multimodal data.
2. The deep learning-based adaptive alignment method for multimodal sensor data according to claim 1, characterized in that, In step S1, when acquiring the spatial location data of each of the multiple heterogeneous sensors, the locations of each heterogeneous sensor are fused and obtained as follows: : ; in, for The spatial point vector of the final output after time-mapping; For the first Sensors at each location Dynamic adjustment coefficient at any given time; The total number of sensors participating in the fusion; It involves summing the reliability of the point vector for each sensor; It is the first A data source or sensor at any time The original vector; The normalization factor represents the spatial point vector of each sensor.
3. The deep learning-based adaptive alignment method for multimodal sensor data according to claim 1, characterized in that, The formula for calculating the alignment loss of the optimal position offset is: ; in, For the first Modality and the first Alignment loss of modal features is used to measure the degree of spatial alignment between two modal features; and The first and the Modality; and The spatial dimension of the feature is represented as ; For the first Modal features in spatial location The eigenvector at that location; For the first Mode relative to the first Spatial offset of the mode, This represents the spatial offset in the horizontal direction. This represents the spatial offset in the vertical direction. The square of the Euclidean norm or the square of the L2 norm is used to calculate the distance between two eigenvectors. The smaller the distance, the more similar the two features are. No. Modal features With the Features after mode shift The square of the L2 norm is ; Specifically, the optimal offset The calculation is as follows: ; in, , and These are the offsets along the x, y, and z axes, respectively. The argmin operation finds the variable values that minimize the objective function, and finds a set of... Let the loss Minimum.
4. The deep learning-based adaptive alignment method for multimodal sensor data according to claim 1, characterized in that, In step S3, the dimension normalization of the position vector feature sequence is an ordered set composed of temporally sampled position vectors, denoted as... ,in, The sequence length; Specifically, the dimension normalization method for the position vector feature sequence is as follows: ; in, For the first The mean of the nearest neighbor distances of each convolution kernel mode; For the first The set of points representing the modalities of each convolutional kernel; for The number of points included; for The first in One point, for The first in One point, for Nearest neighbor points The nearest neighbor set; For point To the nearest point The distance; Point Nearest neighbor search: find the points that are closest to it. , forming a set ; Traverse the set Each point in Accumulate; By calculating the sum of the average nearest distances of all points, the core concept is to quantify the density distribution characteristics of the points. This involves calculating the average local nearest distance of a single point to reflect the density of its surroundings. To make the result independent of the number of nearest neighbors; For small point values, the points are densely packed; for large point values, the points are sparsely packed. This is used for scene recognition. This makes the characteristics of the points comparable.
5. The deep learning-based adaptive alignment method for multimodal sensor data according to claim 1, characterized in that, In step S4, the micro-position offset alignment technique is based on the differences in spatial feature representation of multimodal data. It uses model learning to predict the position offset caused by these differences, and then transforms the data according to the predicted offset so that the data of different modalities can be better matched in the position domain.
6. The deep learning-based adaptive alignment method for multimodal sensor data according to claim 5, characterized in that, The multimodal data undergoes continuous location domain transformation, and the specific steps are as follows: Step a: Feature extraction and offset prediction; extract features from multimodal data and then predict the positional offsets between features; Step b: Interpolation sampling and transformation; based on the predicted offset, the original features are interpolated to achieve fine-grained spatial alignment; Step c: Progressive alignment; Alignment is performed in a multi-stage, progressive manner. The offset result of the previous stage will affect the offset prediction of the next stage, making the alignment more refined and stable, and avoiding the loss of high-frequency details and the accumulation of cascading errors.
7. A deep learning-based adaptive alignment system for multimodal sensor data, employing the deep learning-based adaptive alignment method for multimodal sensor data according to any one of claims 1-6, characterized in that... include: The data acquisition module is used to acquire spatial point data collected by multiple heterogeneous sensors, and integrate the spatial point data after preprocessing to form a multimodal sensor dataset; the preprocessing includes data denoising, outlier removal and format standardization. The multi-path feature extraction module, connected to the data acquisition module, has a built-in multi-path feature extraction network model. Each path of the multi-path feature extraction network model is configured with a vector-aligned point convolutional network. The vector-aligned point convolutional network can adapt to the positional change characteristics of the corresponding heterogeneous sensors. The multi-path feature extraction module extracts features from the sensor data of different modalities in the multimodal sensor dataset through the multi-path feature extraction network model to obtain the spatial convolutional feature representation of each modal sensor data, and outputs the optimal positional offset between each modal data. The dynamic adaptive alignment module is connected to the multi-path feature extraction module and constructs a dynamic adaptive alignment model. The dynamic adaptive alignment model receives the spatial convolution features output by the multi-path feature extraction module, uses the position vector feature sequence dimension normalization method to unify the dimensions of the spatial convolution features of different modalities, calculates the cross-modal correlation between the features of different modalities based on the dimension-unified features, and dynamically adjusts the alignment offset of the multimodal features in the spatial position dimension according to the cross-modal correlation. The fine-grained alignment execution module, connected to the dynamic adaptive alignment module, employs a micro-position offset alignment technique. Based on the alignment offset adjusted by the dynamic adaptive alignment module, it performs continuous position domain transformation on the multimodal data in the multimodal sensor dataset to achieve fine-grained alignment of the multimodal data.
Citation Information
Patent Citations
Fine-grained action recognition method based on cross-modal knowledge alignment
CN118196888A
High-precision multi-mode sensor space alignment method in road monitoring
CN120164173A