A cross-modal target perception method and system based on physical constraint driving

CN122595240BActive Publication Date: 2026-09-29GUANGZHOU INSTITUTE OF TECHNOLOY XIDIAN UNIVERSITY +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202611087983.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-22
Publication Date
2026-09-29
Estimated Expiration
2046-07-22

AI Technical Summary

Technical Problem

[0004]然而,现有的跨模态多源信息融合方法仍然存在若干亟待解决的关键问题

Benefits of technology

本发明通过获取目标的多模态传感器数据并进行特征提取与统一语义空间映射,进一步基于空间结构特征构建空间约束掩码,并在空间约束掩码的引导下进行多路径差分运算,主动抑制了复杂环境下的跨模态非一致性共模噪声;同时,结合基于信号传播机理的物理约束模型对退化、畸变的频域特征进行反演与一致性修正,显著提升了模型在雨雾、低信噪比等极端恶劣场景下的特征鲁棒性与抗干扰能力,适用于复杂动态环境下的智能驾驶感知与多模态避障任务;最后基于具备高度物理一致性的最终融合特征进行目标感知与辨识,得到高置信度的目标识别结果,进而大幅提高了自动驾驶系统的避障安全性和感知精度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122595240B_ABST
    Figure CN122595240B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of intelligent driving perception and multi-modal information fusion, and relates to a cross-modal target perception method and system based on physical constraint driving, which comprises the following steps: acquiring multi-modal sensor data, respectively extracting features according to modes and mapping to a unified semantic space to obtain spatial structure features and frequency domain signal features; constructing a spatial constraint mask based on the spatial structure features, and performing difference operation on the multi-path mapping responses of the spatial structure features and the frequency domain signal features under the guidance of the spatial constraint mask to obtain preliminary difference results; based on a physical constraint model of signal propagation mechanism, the preliminary difference results are inverted and consistency corrected to obtain final fusion features; and the final fusion features are subjected to cross-modal target perception and recognition to output target recognition results. The method can improve the perception robustness and anti-interference ability in complex traffic scenes through spatial-guided difference fusion and numerical-physical dual-driven physical correction mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent driving perception and multimodal information fusion technology, specifically involving a cross-modal target perception method and system based on physical constraints. Background Technology

[0002] The core objective of target perception is to continuously and accurately detect and identify specific obstacle targets in complex and dynamic environments. Existing environmental perception methods mainly rely on purely data-driven deep learning models, combining features extracted from single sensors or conventional multi-source data to achieve target recognition. However, purely data-driven methods have poor robustness in complex scenarios, especially in extreme weather conditions such as rain, snow, and dense fog, or in environments with low signal-to-noise ratios. They are highly susceptible to interference from environmental noise, signal attenuation, and target occlusion, resulting in weak generalization ability and a high likelihood of false alarms or missed alarms.

[0003] Cross-modal multi-source information fusion, as an important research direction in the field of target perception, provides a new solution to the challenges faced by traditional pure data-driven perception by introducing multi-dimensional heterogeneous information provided by spatial structural data and frequency domain signal data. Existing multi-modal fusion methods can be mainly divided into three categories: early pre-fusion methods based on the data level, which directly stitch together the original sensor data through spatial coordinate alignment; mid-fusion methods based on the feature level, which use deep neural networks to extract features from each modality and then perform stitching or cross-attention mechanism calculations; and post-fusion methods based on the result level, which perform logical judgment and cascade on the detection results output independently by each modality. These methods improve the ability of the perception system to capture target features to varying degrees. Especially when there are drastic changes in illumination or when a single sensor is limited, the complementarity of multi-source heterogeneous data can effectively assist target perception decision-making.

[0004] However, existing cross-modal multi-source information fusion methods still face several critical issues that urgently need to be addressed. First, regarding feature fusion mechanisms, most methods employ simple feature channel splicing or shallow attention mechanisms, failing to fully explore the deep physical mechanisms linking spatial structure data and frequency domain signal data. This makes them highly susceptible to non-uniform common-mode noise interference in complex dynamic scenarios. Second, when facing degraded signals, existing purely data-driven deep learning models lack explicit physical constraints. When frequency domain signals experience severe energy attenuation or waveform distortion due to harsh environments, the network cannot actively compensate for and recover features, resulting in weak system generalization capabilities and a high likelihood of false positives or false negatives. Furthermore, existing methods often lack mechanisms to use geometric spatial structure information as prior guidance, making it difficult to accurately constrain the fusion range of frequency domain features using spatial occupancy boundaries. This leads to severely insufficient efficiency and specificity in background noise suppression, making the system ill-suited for the stringent requirements of high-reliability obstacle avoidance in scenarios such as autonomous driving. Summary of the Invention

[0005] To address the aforementioned problems in the existing technology, this invention provides a cross-modal target perception method and system based on physical constraint driving. The technical problem to be solved by this invention is achieved through the following technical solution: This invention provides a cross-modal target perception method based on physical constraints, comprising the following steps: Multimodal sensor data representing the spatial and frequency domain attributes of the target are acquired, and features are extracted according to the modality and mapped to a unified semantic space to obtain spatial structure features and frequency domain signal features. Based on the geometry and position of the spatial structural features, a spatial constraint mask is constructed, and under the guidance of the spatial constraint mask, a differential operation is performed on the multipath mapping response of the spatial structural features and the frequency domain signal features to obtain preliminary differential results. Based on the physical constraint model of signal propagation mechanism, the preliminary difference results are inverted and consistency corrected to obtain the final fusion features; The final fused features are used to perceive and identify targets across modalities, and the target recognition results are output.

[0006] In one embodiment of the present invention, multimodal sensor data characterizing the spatial and frequency domain attributes of a target are acquired, and features are extracted according to each mode and mapped to a unified semantic space to obtain spatial structure features and frequency domain signal features, including: Acquire spatial structural data characterizing the target's geometry and position, and frequency domain signal data characterizing the target's physical properties; The spatial structure data is feature-encoded using a spatial feature extraction module to obtain spatial extracted features. The frequency domain signal data is feature-encoded using a frequency domain feature extraction module to obtain frequency domain extracted features; By performing feature reduction and projection operations, the spatial extracted features and the frequency domain extracted features are aligned to a feature space with the same channel dimension to obtain the spatial structure features and the frequency domain signal features.

[0007] In one embodiment of the present invention, a spatial constraint mask is constructed based on the geometry and position of the spatial structural features, and a differential operation is performed on the multipath mapping response of the spatial structural features and the frequency domain signal features under the guidance of the spatial constraint mask to obtain a preliminary differential result, including: Based on the spatial structure features, the three-dimensional spatial occupancy state and geometric shape distribution prior of the target are extracted, and a spatial constraint mask is constructed to indicate the potential existence area of ​​the target. The spatial structure features and the frequency domain signal features are subjected to nonlinear feature transformation using the first feature mapping function to generate a hybrid correlation feature response that includes the target matching signal and environmental background interference; The spatial structure features and the frequency domain signal features are subjected to nonlinear feature transformation using the second feature mapping function to generate a common-mode noise feature response for capturing cross-modal common-mode noise distribution in complex environments; The spatial constraint mask is used to perform spatial guided modulation on the hybrid correlation feature response, and the spatially guided modulated hybrid correlation feature response is differentially subtracted from the common-mode noise feature response to obtain the preliminary difference result.

[0008] In one embodiment of the present invention, the spatial constraint mask is obtained by extracting geometric and topological features from spatial structural features and concatenating them with spatial structural features, then performing convolutional mapping on the concatenated features through a learnable spatial convolutional network, and finally outputting the result after passing an activation function. The first feature mapping function performs a nonlinear transformation by concatenating spatial structural features and frequency domain signal features through channels, and outputs the hybrid correlation feature response through an activation function; The second feature mapping function performs a nonlinear transformation after calculating the absolute difference between the spatial structure features and the frequency domain signal features, and outputs the common-mode noise feature response through an activation function; The preliminary difference result is obtained by performing element-wise multiplication of the hybrid correlation feature response output by the first feature mapping function with the enhancement factor of the spatial constraint mask, weighting the common-mode noise feature response output by the second feature mapping function with the noise cancellation coefficient obtained by dynamic learning, and subtracting the weighted feature from the feature after multiplication.

[0009] In one embodiment of the present invention, a physical constraint model based on the signal propagation mechanism is used to invert and correct the consistency of the preliminary difference results to obtain the final fusion features, including: Obtain the current environmental context information and target surface reflector properties, and call the preset physical constraint model that includes signal attenuation characteristics; The local feature energy in the preliminary difference results is analyzed and the feature distortion degree index is calculated. Based on the feature distortion degree index, the abnormal frequency response region caused by complex environmental interference is located. Combining the preliminary difference results, the current environmental context information, and the target surface reflector properties, the physical constraint model is used to perform feature inversion and compensation processing on the abnormal frequency response region to generate the final fused features.

[0010] In one embodiment of the present invention, the local feature energy is characterized by the sum of squares of the L2 norm of the preliminary difference results of each coordinate point within a local sensing window centered on the spatial coordinates. The feature distortion index is calculated based on the ratio of the local feature energy to the lossless distribution of the theoretical target frequency domain feature energy based on spatial prior estimation. When the feature distortion index is greater than the preset hard threshold for distortion judgment, the region corresponding to the local sensing window at the center of the spatial coordinates is taken as the abnormal frequency response region. The final fusion feature is obtained by multiplying the preliminary difference result by an exponential compensation term, which is calculated based on environmental context information, the retroreflection coefficient of the target surface, the radial physical distance, and the detection wavelength of the multimodal sensing device.

[0011] In one embodiment of the present invention, the final fused features are used for cross-modal target perception and identification, and a target recognition result is output, including: The final fused features are restored to tensor input format through feature reshaping and dimension transformation operations to obtain the transformed feature tensor. The transformed feature tensors are stacked using a multi-layer network structure, and the stacked features are input into the localization head detection module. The module outputs a classification score map representing the probability of the target's existence, a bounding box size map for predicting the target's 3D size, and a local position offset map for eliminating downsampling discretization errors through three parallel decoding branches. The location with the highest score is selected as the target center based on the classification score map. Combined with the bounding box size map and the local position offset map, the bounding box pose and category attributes of the target in the real physical environment are restored through spatial geometric transformation equations to obtain the target recognition result.

[0012] In one embodiment of the present invention, the three parallel decoding branches include a classification branch, a regression branch, and an offset branch; The classification branch maps the transformed feature tensor through a convolutional network and outputs a classification score map representing the probability of the target's existence through a Sigmoid activation function. The regression branch maps the transformed feature tensor through a convolutional network and outputs a bounding box size map for predicting the three-dimensional size of the target. The offset branch maps the transformed feature tensor through a convolutional network, outputting a local position offset map to eliminate downsampling discretization errors.

[0013] Another embodiment of the present invention provides a cross-modal target perception system based on physical constraints, comprising: The data acquisition and mapping module is used to acquire multimodal sensor data that characterizes the spatial and frequency domain attributes of the target, extract features according to the mode and map them to a unified semantic space to obtain spatial structure features and frequency domain signal features. The differential fusion module is used to construct a spatial constraint mask based on the geometry and position of the spatial structural features, and perform differential operations on the multipath mapping response of the spatial structural features and the frequency domain signal features under the guidance of the spatial constraint mask to obtain preliminary differential results. The physical constraint module is used to perform inversion and consistency correction on the preliminary difference results based on the physical constraint model of signal propagation mechanism to obtain the final fusion features; The positioning head detection module is used to perceive and identify cross-modal targets based on the final fused features, and output the target recognition result.

[0014] In one embodiment of the present invention, the overall joint optimization objective function of the cross-modal target perception system is composed of a weighted sum of a data-driven loss term and a physical constraint loss term; The data-driven loss term is composed of a weighted combination of the focus loss function for target classification and the geometric regression loss function for the three-dimensional size and position offset of the bounding box; The physical constraint loss term is obtained by calculating the difference between the squared L2 norm of the final fusion feature output by the network forward propagation and the preset analytical benchmark function of wave attenuation and optical scattering theory, and then weighted by dynamic penalty weight control.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention acquires multimodal sensor data of the target and performs feature extraction and unified semantic space mapping. It further constructs a spatial constraint mask based on spatial structural features and performs multipath differential operations under the guidance of the spatial constraint mask, actively suppressing cross-modal non-uniform common-mode noise in complex environments. Simultaneously, it combines a physical constraint model based on signal propagation mechanisms to invert and correct the consistency of degraded and distorted frequency domain features, significantly improving the model's feature robustness and anti-interference capability in extreme and harsh scenarios such as rain, fog, and low signal-to-noise ratio. This makes it suitable for intelligent driving perception and multimodal obstacle avoidance tasks in complex dynamic environments. Finally, it performs target perception and identification based on the highly physically consistent final fused features, obtaining high-confidence target recognition results, thereby significantly improving the obstacle avoidance safety and perception accuracy of the autonomous driving system. Attached Figure Description

[0016] Figure 1 A flowchart illustrating the cross-modal target perception method based on physical constraints provided in an embodiment of the present invention; Figure 2A framework diagram of a cross-modal target perception method based on physical constraints provided in an embodiment of the present invention; Figure 3 The diagram shows the structure of a cross-modal target perception system based on physical constraints, as provided in an embodiment of the present invention. Detailed Implementation

[0017] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.

[0018] Example 1 This invention proposes a cross-modal target perception method based on physical constraints. This method maps multimodal sensor data to a unified semantic space, designs a multi-path differential fusion mechanism guided by prior spatial structure, and couples a physical constraint model based on signal propagation mechanism. It effectively solves the shortcomings of existing pure data-driven methods in terms of missing physical mechanisms, cross-modal non-uniform noise interference, and degradation signal repair. It has stronger scene adaptability and extremely high perception confidence, especially showing significant robustness advantages in complex and ever-changing real-world autonomous driving obstacle avoidance scenarios such as extreme weather.

[0019] Please see Figure 1 and Figure 2 , Figure 1 This is a flowchart illustrating the cross-modal target perception method based on physical constraints provided in an embodiment of the present invention. Figure 2 This is a framework diagram of a physical constraint-driven cross-modal target perception method provided in an embodiment of the present invention. The method includes the following steps: S100. Acquire multimodal sensor data representing the spatial and frequency domain attributes of the target, extract features according to the mode and map them to a unified semantic space to obtain spatial structure features and frequency domain signal features.

[0020] S110. Acquire spatial structural data characterizing the geometry and position of the target and frequency domain signal data characterizing the physical properties of the target.

[0021] Specifically, multi-source heterogeneous sensing devices are deployed to simultaneously collect sensing data of the target's environment. The first type of sensing device acquires spatial structure data representing the target's geometry and position, which can be a three-dimensional point cloud or depth image. The second type of sensing device acquires frequency domain signal data representing the target's physical properties, which can be an active near-infrared emission coded waveform or a frequency domain waveform.

[0022] S120. Use the spatial feature extraction module to encode the spatial structure data to obtain spatial extracted features; use the frequency domain feature extraction module to encode the frequency domain signal data to obtain frequency extracted features; through feature dimensionality reduction and projection operations, align the spatial extracted features and frequency extracted features to a feature space with the same channel dimension to obtain spatial structure features and frequency domain signal features.

[0023] Specifically, for spatial structure data, pre-trained 3D sparse convolutional networks or point cloud Transformers can be used to extract local geometric and global topological features to obtain spatial extraction features; for frequency domain signal data, one-dimensional convolution can be used to extract time-frequency domain sequence features to obtain frequency domain extraction features.

[0024] Assume the multimodal sensor data is spatial structure data. With frequency domain signal data The corresponding feature extraction module and projection function are used to align them to the same dimensional space, as shown in the following formula:

[0025]

[0026] in, These are the spatial structural features and frequency domain signal features mapped to a unified semantic space, respectively. For the set of real numbers, For feature map height, The width of the feature map. The number of feature map channels. This represents a feature extraction operator for spatial structure data. This represents a feature extraction operator for frequency domain signal data. This represents the feature dimensionality reduction and spatial alignment projection operations performed on the extracted spatial features. This represents the feature dimensionality reduction and spatial alignment projection operations performed on the features extracted in the frequency domain. and This is used to ensure that the two types of heterogeneous features are consistent across the channel dimension C and the spatial resolution dimension. Strict tensor alignment is achieved on top.

[0027] In addition, during the training process, the collected sample data is enhanced by random occlusion, point cloud jitter, and signal noise addition to improve the model's generalization ability in complex traffic scenarios.

[0028] S200. Construct a spatial constraint mask based on the geometry and position of the spatial structure features, and perform differential operations on the multipath mapping response of the spatial structure features and frequency domain signal features under the guidance of the spatial constraint mask to obtain preliminary differential results.

[0029] S210. Based on the spatial structure features, extract the three-dimensional spatial occupancy state and geometric shape distribution prior of the target, and construct a spatial constraint mask to indicate the potential existence area of ​​the target.

[0030] Specifically, geometric feature analysis is performed on the spatial structure features mapped to the unified semantic space to determine the three-dimensional physical boundary of the target in the real environment, thus obtaining the target's three-dimensional spatial occupancy state; and local surface geometric features and topological relationships of the target are extracted to obtain the target's geometric shape distribution prior; and a multi-dimensional spatial constraint mask is generated by combining the three-dimensional spatial occupancy state and the geometric shape distribution prior. The spatial constraint mask is obtained by extracting geometric and topological features from spatial structural features, concatenating these features with other spatial structural features, performing convolutional mapping on the concatenated features through a learnable spatial convolutional network, and then outputting the result after an activation function. The expression is:

[0031] in, For spatial constraint mask, It is a spatial structural feature; Operators for extracting geometric and topological features are used to further extract local surface normals and spatial connectivity priors of the target from spatial structural features; This indicates a splicing operation along the channel dimension. For learnable spatial convolutional network weights, For learnable spatial convolutional network bias parameters, Represents spatial convolution operation; The sigmoid activation function is used to normalize the output tensor to 0. This interval is used to generate a soft mask that represents the probability of the target's potential spatial occupancy. .

[0032] In this embodiment, by processing multimodal spatial structure data such as lidar point clouds or millimeter-wave radar, the perception model can initially capture the approximate spatial outline of small obstacle avoidance cones or construction vehicles ahead. The generated spatial constraint mask is equivalent to applying an attention focusing box to the feature map, pre-delineating the high-value physical space range for subsequent cross-modal feature fusion, thereby effectively avoiding the waste of computing power caused by global computation and isolating a large amount of background clutter interference from the spatial source.

[0033] S220. Input the spatial structure features and frequency domain signal features into the differential fusion model, and generate the hybrid correlation feature response and common mode noise feature response respectively through the multi-path mapping mechanism.

[0034] Specifically, two parallel feature mapping network structures, the first and the second, are constructed within the differential fusion model. In the first feature mapping network structure, a first feature mapping function is used to perform nonlinear feature transformation on spatial structural features and frequency domain signal features to generate a hybrid correlated feature response that includes the target matching signal and environmental background interference. In the second feature mapping network structure, a second feature mapping function is used to perform nonlinear feature transformation on spatial structural features and frequency domain signal features to generate a common-mode noise feature response for capturing cross-modal common-mode noise distribution in complex environments.

[0035] The first feature mapping function performs a nonlinear transformation by concatenating spatial structure features and frequency domain signal features through channels, and outputs a hybrid correlation feature response after activation. The second feature mapping function performs a nonlinear transformation by calculating the absolute difference between spatial structure features and frequency domain signal features, and outputs a common-mode noise feature response after activation. The expression is:

[0036]

[0037] in, Let represent the first feature mapping function, which aims to extract cross-modal joint feature responses (which are mixed with real target matching signals and environmental background interference) through nonlinear mapping. For frequency domain signal characteristics, This indicates a cross-modal feature concatenation operation. For activation function, The weight parameters of the first feature mapping function are... The bias parameter of the first feature mapping function; This represents the second feature mapping function, which calculates the absolute value of the feature difference between spatial structure features and frequency domain signal features. It captures cross-modal non-uniform common-mode noise features caused by occlusion or attenuation due to extreme environments; For activation function, The weight parameters of the second feature mapping function are... is the bias parameter of the second feature mapping function.

[0038] In this embodiment, the first feature mapping network structure is mainly used to extract conventional feature activation states, i.e., signal plus noise, while the second feature mapping network structure is designed to focus on common-mode noise in traffic environments, such as scattering and obstruction caused by rain and fog, or radar clutter caused by reflections from ground water stains. This dual design of multi-path mapping provides a preliminary feature foundation for achieving "signal and noise separation" in a high-dimensional space with extremely low signal-to-noise ratio.

[0039] S230. Spatial guidance modulation is performed on the hybrid correlation feature response using a spatial constraint mask, and the hybrid correlation feature response after spatial guidance modulation is differentially subtracted from the common-mode noise feature response to obtain a preliminary difference result.

[0040] Specifically, spatial guidance modulation is applied to the hybrid correlation feature response using a spatial constraint mask to explicitly enhance the feature weights of target candidate regions within the spatial constraint mask range and suppress the weights of non-target regions. Subsequently, the spatially guided modulated hybrid correlation feature response is subtracted from the common-mode noise feature response to actively remove redundant background noise components. The hybrid correlation feature response output by the first feature mapping function is multiplied element-wise with the enhancement factor of the spatial constraint mask. The common-mode noise feature response output by the second feature mapping function is then weighted using dynamically learned noise cancellation coefficients. The preliminary difference result is obtained by subtracting the weighted feature from the multiplied feature. The dynamic calculation formula for this step is as follows:

[0041] in, This indicates the preliminary difference results. This represents element-wise multiplication. This represents the noise reduction coefficient for dynamic learning.

[0042] In this embodiment, the model strictly limits the scope of the differential operation through a spatial guidance mechanism. During differential subtraction, the model actively cancels out common-mode background noise that disrupts consistency. Due to the guidance and protection of the spatially constrained mask, real obstacle avoidance target signals, such as specifically coded near-infrared signals or high-reflectivity radar echoes, are fully preserved. This step ultimately outputs a clean reconstructed feature with an extremely high signal-to-noise ratio, providing reliable input for subsequent physical corrections.

[0043] S300, a physical constraint model based on signal propagation mechanism, performs inversion and consistency correction on the preliminary difference results to obtain the final fusion features.

[0044] S310. Obtain the current environmental context information and target surface reflector attributes, and call the preset physical constraint model that includes signal attenuation characteristics.

[0045] Specifically, environmental context information in the current scene is extracted, including rain and snow concentration, as well as physical reflection properties of the target surface, such as retroreflection coefficient. These multidimensional parameters are used together as external prior conditions and input into a pre-defined physical constraint model that includes wave signal attenuation curves.

[0046] Context information It can be quantified as follows:

[0047] in, Atmospheric visibility (km) The rainfall or snowfall rate is denoted as mm / h, and T is the ambient temperature.

[0048] Meanwhile, the physical reflection properties of the target surface (retroreflection coefficient) The extraction of ) is calculated based on the radar or optical sensor echo equation:

[0049] in, The return intensity or echo power of the point cloud recorded in the spatial structure data. These are the factory calibration constants for the sensing device system. To detect the radial physical distance of the target relative to the sensing device. The angle between the target surface normal vector and the incident signal ray is denoted as .

[0050] In this embodiment, the physical constraint model serves as the theoretical benchmark for the sensing algorithm. It quantifies the theoretical attenuation of active near-infrared signals or millimeter-wave radar signals under specific adverse weather conditions based on physical laws. By inputting the real-time environmental context into this model, the system can dynamically calculate the current theoretical attenuation factor, thereby preparing accurate compensation parameters for the distorted feature representation.

[0051] S320. Analyze the local feature energy in the preliminary difference results and calculate the feature distortion index. Based on the feature distortion index, locate the abnormal frequency response region caused by complex environmental interference.

[0052] Specifically, a feature-level scan is performed on the preliminary differential results after differential noise reduction to extract the frequency domain signal feature components that have experienced sharp energy attenuation, phase shift, or severe waveform distortion due to extreme weather, signal blockage, or low signal-to-noise ratio environment, thereby locating the abnormal frequency response region.

[0053] For determining severe distortion, the system quantifies the degree of attenuation of local feature energy by calculating it through the frequency anomaly detection module. For the lossless distribution of the frequency domain characteristic energy of the theoretical target based on spatial prior estimation, For the preliminary difference results extracted The local feature energy is represented by the sum of squared L2 norms of the preliminary difference results for each coordinate point within a local sensing window centered on the spatial coordinates. The expression is:

[0054] in, Represented by spatial coordinates A local perception window centered on the subject. Representing coordinates Preliminary difference results.

[0055] Furthermore, a feature distortion degree index is defined. The ratio of the lossless distribution of the theoretical target frequency domain characteristic energy based on local feature energy to that based on spatial prior estimation is calculated, and the expression is:

[0056] in, To prevent smoothing terms with a denominator of zero, when the characteristic distortion index... At that time, in spatial coordinates The region corresponding to the central local sensing window is the abnormal frequency response region. The frequency domain characteristics of this region are severely distorted due to complex environmental interference. It is necessary to extract this feature component and feed it into the downstream physical constraint model for correction. This is a preset hard threshold for distortion determination, typically set between 0.5 and 0.7.

[0057] In this embodiment, a frequency anomaly detection module is used to screen the preliminary differential results. Since the preceding differential fusion step has already cleared most of the common-mode environmental clutter, the remaining features are mainly caused by energy loss during physical propagation, such as active infrared pulses absorbed by fog or attenuated radar waves. Accurately locating these areas ensures that subsequent physical compensation is more targeted and avoids over-correction of normal features.

[0058] S330. Combining the preliminary difference results, current environmental context information, and target surface reflector properties, the physical constraint model is used to perform feature inversion and compensation processing on the abnormal frequency response region, generating the final fused features with physical consistency.

[0059] Specifically, the signal inversion and consistency correction algorithm within the physical constraint model is invoked to perform energy compensation and phase correction processing on the located abnormal frequency domain response regions. Considering that active near-infrared light waves or millimeter-wave radar signals exhibit severe exponential attenuation under complex traffic environments such as rain and fog, this embodiment uses a physical inversion compensation formula that includes a composite atmospheric extinction coefficient for consistency correction. After feature inversion and compensation processing, the final fused features are obtained. The final fused features are obtained by multiplying the preliminary difference results by an exponential compensation term. The exponential compensation term is jointly calculated based on environmental context information, the retroreflection coefficient of the target surface, the radial physical distance, and the detection wavelength of the multimodal sensing device, and its expression is:

[0060]

[0061]

[0062] in, This represents the final fused feature output after physical mechanism compensation. This represents the initial difference result of the input, which contains the anomalous frequency response region. This represents the exponential compensation term for constructing physical decay. This represents the adjustable compensation gain coefficient of the network. This represents the theoretical attenuation coefficient calculated by combining environmental and material properties. This indicates the radial physical distance of the target relative to the sensing device (i.e., the relative physical distance of the target). This represents the retroreflection coefficient of the target surface. This represents the atmospheric extinction coefficient function. Represents environmental context information Atmospheric visibility parameters in the data. This indicates the detection wavelength of the multimodal sensing device. This indicates the preset reference center wavelength. This represents the dynamic coefficient of the scattering particle size distribution related to visibility conditions. Represents environmental context information The precipitation or snowfall rate parameter in the data. This represents the first precipitation scattering empirical coefficient related to the operating frequency band of the multimodal sensing device. This represents the second precipitation scattering empirical coefficient related to the operating frequency band of the multimodal sensing device.

[0063] In this embodiment, the physical constraint model not only plays a compensatory role in forward reasoning, but also, during the training phase of the perceptual network, the analytical solution of its physical decay law is explicitly embedded into the total loss function, forming a joint optimization architecture driven by both data and physical factors.

[0064] S400: Perform cross-modal target perception and identification on the final fused features, and output the target recognition result.

[0065] S410. The final fused features are restored to tensor input format through feature reshaping and dimension transformation operations to obtain the transformed feature tensor.

[0066] Specifically, the high-dimensional final fused feature sequence output from the previous stage is received, and through feature reshaping and dimensional transformation operations, it is restored from a one-dimensional sequence or unstructured representation to a two-dimensional or three-dimensional mesh feature map with spatial topology to adapt to the tensor input format of the positioning head detection module.

[0067] The aforementioned feature reshaping and dimension transformation operations specifically refer to the final fusion of 3D voxelized or serialized features. Compression along the spatial height dimension maps to a regular two-dimensional bird's-eye view (BEV) feature tensor. The calculation formula is as follows:

[0068] in, For vertical physical height range, These are mapping parameters.

[0069] The above formula utilizes parameters with mapping. The convolutional kernel combined with max-pooling operation, in the vertical physical height range Feature compression is completed internally, effectively preserving strong response physical features while reducing the computational dimension of downstream detection heads.

[0070] In this embodiment, spatial transformation and feature pooling operations are used to map complex cross-modal fusion features into regular bird's-eye view features or dense two-dimensional feature maps. This step effectively bridges the data flow gap between the high-dimensional physical feature space and the standard convolutional / attention detection head, ensuring the lossless transmission of perceptual information.

[0071] S420: The transformed feature tensors are stacked through a multi-layer network structure, and the stacked features are input into the localization head detection module. The three parallel decoding branches output a classification score map representing the probability of the target's existence, a bounding box size map for predicting the target's three-dimensional size, and a local position offset map for eliminating downsampling discretization errors.

[0072] Specifically, a deep stacking operation is performed on the reshaped feature tensor using a multi-layer convolutional neural network, batch normalization, and activation functions. The three parallel decoding branches of the localization head detection module are a classification branch, a regression branch, and an offset branch. The stacked features are input into the localization head detection module, and the three parallel decoding branches output a classification score map representing the probability of target presence, a bounding box size map for predicting the target's 3D dimensions, and a local position offset map for eliminating downsampling discretization errors, respectively.

[0073] The classification branch maps the transformed feature tensor through a convolutional network and outputs a classification score map representing the probability of target presence after passing through a sigmoid activation function; the regression branch maps the transformed feature tensor through a convolutional network and outputs a bounding box size map used to predict the 3D size of the target; the offset branch maps the transformed feature tensor through a convolutional network and outputs a local position offset map used to eliminate downsampling discretization errors. The expressions for the three parallel decoding branches are:

[0074]

[0075]

[0076] in, The classification scores for the corresponding K target categories are represented by a classification score map. The bounding box size map represents the regression prediction features that include the target's length, width, height, and three-dimensional heading angle. The offset is corrected for the horizontal coordinates of the predicted grid center point, representing a local position offset map; Each branch represents a separate set of parameters for the convolutional network. For the transformed feature tensor, For activation function, The number of feature map channels. It is the set of real numbers.

[0077] In this embodiment, the localization head detection module adopts a lightweight and efficient parallel prediction architecture. The classification branch focuses on the semantic recognition of the target, including distinguishing between obstacle avoidance cones and ordinary road debris, while the regression branch focuses on high-precision geometric localization. Thanks to the deep correction and enhancement of the input features by the physical constraint model, the discriminative difference between the target center point and the background region on the feature map is significantly amplified, greatly reducing the difficulty of network prediction.

[0078] S430. Select the position with the highest score as the target center based on the classification score map, and combine it with the bounding box size map and the local position offset map. Then, restore the bounding box pose and category attributes of the target in the real physical environment through the spatial geometric transformation equation to obtain the target recognition result.

[0079] Specifically, in the output classification score map, the highest local spatial coordinates where the response value exceeds a preset threshold are selected and used as the geometric center point of the obstacle avoidance target; at the same time, the predicted values ​​of the corresponding coordinate positions in the bounding box size map and the local position offset map are extracted, and the bounding box pose and category attributes of the target in the real physical world are restored through the spatial geometric transformation equation to complete the final target recognition output.

[0080] In this embodiment, the physically corrected high-confidence features are directly mapped to the bounding box pose and category attributes required for the autonomous driving system's decision-making. Since the input features have eliminated common-mode noise such as rain and fog, and actively compensated for the attenuation energy of infrared or radar signals, this positioning head detection module can still accurately output high-precision detection boxes for tiny obstacle avoidance cones with an extremely low false positive rate, even in complex traffic conditions with extremely low visibility and drastic changes in lighting. This provides a solid perception guarantee for the safe obstacle avoidance of intelligent driving vehicles.

[0081] In summary, this invention first acquires multimodal sensor data characterizing the spatial and frequency domain attributes of the target, and maps independently extracted heterogeneous features to a unified semantic space; it then constructs a spatial constraint mask using the extracted spatial structural features; guided by the spatial constraint mask, a differential fusion model performs multipath differential operations to actively filter out cross-modal non-consistent common-mode noise; a physical constraint model based on signal propagation mechanisms is used to perform feature inversion and consistency correction on the preliminary differential results distorted by environmental interference; finally, the positioning head detection module outputs accurate cross-modal target recognition results based on the final fused features. This invention's method significantly improves the perception robustness and anti-interference capability of intelligent driving systems in extremely complex traffic scenarios such as rain, fog, and low signal-to-noise ratios through spatially guided differential noise reduction and a dual-drive physical correction mechanism, making it suitable for high-reliability autonomous driving obstacle avoidance perception tasks in complex dynamic environments.

[0082] Please see Figure 3 , Figure 3 A structural block diagram of a physical constraint-driven cross-modal target sensing system provided in an embodiment of the present invention. The cross-modal target sensing system includes: The data acquisition and mapping module is used to acquire multimodal sensor data that characterizes the spatial and frequency domain attributes of the target, extract features according to the mode and map them to a unified semantic space to obtain spatial structure features and frequency domain signal features. The differential fusion module is used to construct a spatial constraint mask based on the geometry and position of spatial structural features, and to perform differential operations on the multipath mapping response of spatial structural features and frequency domain signal features under the guidance of the spatial constraint mask to obtain preliminary differential results. The physical constraint module is used to perform inversion and consistency correction on the preliminary difference results based on the physical constraint model of signal propagation mechanism, and obtain the final fusion features. The positioning head detection module is used to perceive and identify cross-modal targets based on the final fused features, and output the target recognition results.

[0083] Furthermore, the overall joint optimization objective function of this cross-modal target perception system is composed of a weighted sum of a data-driven loss term and a physical constraint loss term. The data-driven loss term is composed of a weighted combination of the focus loss function for target classification and the geometric regression loss function for the 3D dimensions and positional offset of the bounding box. The physical constraint loss term is obtained by calculating the squared L2 norm difference between the final fused features output by the network forward propagation and the preset theoretical analytical benchmark functions for wave attenuation and optical scattering, and then weighting them through dynamic penalty weights. The expression is as follows:

[0084]

[0085] in, For the overall joint optimization objective function, For data-driven loss terms, This represents the dynamic penalty weight that controls the strength of physical constraints. This represents the total number of samples in a single batch during network training. Indicates the first in a single batch A sample index, Indicates the first The final fused features of each sample are output from the network's forward propagation. This represents the pre-defined theoretical analytical reference function for wave attenuation and optical scattering. Indicates the first Each sample input contains abnormal feature components that are distorted due to environmental interference. Indicates the first Environmental context information for each sample Indicates the first The relative physical distance of each sample target Indicates the first Retroreflection coefficient of a sample target surface, The squared L2 norm represents the ratio between eigenvectors. The focus loss function represents the target classification. The geometric regression loss function represents the 3D dimensions and positional offset of the bounding box. The weights for the classification loss, The weights for the regression loss.

[0086] This embodiment forces the neural network to strictly adhere to the laws of electromagnetic scattering and optical propagation during backpropagation optimization by calculating the residual between the reconstructed features of the network and the theoretical solutions of physical laws. This dual-drive mechanism of data and physics greatly improves the model's generalization ability when facing non-standard operating conditions and extreme weather.

[0087] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such modifications and substitutions should be considered within the scope of protection of the present invention.

Claims

1. A cross-modal target perception method based on physical constraints, characterized in that, Including the following steps: Multimodal sensor data representing the spatial and frequency domain attributes of the target are acquired, and features are extracted according to the modality and mapped to a unified semantic space to obtain spatial structure features and frequency domain signal features. Based on the geometry and position of the spatial structural features, a spatial constraint mask is constructed, and under the guidance of the spatial constraint mask, a differential operation is performed on the multipath mapping response of the spatial structural features and the frequency domain signal features to obtain preliminary differential results. The physical constraint model based on the signal propagation mechanism performs inversion and consistency correction on the preliminary difference results to obtain the final fusion features, including: obtaining the current environmental context information and target surface reflector attributes, and calling a preset physical constraint model containing signal attenuation characteristics. The local feature energy in the preliminary difference results is analyzed and the feature distortion index is calculated. The abnormal frequency response region caused by complex environmental interference is located based on the feature distortion index. The local feature energy is represented by the sum of squares of the L2 norm of the preliminary difference results of each coordinate point within the local sensing window centered on the spatial coordinates. The feature distortion index is calculated based on the ratio of the local feature energy to the lossless distribution of the theoretical target frequency domain feature energy based on spatial prior estimation. When the feature distortion index is greater than a preset hard threshold for distortion judgment, the region corresponding to the local sensing window centered on the spatial coordinates is the abnormal frequency response region. Combining the preliminary difference results, the current environmental context information, and the target surface reflector properties, the physical constraint model is used to perform feature inversion and compensation processing on the abnormal frequency response region to generate the final fused features. The final fused features are obtained by multiplying the preliminary difference results by an exponential compensation term, which is calculated based on the environmental context information, the target surface retroreflection coefficient, the radial physical distance, and the detection wavelength of the multimodal sensing device. The final fused features are used to perceive and identify targets across modalities, and the target recognition results are output.

2. The cross-modal target perception method based on physical constraints as described in claim 1, characterized in that, Multimodal sensor data representing the spatial and frequency domain attributes of the target are acquired, and features are extracted separately for each mode and mapped to a unified semantic space to obtain spatial structure features and frequency domain signal features, including: Acquire spatial structural data characterizing the target's geometry and position, and frequency domain signal data characterizing the target's physical properties; The spatial structure data is feature-encoded using a spatial feature extraction module to obtain spatial extracted features. The frequency domain signal data is feature-encoded using a frequency domain feature extraction module to obtain frequency domain extracted features; By performing feature reduction and projection operations, the spatial extracted features and the frequency domain extracted features are aligned to a feature space with the same channel dimension to obtain the spatial structure features and the frequency domain signal features.

3. The cross-modal target perception method based on physical constraints according to claim 1, characterized in that, Based on the geometry and position of the spatial structural features, a spatial constraint mask is constructed. Guided by the spatial constraint mask, a difference operation is performed on the multipath mapping response of the spatial structural features and the frequency domain signal features to obtain preliminary difference results, including: Based on the spatial structure features, the three-dimensional spatial occupancy state and geometric shape distribution prior of the target are extracted, and a spatial constraint mask is constructed to indicate the potential existence area of ​​the target. The spatial structure features and the frequency domain signal features are subjected to nonlinear feature transformation using the first feature mapping function to generate a hybrid correlation feature response that includes the target matching signal and environmental background interference; The spatial structure features and the frequency domain signal features are subjected to nonlinear feature transformation using the second feature mapping function to generate a common-mode noise feature response for capturing cross-modal common-mode noise distribution in complex environments; The spatial constraint mask is used to perform spatial guided modulation on the hybrid correlation feature response, and the spatially guided modulated hybrid correlation feature response is differentially subtracted from the common-mode noise feature response to obtain the preliminary difference result.

4. The cross-modal target perception method based on physical constraints according to claim 3, characterized in that, The spatial constraint mask is obtained by extracting geometric and topological features from spatial structural features, concatenating them with spatial structural features, performing convolutional mapping on the concatenated features through a learnable spatial convolutional network, and then outputting the result after passing an activation function. The first feature mapping function performs a nonlinear transformation by concatenating spatial structural features and frequency domain signal features through channels, and outputs the hybrid correlation feature response through an activation function; The second feature mapping function performs a nonlinear transformation after calculating the absolute difference between the spatial structure features and the frequency domain signal features, and outputs the common-mode noise feature response through an activation function; The preliminary difference result is obtained by performing element-wise multiplication of the hybrid correlation feature response output by the first feature mapping function with the enhancement factor of the spatial constraint mask, weighting the common-mode noise feature response output by the second feature mapping function with the noise cancellation coefficient obtained by dynamic learning, and subtracting the weighted feature from the feature after multiplication.

5. The cross-modal target perception method based on physical constraints according to claim 1, characterized in that, The final fused features are used for cross-modal target perception and identification, and the target recognition results are output, including: The final fused features are restored to tensor input format through feature reshaping and dimension transformation operations to obtain the transformed feature tensor. The transformed feature tensors are stacked using a multi-layer network structure, and the stacked features are input into the localization head detection module. The module outputs a classification score map representing the probability of the target's existence, a bounding box size map for predicting the target's 3D size, and a local position offset map for eliminating downsampling discretization errors through three parallel decoding branches. The location with the highest score is selected as the target center based on the classification score map. Combined with the bounding box size map and the local position offset map, the bounding box pose and category attributes of the target in the real physical environment are restored through spatial geometric transformation equations to obtain the target recognition result.

6. The cross-modal target perception method based on physical constraints according to claim 5, characterized in that, The three parallel decoding branches include a classification branch, a regression branch, and an offset branch; The classification branch maps the transformed feature tensor through a convolutional network and outputs a classification score map representing the probability of the target's existence through a Sigmoid activation function. The regression branch maps the transformed feature tensor through a convolutional network and outputs a bounding box size map for predicting the three-dimensional size of the target. The offset branch maps the transformed feature tensor through a convolutional network, outputting a local position offset map to eliminate downsampling discretization errors.

7. A cross-modal target perception system based on physical constraints, characterized in that, include: The data acquisition and mapping module is used to acquire multimodal sensor data that characterizes the spatial and frequency domain attributes of the target, extract features according to the mode and map them to a unified semantic space to obtain spatial structure features and frequency domain signal features. The differential fusion module is used to construct a spatial constraint mask based on the geometry and position of the spatial structural features, and perform differential operations on the multipath mapping response of the spatial structural features and the frequency domain signal features under the guidance of the spatial constraint mask to obtain preliminary differential results. The physical constraint module is used to perform inversion and consistency correction on the preliminary difference results based on the physical constraint model of signal propagation mechanism to obtain the final fusion features; including: obtaining the current environmental context information and target surface reflector attributes, and calling the preset physical constraint model containing signal attenuation characteristics; The local feature energy in the preliminary difference results is analyzed and the feature distortion index is calculated. The abnormal frequency response region caused by complex environmental interference is located based on the feature distortion index. The local feature energy is represented by the sum of squares of the L2 norm of the preliminary difference results of each coordinate point within the local sensing window centered on the spatial coordinates. The feature distortion index is calculated based on the ratio of the local feature energy to the lossless distribution of the theoretical target frequency domain feature energy based on spatial prior estimation. When the feature distortion index is greater than a preset hard threshold for distortion judgment, the region corresponding to the local sensing window centered on the spatial coordinates is the abnormal frequency response region. Combining the preliminary difference results, the current environmental context information, and the target surface reflector properties, the physical constraint model is used to perform feature inversion and compensation processing on the abnormal frequency response region to generate the final fused features. The final fused features are obtained by multiplying the preliminary difference results by an exponential compensation term, which is calculated based on the environmental context information, the target surface retroreflection coefficient, the radial physical distance, and the detection wavelength of the multimodal sensing device. The positioning head detection module is used to perceive and identify cross-modal targets based on the final fused features, and output the target recognition result.

8. The cross-modal target perception system based on physical constraints according to claim 7, characterized in that, The overall joint optimization objective function of the cross-modal target perception system is composed of a weighted sum of a data-driven loss term and a physical constraint loss term; The data-driven loss term is composed of a weighted combination of the focus loss function for target classification and the geometric regression loss function for the three-dimensional size and position offset of the bounding box; The physical constraint loss term is obtained by calculating the difference between the squared L2 norm of the final fusion feature output by the network forward propagation and the preset analytical benchmark function of wave attenuation and optical scattering theory, and then weighted by dynamic penalty weight control.

Citation Information

Patent Citations

  • Noise suppression and frequency response balance dynamic optimization method for sound system in complex sound field environment

    CN122340424A