AEB enhancement method based on fusion of wifi imaging and multi-modal perception

CN122347791BActive Publication Date: 2026-08-11WUHAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-09
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

然而,现有WiFi成像研究多集中于室内静态环境,尚未有效应用于高速移动的车辆场景,且缺乏与视觉、雷达数据的深度融合机制,无法直接服务于AEB等主动安全功能

Benefits of technology

1、针对现有视觉/雷达传感器在遭遇遮挡时完全失效,而本发明利用WiFi信号的穿透特性,在视觉失效的场景下仍能探测到数十米范围内的微弱运动扰动。虽然WiFi感知的距离分辨率有限,无法精确描绘目标轮廓,但足以在“鬼探头”目标进入视觉视野前0.5-1.5秒提供宝贵的预警线索,为AEB系统争取关键反应时间。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122347791B_ABST
    Figure CN122347791B_ABST
Patent Text Reader

Abstract

This invention proposes an AEB (Autonomous Emergency Braking) enhancement method based on WiFi imaging and multimodal perception fusion, belonging to the field of intelligent transportation and autonomous driving technology. The method includes the following steps: acquiring channel state information, images, and target point cloud data within the coverage area, and synchronizing them with timestamps to ensure precise temporal alignment of data from different sources; acquiring a first perception result from the image; acquiring a second perception result from the target point cloud data acquired by millimeter-wave radar; acquiring a third perception result based on the channel state information; projecting the three perception results onto the same coordinate system and performing weighted fusion using a cross-modal adaptive gating fusion mechanism to obtain a fused feature map; performing Kalman filtering tracking on the fused feature map to maintain the target's tracking state and estimate its motion state; predicting the target's potential location area over a future period, evaluating the dynamic collision risk assessment index, and executing a graded AEB control strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of intelligent transportation and autonomous driving technology, and in particular to an AEB enhancement method based on WiFi imaging and multimodal perception fusion. Background Technology

[0002] With the rapid development of intelligent connected vehicles and autonomous driving technologies, environmental perception systems, as the foundation of vehicle decision-making and control, directly impact driving safety. Current mainstream autonomous driving perception solutions primarily rely on sensors such as cameras, millimeter-wave radar, and lidar. Cameras can provide rich semantic information, enabling target classification and lane line recognition, but their performance degrades in adverse weather conditions such as strong light, rain, and fog, and they cannot detect targets in obscured areas. Millimeter-wave radar offers advantages such as high ranging and speed measurement accuracy and strong all-weather adaptability, but its angular resolution is low, limiting its ability to detect stationary or laterally moving targets. LiDAR can generate high-precision 3D point clouds, achieving accurate target localization and contour drawing, but it is costly, and its performance significantly degrades in rain and snow.

[0003] All of the aforementioned sensors share a common physical limitation—"line-of-sight propagation" and "optical obstruction." When large vehicles, buildings, or road infrastructure obstruct the line of sight, vision, radar, and lidar cannot detect moving targets behind the obstruction, such as pedestrians or non-motorized vehicles suddenly appearing, i.e., the "ghost peek" scenario. Existing Automatic Emergency Braking (AEB) systems often fail in such scenarios due to perception lag or target loss, leading to collisions. In recent years, sensing technology based on WiFi Channel State Information (CSI) has made significant progress. Research shows that WiFi signals can penetrate non-metallic obstacles, and by analyzing amplitude and phase changes in CSI, human movements, postures, and trajectories can be reconstructed and imaged. However, existing WiFi imaging research is mostly focused on indoor static environments and has not yet been effectively applied to high-speed moving vehicle scenarios. Furthermore, it lacks a deep fusion mechanism with vision and radar data, and cannot directly serve active safety functions such as AEB. WiFi sensing technology differs fundamentally from traditional vehicle sensors, such as millimeter-wave radar, lidar, and cameras, in terms of physical characteristics. The latter pursues high distance resolution and accurate contour depiction, but is limited by line-of-sight propagation; the former, while having a theoretical limit in distance resolution related to signal bandwidth, possesses natural non-line-of-sight penetration capability and a wider coverage area. This difference in technical characteristics provides a completely new solution for safety early warning in obstructed scenarios such as "ghost peeking"—no need to "see clearly," only "perceive."

[0004] Therefore, this paper proposes an AEB enhancement method based on the fusion of WiFi imaging and multimodal perception, which effectively combines high-precision sensors with WiFi perception. This method can overcome the perception blind zone and still detect the target contour within a certain distance in scenarios where vision fails, thus gaining valuable reaction time for AEB and has important application significance. Summary of the Invention

[0005] In view of this, the present invention proposes an AEB enhancement method based on WiFi imaging and multimodal perception fusion, which uses a mobile WiFi imaging sensing unit and combines multi-source data of visual images and radar point clouds to extract CSI features, identify weak signal disturbances behind obstacles, and invert the motion trajectory of moving targets.

[0006] On one hand, the present invention provides an AEB enhancement method based on WiFi imaging and multimodal sensing fusion, comprising the following steps: S1: Configure the WiFi imaging sensing module to continuously collect Channel State Information (CSI) within the coverage area; configure the visual acquisition module to acquire RGB images; configure the millimeter-wave radar module to collect target point cloud data. S2: Perform multipath-aware adaptive filtering on the Channel State Information (CSI) to obtain filtered CSI data C. clean ; Obtain the theoretical background channel H at the current moment bg The filtered CSI data C clean With theoretical background channel H bg Subtraction yields the residual perturbation signal ΔC(t). This residual perturbation signal ΔC(t) is then organized into a three-dimensional tensor. A spatial-temporal flow spatiotemporal feature extraction network is constructed to extract spatial and temporal feature maps from the three-dimensional tensor. These spatial and temporal feature maps are then fused using a cross-attention mechanism to generate a deep feature vector F. fusion ; S3: Based on the RGB image, target point cloud data, and the depth feature vector F obtained in step S2 fusion Three different perception results are obtained in a one-to-one correspondence; a unified bird's-eye view fusion space is constructed, the three different perception results are projected onto the same coordinate system, and a cross-modal adaptive gating fusion mechanism is used for weighted fusion to obtain the fused BEV feature map; S4: Perform Kalman filtering on the fused BEV feature map to track the target, maintain the target's tracking state and estimate its motion state; combine the motion trend information of the occluded target with its historical trajectory to predict the potential location area of ​​the target in the future and assess the dynamic collision risk assessment index. S5: When the WiFi imaging sensing module first detects a moving target behind an obstruction, but the visual acquisition module or millimeter-wave radar module has not yet confirmed it, the early warning mechanism is activated, and a graded AEB control strategy is executed based on the dynamic collision risk assessment index.

[0007] Based on the above technical solution, preferably, step S1 involves the WiFi imaging sensing module continuously acquiring Channel State Information (CSI) within the coverage area at a sampling rate of 50Hz. The CSI includes the amplitude and phase information of 32 subcarriers, and is stored in a vector C of length 32. raw In the vector, each element represents the amplitude and phase information of the corresponding subcarrier; the visual acquisition module acquires RGB images with a resolution of 1920×1080; the millimeter-wave radar module acquires target point cloud data; and the three types of data, namely channel status information (CSI), RGB images, and target point cloud data, are synchronized through GPS / INS high-precision timestamps to ensure that the multimodal data are accurately aligned in time.

[0008] Preferably, in step S2, the channel state information (CSI) is subjected to multipath-aware adaptive filtering to obtain filtered CSI data C. clean Specifically, it includes: Using vehicle speed V and signal angle of arrival θ, the Doppler frequency shift is calculated and phase compensation is performed on the channel state information (CSI); the phase-compensated CSI is C comp (t), for C comp (t) First, perform an inverse Fourier transform along the frequency dimension to obtain the time delay response, then perform a Fourier transform along the time dimension to obtain the Doppler response, and construct a two-dimensional response matrix. R ( τ , ν ), τ For time delay, ν Set an adaptive threshold T for the Doppler frequency. multipath For two-dimensional response matrix R ( τ , ν Peak detection is performed, and energy ≥ adaptive threshold T is retained. multipath The path is used as the effective signal path, and the energy < adaptive threshold T multipath The path energy is set to zero, resulting in a clean spectral matrix R. clean For the clean spectral matrix R clean Perform an inverse two-dimensional Fourier transform to recover the time-domain CSI sequence C. filtered ; For time-domain CSI sequences C filtered Set a sliding window of length L. Within each window, at the window center point p... iCalculate the Local Outlier Factor (LOF) k (p i Set the impulse noise threshold B, and when the local outlier factor LRD k When (pi) > B, the center point pi of the window is determined to be an impulse noise point. A cubic spline interpolation method is used, employing the values ​​of the preceding and following normal points to remove and repair the impulse noise point. After traversing all windows, the final filtered CSI data C is obtained. clean .

[0009] Preferably, in step S2, the residual perturbation signal ΔC(t) is organized into a three-dimensional tensor, and a spatial-temporal flow spatiotemporal feature extraction network is constructed to extract spatial feature maps and temporal feature maps from the three-dimensional tensor, respectively. The spatial feature maps and temporal feature maps are then fused through a cross-attention mechanism to generate a depth feature vector F. fusion Specifically, it includes: Acquire the residual perturbation signal ΔC(t) of 50-100 consecutive frames and generate the corresponding three-dimensional tensor T; Let the spatial flow network structure include deformable convolutional layers, batch normalization layers, max pooling layers, and ordinary convolutional layers arranged in sequence, and let the spatial slice T of the three-dimensional tensor T be processed. t The input consists of a deformable convolutional layer with a 3×3 kernel, a stride of 1, and 64 output channels, used to focus on key spatial regions where signal energy is concentrated; a batch normalization layer is used to force normalization of the input distribution, and the output of the batch normalization layer introduces nonlinearity through the ReLU activation function; a max pooling layer with a 2×2 kernel and a stride of 2; and a regular convolutional layer with a 3×3 kernel and 128 output channels, used to output the spatial feature vector F. spatial ; Let the time-series flow consist of a first one-dimensional dilated convolutional layer, a gated linear unit (GLU), a second one-dimensional dilated convolutional layer, and a global average pooling layer, arranged sequentially. The three-dimensional tensor T is reorganized into a sequence. A time-series vector is defined for each antenna pair corresponding to a subcarrier location. The sequence-form three-dimensional tensor is input into the first one-dimensional dilated convolutional layer, which expands the receptive field. The first one-dimensional dilated convolutional layer has a kernel size of 5, a dilation rate of 2, and 64 output channels. The GLU selectively transmits time-series information through a gating mechanism. The second one-dimensional dilated convolutional layer also expands the receptive field, with a kernel size of 5, a dilation rate of 4, and 128 output channels. The global average pooling layer aggregates the time-series features of all location points, outputting a 128-dimensional time-series feature vector F. temporal ; The spatial feature vector F spatial and time series eigenvectors F temporal The corresponding query Q is obtained through linear transformation. s With Qt Key K s and K t Value V s and V t Then, the attention weight α of spatial features on temporal features is calculated. s→t Attention weight α of temporal features to spatial features t→s Finally, the two attention weights α s→t and α t→s The data is concatenated and fused through a fully connected (FC) layer to obtain a 256-dimensional deep feature vector F. fusion .

[0010] Preferably, the method described in step S3 is based on the RGB image, target point cloud data, and the depth feature vector F obtained in step S2. fusion This yields three different perceptual results in a one-to-one correspondence, specifically including: First, a multi-scale semantic-geometric joint detection head is constructed, using a deep residual network with ResNet-50 architecture as the backbone network to extract multi-scale features from the image. The feature maps output from layers 3, 4, and 5 of the backbone network are taken and denoted as C3, C4, and C5, respectively. Their sizes are 1 / 8, 1 / 16, and 1 / 32 of the original RGB image, respectively. Through the top-down path and lateral connections of the Feature Pyramid Network (FPN), feature maps P3, P4, P5, P6, and P7 at five scales are constructed, corresponding to targets of different sizes. Two parallel semantic and geometric branches are set at each pyramid level. The semantic branch is used for target classification, including a 3×3 convolutional layer, global average pooling, and a fully connected layer, to output the target's class probability vector. The geometric branch is used for bounding box regression, including a deformable convolutional layer and a fully connected layer, to output the bounding box coordinates [x, y, w, h] and the confidence score, where confidence ∈ [0, 1]. For bounding boxes with an overlap IoU greater than the threshold of 0.5, the occlusion score of each bounding box is calculated. The confidence score output by the geometric branch is weighted and fused with the occlusion score as the retention priority, and the bounding box with the highest priority is retained. The category, bounding box and confidence score of the visible target are output as the first perception result. An improved DBSCAN clustering method is used based on joint features of distance, velocity, and angle to construct the clustering of each point p in the radar point cloud data. i eigenvectors (d) i , v i θ i ), d i The distance to point i is directly measured by the millimeter-wave radar module; v iLet θ be the radial velocity of point i, obtained from Doppler information from the millimeter-wave radar module; i Let d be the azimuth angle of point i, estimated by the antenna array of the millimeter-wave radar module; let d j、 v j and θ j Let v be the distance, radial velocity, and azimuth of point j, respectively. Define a distance metric and perform clustering. Calculate the average velocity for the clustered point cloud clusters. If the average velocity |v| avg If the speed is less than the static speed, it is determined to be a static obstacle; otherwise, it is a dynamic target. A filter is established to track the dynamic target's state x=[d, v, θ, a], where a is the dynamic target's acceleration, and d, v, and θ are the dynamic target's distance, radial velocity, and azimuth angle, respectively. The nearest neighbor data association method is used to match the current frame detection result with the existing trajectory and update the target state as the second perception result. Construct an occluded target detection network, including an input layer, three sequentially arranged fully connected layers, and a multi-task output layer. The input layer receives the depth feature vector F obtained in step S2. fusion The three sequentially arranged fully connected layers are used to further extract discriminative features related to occluded target detection. The multi-task output layer has three parallel branches. The first branch performs motion presence detection, determining whether a moving target exists within the occluded area. It uses the Sigmoid activation function and outputs a probability value P. exist ∈[0,1], when P exist When the value is greater than 0.5, a moving target is identified; the second branch performs a coarse azimuth estimation, estimating the approximate azimuth angle θ of the occluded target. est For the first branch, ∈ [-90°, 90°], a linear activation function is used to output the estimated azimuth angle. The third branch classifies the motion trend to determine the motion trend of the occluded target: moving away, stationary, or approaching. A softmax activation function is used to output the probability distribution of the three categories. The presence of motion disturbance in the occluded area, the rough azimuth of the occluded target, and the motion trend are used as the third perception result.

[0011] Preferably, the network training process of the occluded target detection network adopts an end-to-end training method, and the loss function of the network training of the occluded target detection network is set to be the sum of three branches: binary cross-entropy loss, mean squared error loss function, and multi-component cross-entropy loss.

[0012] Preferably, step S4 involves performing Kalman filtering tracking on the fused target, maintaining the target ID, and estimating the motion state. The parameters of the Kalman filter are set as follows: state transition matrix A, observation matrix H, process noise covariance matrix Q, and observation noise covariance matrix R. The motion trend information of the occluded target and its historical trajectory are input into the time-series prediction model. This model first extracts local temporal features from the long-term sequence through dilated convolution, then applies multi-head self-attention to capture global long-range dependencies, and finally outputs the Gaussian distribution parameters of the predicted positions at each time point within the next 2-3 seconds—namely, the mean μ and the covariance matrix Σ—to quantify the prediction uncertainty. The dynamic collision risk assessment index R is calculated. risk R risk The calculation formula is: R risk =α0×P collision +β0×V rel +γ0×T time-to-collision , where P collision V represents the collision probability. rel T represents relative velocity. time-to-collision The collision time is represented by α0, β0, and γ0, which are weight terms.

[0013] Preferably, the values ​​of α0, β0, and γ0 are 0.5, 0.3, and 0.2, respectively.

[0014] Preferably, the step S5 of implementing a graded AEB control strategy based on the dynamic collision risk assessment index includes the following: When R risk When the value is less than 0.3, the system determines it to be low risk and issues a warning, alerting the driver to the obstructed area ahead via the vehicle's display screen and sound. When 0.3≤R risk When the speed is less than 0.6, the system determines it to be of medium risk and applies partial braking to reduce the vehicle speed to a safe level to reduce the risk of collision. When R risk When the value is ≥0.6, the system determines it as a high risk and performs emergency braking to quickly decelerate the vehicle to a stop in order to avoid a collision.

[0015] On the other hand, the present invention also provides an AEB enhancement system based on WiFi imaging and multimodal sensing fusion, for implementing the above-mentioned method, comprising: The data acquisition unit includes a WiFi imaging sensing module, a visual acquisition module, and a millimeter-wave radar module, which are used to continuously acquire channel state information (CSI), RGB images, and target point cloud data within the coverage area, respectively. The spatiotemporal synchronization submodule communicates with the data acquisition unit and is used to timestamp synchronize the channel state information (CSI), RGB image, and target point cloud data, so that the data acquired from different sources are precisely aligned in time. The data preprocessing submodule, communicatively connected to the spatiotemporal synchronization submodule, is used to perform multipath-aware adaptive filtering on the Channel State Information (CSI) to obtain filtered CSI data. clean ; Obtain the theoretical background channel H at the current moment bg The filtered CSI data C clean With theoretical background channel H bg Subtraction yields the residual perturbation signal ΔC(t). This residual perturbation signal ΔC(t) is organized into a three-dimensional tensor of "antenna pair-subcarrier-time". A spatial-temporal flow spatiotemporal feature extraction network is constructed to extract spatial and temporal feature maps from the three-dimensional tensor. These spatial and temporal feature maps are then fused using a cross-attention mechanism to generate a deep feature vector F. fusion The RGB image is input into a multi-scale semantic-geometric joint detection head, which outputs the first perception result. The target point cloud data acquired by the millimeter-wave radar is input into a dynamic target-static obstacle joint clustering and association algorithm to obtain the second perception result. The depth feature vector F is then obtained. fusion It inputs the occluded target detection network and outputs the third perception result; The multimodal fusion submodule is used to construct a unified bird's-eye view fusion space, project the three perception results onto the same coordinate system, and use a cross-modal adaptive gating fusion mechanism to perform weighted fusion to obtain the fused BEV feature map; The decision control submodule is used to perform Kalman filtering tracking on the fused BEV feature map, maintain the target's tracking state and estimate its motion state; combine the motion trend information of the occluded target with its historical trajectory to predict the potential location area of ​​the target in the future and evaluate the dynamic collision risk assessment index; and execute a graded AEB control strategy based on the dynamic collision risk assessment index. The execution unit is used to connect with the decision control submodule, the vehicle AEB system, and the ESC vehicle stability system to execute braking or warning commands.

[0016] The AEB enhancement method based on WiFi imaging and multimodal sensing fusion provided by this invention has the following advantages compared with the prior art: 1. Addressing the issue that existing visual / radar sensors become completely ineffective when obstructed, this invention utilizes the penetrating properties of WiFi signals to detect subtle motion disturbances within a range of tens of meters, even in scenarios where visual perception fails. Although WiFi sensing has limited distance resolution and cannot accurately depict target outlines, it is sufficient to provide valuable early warning clues 0.5-1.5 seconds before a "ghostly" target enters the visual field, buying crucial reaction time for the AEB system.

[0017] 2. The WiFi imaging sensing module acts as a "sentinel," continuously monitoring the obscured area. Upon detecting motion disturbance, it immediately triggers system alert. The visual acquisition module and millimeter-wave radar module act as "snipers," accurately identifying, locating, and tracking targets once they emerge from behind obstructions and enter the line-of-sight range. Through a dynamic attention mechanism, the system enters "alert mode" when data from the WiFi imaging sensing module provides a warning but visual confirmation is not yet available. Once target confirmation is achieved based on data from the visual acquisition module and millimeter-wave radar module, the system switches to "precision braking." This hierarchical mechanism ensures security while preventing false triggers due to insufficient WiFi resolution.

[0018] 3. Unlike existing research that focuses solely on WiFi "high-resolution imaging", this invention explicitly positions WiFi sensing as a "penetrating motion disturbance detector", giving full play to its physical characteristics of wide signal coverage (theoretically reaching tens to hundreds of meters) and strong penetration ability, forming a perfect complement to millimeter-wave radar / lidar / camera in terms of technical characteristics, rather than a replacement relationship.

[0019] 4. The WiFi sensing mechanism adopted in this solution conforms to the single-site / dual-site sensing architecture defined in the newly released IEEE 802.11bf standard. The variable bandwidth of 20MHz to 160MHz defined in the standard corresponds precisely to the trade-off between "long-range coarse-grained detection" and "short-range fine-grained imaging." The preferred bandwidth mode is 20MHz / 40MHz, which ensures signal coverage while using the amplitude and phase changes of CSI to detect motion disturbances, achieving good compatibility with the standard technical approach. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1This is a structural block diagram of the AEB enhancement method based on WiFi imaging and multimodal sensing fusion of the present invention; Figure 2 This is a schematic diagram illustrating the multi-source data acquisition and preprocessing process of the AEB enhancement method based on WiFi imaging and multimodal sensing fusion according to the present invention. Figure 3 This is a schematic diagram of the mobile channel compensation and feature extraction process of the AEB enhancement method based on WiFi imaging and multimodal sensing fusion of the present invention; Figure 4 This is a schematic diagram of the multimodal perception fusion and target detection process of the AEB enhancement method based on WiFi imaging and multimodal perception fusion according to the present invention; Figure 5 This is a schematic diagram illustrating the trajectory prediction and risk assessment process of the AEB enhancement method based on WiFi imaging and multimodal sensing fusion according to the present invention. Figure 6 This is a schematic diagram of the multi-scale semantic-geometric joint detection head structure of the AEB enhancement method based on WiFi imaging and multimodal perception fusion of the present invention; Figure 7 This is a schematic diagram of the occluded target detection network structure of the AEB enhancement method based on WiFi imaging and multimodal perception fusion of the present invention. Detailed Implementation

[0022] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0023] Existing WiFi imaging research is mostly focused on indoor static environments and has not been effectively applied to high-speed moving vehicle scenarios. Furthermore, it lacks a deep fusion mechanism with visual and radar data, and cannot directly serve active safety functions such as AEB.

[0024] In view of this, such as Figure 1 As shown, on one hand, the present invention provides an AEB enhancement method based on WiFi imaging and multimodal sensing fusion, comprising the following steps: S1: Configure a WiFi imaging sensing module to continuously collect Channel State Information (CSI) within the coverage area; configure a visual acquisition module to acquire RGB images; configure a millimeter-wave radar module to collect target point cloud data; and perform time stamp synchronization on the Channel State Information (CSI), RGB images, and target point cloud data to ensure that the data from different sources are precisely aligned in time.

[0025] like Figure 2 As shown, the WiFi imaging sensing module includes a multi-antenna array with separate transmit and receive antennas (no fewer than four), a CSI extraction chip, and a preprocessing unit. The WiFi imaging sensing module continuously acquires Channel State Information (CSI) within the coverage area at a sampling rate of 50Hz, including the amplitude and phase information of 32 subcarriers. The CSI is stored in a vector C of length 32. raw In the vector, each element represents the amplitude and phase information of the corresponding subcarrier; the visual acquisition module includes a forward-looking high-definition camera for acquiring road environment images. The visual acquisition module acquires RGB images with a resolution of 1920×1080 at a frame rate of 25-30fps. The RGB images are compressed using JPEG encoding with a compression quality of 75% and stored in .jpg format; the millimeter-wave radar module uses the 77GHz or 24GHz frequency band. The millimeter-wave radar module acquires target point cloud data and extracts the target's distance, velocity, and azimuth angle. The distance resolution is 0.1-0.2m, and the angle resolution is 1°-2°. The point cloud data is stored in a vector of length 1024, where each element represents the target reflection intensity at the corresponding angle and distance; the channel status information (CSI), RGB images, and target point cloud data are synchronized using GPS / INS high-precision timestamps to ensure precise temporal alignment of the multimodal data.

[0026] S2: Perform multipath-aware adaptive filtering on the Channel State Information (CSI) to obtain filtered CSI data C. clean ; Obtain the theoretical background channel H at the current moment bg The filtered CSI data C clean With theoretical background channel H bg Subtraction yields the residual perturbation signal ΔC(t). This residual perturbation signal ΔC(t) is organized into a three-dimensional tensor of "antenna pair-subcarrier-time". A spatial-temporal flow spatiotemporal feature extraction network is constructed to extract spatial and temporal feature maps from the three-dimensional tensor. These spatial and temporal feature maps are then fused using a cross-attention mechanism to generate a deep feature vector F. fusion .

[0027] like Figure 3 As shown, firstly, using the vehicle speed V and the signal angle of arrival θ, the Doppler frequency shift Δf = 2×(V / c)×f is calculated. c cosθ, c is the speed of light, f c The carrier frequency is used, and phase compensation is performed on the Channel State Information (CSI). comp (t) = C raw (t)×e -j2πΔft C raw (t) represents the original Channel State Information (CSI), C comp(t) represents the Channel State Information (CSI) after phase compensation; the vehicle speed here is obtained through the CAN bus, and the signal arrival angle θ is estimated based on the multi-antenna array.

[0028] Channel State Information (CSI) after phase compensation, i.e., C comp (t) First, perform an inverse Fourier transform along the frequency dimension to obtain the time delay response, then perform a Fourier transform along the time dimension to obtain the Doppler response, and construct a two-dimensional response matrix. R ( τ , ν ), τ For time delay, ν Calculate the two-dimensional response matrix for the Doppler frequency. R ( τ , ν Given the mean μ(R) and standard deviation σ(R), set an adaptive threshold T. multipath =μ(R)+α·σ(R), where the coefficients α∈[2.0,2.5], for a two-dimensional response matrix R ( τ , ν Peak detection is performed, and energy ≥ adaptive threshold T is retained. multipath The path is used as the effective signal path, and the energy < adaptive threshold T multipath The path energy is set to zero, resulting in a clean spectral matrix R. clean For the clean spectral matrix R clean Perform an inverse two-dimensional Fourier transform to recover the time-domain CSI sequence C. filtered For time-domain CSI sequences C filtered Set a sliding window of length L. Within each window, at the window center point p... i Calculate the Local Outlier Factor (LOF) k (p i ), where k is the number of nearest neighbors, and the LOF calculation method is as follows: first find p i Find the k nearest neighbors and calculate the reach-dist distance. k (p i ,o)=max(k-distance(o), d(p i Then calculate the locally reachable density (LRD). k (p i = 1 / [Average reachable distance]; finally calculate LOF. k (p i )=[Local reachability density (LRD) of neighboring points k (o) sum / p i [LRD] / k, where d(p i (,o) is the calculation point p iThe Euclidean distance between the neighboring point o and the neighboring point o is calculated using k-distance(o), which is the distance to the k-th neighboring point o. An impulse noise threshold B is set, ranging from 2.0 to 3.0, where the local outlier factor (LRD) is... k When (pi) > B, the center point pi of the window is determined to be an impulse noise point. A cubic spline interpolation method is used, employing the values ​​of the preceding and following normal points to remove and repair the impulse noise point. After traversing all windows, the final filtered CSI data C is obtained. clean .

[0029] Then, using the historical CSI sequence and vehicle motion state as input, a channel state transition matrix is ​​constructed through a physical constraint coding layer as prior guidance, and a two-layer LSTM network predicts the theoretical background channel H at the current moment. bg The filtered CSI data C clean With theoretical background channel H bg Subtracting the two yields the residual disturbance signal ΔC(t).

[0030] Obtain the residual perturbation signal ΔC(t) and generate the corresponding three-dimensional tensor T∈R. A×S×T0 Where A represents the number of antenna pairs ≥ 4, S is the number of subcarriers (value 32), and the superscript T0 represents the number of time frames (50-100 consecutive frames, corresponding to a historical window of 0.5-1 second). The spatial flow network structure includes sequentially arranged deformable convolutional layers, batch normalization layers, max pooling layers, and ordinary convolutional layers, which slice the spatial map T of the three-dimensional tensor T. t ∈R A×S The input consists of a deformable convolutional layer with a 3×3 kernel, a stride of 1, and 64 output channels. The position of each sampling point is adaptively adjusted based on the input features. This is achieved by learning an offset through an auxiliary convolutional layer, deforming the sampling grid to focus on key spatial regions where signal energy is concentrated. A batch normalization layer is used to force normalization of the input distribution, and its output is nonlinearized using the ReLU activation function. A max-pooling layer has a 2×2 kernel and a stride of 2. A regular convolutional layer has a 3×3 kernel and 128 output channels, used to output the spatial feature vector F. spatial ∈R 128×T0The gated temporal convolutional network (GTCN) for processing the time-series data includes a first one-dimensional dilated convolutional layer, a gated linear unit (GLU), a second one-dimensional dilated convolutional layer, and a global average pooling layer, arranged sequentially. The three-dimensional tensor T is reorganized into a sequence. A temporal vector is defined for each antenna pair corresponding to a subcarrier location. The sequence-form three-dimensional tensor is input into the first one-dimensional dilated convolutional layer, which expands the receptive field. The first one-dimensional dilated convolutional layer has a kernel size of 5, a dilation rate of 2, and 64 output channels. The GLU selectively transmits temporal information and suppresses noise through a gating mechanism. The second one-dimensional dilated convolutional layer also expands the receptive field, with a kernel size of 5, a dilation rate of 4, and 128 output channels. The global average pooling layer aggregates the temporal features of all location points, outputting a 128-dimensional temporal feature vector F. temporal ∈R 128 ; the spatial feature vector F spatial and time series eigenvectors F temporal The corresponding query Q is obtained through linear transformation. s With Q t Key K s and K t Value V s and V t Q s =W qs ·F s K s =W ks ·F s V s =W vs ·F s Q t =W qt ·F t K t =W kt ·F t V t =W vt ·F t W qs W ks W vs W qt W kt W vt ∈R 64×128 All are learnable weight matrices. The values ​​of these matrices were obtained through end-to-end training on a multimodal dataset containing various traffic scenarios. Then, the attention weight α of spatial features on temporal features was calculated. s→t =Softmax((Q s ·K t T ) / sqrt d k )·V tAttention weight α of temporal features to spatial features t→s =Softmax((Q t ·K s T ) / sqrt d k )·V s , where d k Let α be the dimension of the key vector, sqrt is the square root operation, Softmax is used to normalize the attention weights into a probability distribution, and finally, the two attention weights α are... s→t and α t→s The data is concatenated and fused through a fully connected (FC) layer to obtain a 256-dimensional deep feature vector F. fusion F fusion =FC(Concat(α s→t , α t→s Concat is a vector concatenation operation, FC is a fully connected layer, and a non-linear transformation is introduced to obtain the final deep feature vector F. fusio ∈R 256 It is used for subsequent detection of occluded targets.

[0031] S3: Input the RGB image into the multi-scale semantic-geometric joint detection head to output the first perception result; input the target point cloud data acquired by the millimeter-wave radar into the dynamic target-static obstacle joint clustering and association algorithm to obtain the second perception result; obtain the depth feature vector F obtained in step S2. fusion The occluded target detection network is input and the third perception result is output. A unified bird's-eye view fusion space is constructed, and the three perception results are projected onto the same coordinate system. A cross-modal adaptive gating fusion mechanism is used for weighted fusion to obtain the fused BEV feature map.

[0032] like Figure 4 Combination Figure 6As shown, the specific content for obtaining the three types of perception results includes: First, constructing a multi-scale semantic-geometric joint detection head, using a deep residual network with a ResNet-50 architecture as the backbone network to extract multi-scale features from the image. Feature maps output from layers 3, 4, and 5 of the backbone network are taken, denoted as C3, C4, and C5, respectively, with sizes of 1 / 8, 1 / 16, and 1 / 32 of the original RGB image. Through the top-down path and lateral connections of the Feature Pyramid Network (FPN), five scale feature maps P3, P4, P5, P6, and P7 are constructed, corresponding to targets of different sizes. At each pyramid level, two parallel semantic and geometric branches are set. The semantic branch is used for target classification, including 3×3 convolutional layers, global average pooling, and fully connected layers, outputting the target's class probability vector. This semantic branch introduces the channel attention module SENet to enhance the expression of key semantic features. The geometric branch is used for bounding box regression, including deformable convolutional layers and fully connected layers, outputting the target's bounding box coordinates [x, y, w, ...]. h] and confidence score, where confidence ∈ [0,1]. Deformable convolution can adaptively learn the geometric deformation of the target, improving the localization accuracy of tilted or partially occluded targets. The confidence score is activated by the Sigmoid function and represents the probability that the detection box contains a real target.

[0033] For bounding boxes with an IoU greater than the threshold of 0.5, calculate the occlusion score for each bounding box, where OcclusionScore = 1 - (percentage of target pixels within the box) / (total pixels within the box). The confidence score from the geometric branch output is weighted and fused with the occlusion score to determine the retention priority, where Priority = λ•Confidence + (1-λ) • (1-OcclusionScore), with the coefficient λ ranging from 0.6 to 0.8. The bounding boxes with the highest priority are retained. The category, bounding box, and confidence score of the visible target are output as the first perception result.

[0034] Traditional DBSCAN algorithms only use Euclidean distance and cluster based solely on spatial location, making it difficult to distinguish targets that are spatially close but move at different speeds, such as vehicles and pedestrians traveling side-by-side. This invention improves the distance metric of DBSCAN by introducing velocity and angle features to construct a three-dimensional joint feature space. Specifically, it improves DBSCAN clustering based on distance-velocity-angle joint features, constructing the clustering of each point p in the radar point cloud data. i eigenvectors (d) i , v i θ i ), d i The distance to point i is directly measured by the millimeter-wave radar module; vi Let θ be the radial velocity of point i, obtained from Doppler information from the millimeter-wave radar module; i Let d be the azimuth angle of point i, estimated by the antenna array of the millimeter-wave radar module; let d j、 v j and θ j Let $j$ be the distance, radial velocity, and azimuth of point $j$, respectively. The distance metric is defined as $j$. , where Δd = |d i -d j |,Δv = |v i - v j |,Δθ = |θ i - θ j |,σ d σ v σ θ The normalization coefficient is used to unify features of different dimensions to the same scale; the value of the normalization coefficient is based on: σ d The range sensitivity is determined based on the radar's range resolution, typically ranging from 0.1 to 0.2 meters. This value represents the clustering sensitivity in the range dimension; if the distance difference between two points is less than the resolution, they are considered similar in the range dimension. σ v The threshold is determined based on the typical speed variation range of urban roads, usually ranging from 1 to 5 meters per second. This value represents the clustering sensitivity in the speed dimension; if the speed difference between two points is less than this threshold, they are considered similar in the speed dimension. σ θ The value is determined based on the radar's angular resolution, typically ranging from 5° to 15°. This value represents the clustering sensitivity in the angular dimension.

[0035] Calculate the average velocity for the clustered point cloud clusters, if |v avg If the speed is less than the static velocity, it is determined to be a static obstacle; otherwise, it is a dynamic target. A filter is established to track the dynamic target's state x=[d, v, θ, a], where a is the acceleration. The nearest neighbor data association method is used to match the current frame detection result with the existing trajectory and update the target state as the second perception result.

[0036] like Figure 7 As shown, an occluded target detection network is constructed, including an input layer, three sequentially arranged fully connected layers, and a multi-task output layer. The input layer is used to receive the depth feature vector F obtained in step S2. fusion Three sequentially arranged fully connected layers are used to further extract discriminative features related to occluded target detection; First fully connected layer: 256→128, ReLU activation, Dropout rate 0.3; Second fully connected layer: 128→64, ReLU activation, Dropout rate 0.3; The third fully connected layer: 64→32, ReLU activated.

[0037] The multi-task output layer has three parallel branches. The first branch performs motion presence detection, determining whether there is a moving target within the occluded area. It uses the Sigmoid activation function and outputs a probability value P. exist ∈[0,1], when P exist When the value is greater than 0.5, a moving target is identified; the second branch performs a coarse azimuth estimation, estimating the approximate azimuth angle θ of the occluded target. est The first branch, ∈ [-90°, 90°], uses a linear activation function with 0° directly in front of the vehicle as the reference angle. It outputs an estimated azimuth angle, derived from the phase difference between different antenna pairs in the CSI signal. High precision is not prioritized; the estimated azimuth angle is only used to indicate the direction of the risk source. The third branch classifies motion trends, determining the movement trend of the obscured target: moving away, stationary, or approaching. A Softmax activation function is used to output the probability distribution for the three categories. The motion trend is derived from the Doppler frequency shift of the CSI phase's time change rate, serving as a crucial basis for AEB early warning. The presence of motion disturbances in the obscured area, the approximate azimuth of the obscured target, and its motion trend are used as the third sensing result.

[0038] The training process of the occluded target detection network adopts an end-to-end training method, with a loss function L. occluded Let L be the cross-entropy loss for binary classification. exist Mean square error loss L angle With multi-class cross-entropy loss L trend Sum of: L occluded =L exist +L angle +L trend Network output format: Existence P exist (0-1 scalar), azimuth θ est Movement trend category labels (0: away, 1: stationary, 2: near).

[0039] After detecting occluded targets, the system proceeds to the cross-modal adaptive gating fusion stage. This stage first constructs a unified bird's-eye view (BEV) fusion space, projecting the three types of perception results onto the same coordinate system. Then, it calculates the reliability index r based on the current operating state of each modality. w , r v , r r These metrics are then input into the gating network to generate fusion weights w. w , w v ,w r Finally, the features of each modality in the BEV space are weighted and summed to generate a fused BEV feature map, which serves as the input for subsequent object detection and trajectory prediction.

[0040] S4: Perform Kalman filtering on the fused BEV feature map to track the target, maintain the target's tracking state and estimate its motion state; combine the motion trend information of the occluded target with its historical trajectory to predict the potential location area of ​​the target in the future and assess the dynamic collision risk assessment index.

[0041] like Figure 5 As shown, the specific content involves performing Kalman filtering tracking on the fused target, maintaining the target ID, and estimating the motion state. The parameters of the Kalman filter are set as follows: state transition matrix A, observation matrix H, process noise covariance matrix Q, and observation noise covariance matrix R. The motion trend information of the occluded target and its historical trajectory are input into a time-series prediction model. This model first extracts local temporal features from the long-term sequence through dilated convolution, then applies multi-head self-attention to capture global long-range dependencies, and finally outputs the Gaussian distribution parameters of the predicted positions at each time point within the next 2-3 seconds—namely, the mean μ and the covariance matrix Σ—to quantify the prediction uncertainty. The dynamic collision risk assessment index R is calculated. risk R risk The calculation formula is: R risk =α0×P collision +β0×V rel +γ0×T time-to-collision , where P collision V represents the collision probability. rel T represents relative velocity. time-to-collision The collision time is represented by α0, β0, and γ0, which are weight terms.

[0042] The weights α0, β0, and γ0 are 0.5, 0.3, and 0.2, respectively.

[0043] S5: When the WiFi imaging sensing module first detects a moving target behind an obstruction, but the visual acquisition module or millimeter-wave radar module has not yet confirmed it, the early warning mechanism is activated, and a graded AEB control strategy is executed based on the dynamic collision risk assessment index.

[0044] like Figure 5 As shown, a graded AEB control strategy is implemented based on the dynamic collision risk assessment index, including the following: When R risk When the value is less than 0.3, the system determines it to be low risk and issues a warning, alerting the driver to the obstructed area ahead via the vehicle's display screen and sound. When 0.3≤R risk When the speed is less than 0.6, the system determines it to be of medium risk and applies partial braking to reduce the vehicle speed to a safe level to reduce the risk of collision. When R riskWhen the value is ≥0.6, the system determines it as a high risk and performs emergency braking to quickly decelerate the vehicle to a stop in order to avoid a collision.

[0045] On the other hand, the present invention also provides an AEB enhancement system based on WiFi imaging and multimodal sensing fusion, for implementing the above-mentioned method, comprising: The data acquisition unit includes a WiFi imaging sensing module, a visual acquisition module, and a millimeter-wave radar module, which are used to continuously acquire channel state information (CSI), RGB images, and target point cloud data within the coverage area, respectively. The spatiotemporal synchronization submodule communicates with the data acquisition unit and is used to timestamp synchronize the channel state information (CSI), RGB image, and target point cloud data, so that the data acquired from different sources are precisely aligned in time. The data preprocessing submodule, communicatively connected to the spatiotemporal synchronization submodule, is used to perform multipath-aware adaptive filtering on the Channel State Information (CSI) to obtain filtered CSI data. clean ; Obtain the theoretical background channel H at the current moment bg The filtered CSI data C clean With theoretical background channel H bg Subtraction yields the residual perturbation signal ΔC(t). This residual perturbation signal ΔC(t) is organized into a three-dimensional tensor of "antenna pair-subcarrier-time". A spatial-temporal flow spatiotemporal feature extraction network is constructed to extract spatial and temporal feature maps from the three-dimensional tensor. These spatial and temporal feature maps are then fused using a cross-attention mechanism to generate a deep feature vector F. fusion The RGB image is input into a multi-scale semantic-geometric joint detection head, which outputs the first perception result. The target point cloud data acquired by the millimeter-wave radar is input into a dynamic target-static obstacle joint clustering and association algorithm to obtain the second perception result. The depth feature vector F is then obtained. fusion It inputs the occluded target detection network and outputs the third perception result; The multimodal fusion submodule is used to construct a unified bird's-eye view fusion space, project the three perception results onto the same coordinate system, and use a cross-modal adaptive gating fusion mechanism to perform weighted fusion to obtain the fused BEV feature map; The decision control submodule is used to perform Kalman filtering tracking on the fused BEV feature map, maintain the target's tracking state and estimate its motion state; combine the motion trend information of the occluded target with its historical trajectory to predict the potential location area of ​​the target in the future and evaluate the dynamic collision risk assessment index; and execute a graded AEB control strategy based on the dynamic collision risk assessment index. The execution unit is used to connect with the decision control submodule, the vehicle AEB system, and the ESC vehicle stability system to execute braking or warning commands.

[0046] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. An AEB enhancement method based on WiFi imaging and multimodal sensing fusion, characterized in that, Includes the following steps: S1: Configure a WiFi imaging sensing module to continuously collect Channel State Information (CSI) within the coverage area; configure a visual acquisition module to acquire RGB images; configure a millimeter-wave radar module to collect target point cloud data. S2: Perform multipath-aware adaptive filtering on the Channel State Information (CSI) to obtain filtered CSI data C. clean ; Obtain the theoretical background channel H at the current moment bg The filtered CSI data C clean With theoretical background channel H bg Subtraction yields the residual perturbation signal ΔC(t). This residual perturbation signal ΔC(t) is then organized into a three-dimensional tensor. A spatial-temporal flow spatiotemporal feature extraction network is constructed to extract spatial and temporal feature maps from the three-dimensional tensor. These spatial and temporal feature maps are then fused using a cross-attention mechanism to generate a deep feature vector F. fusion ; S3: obtaining a depth feature vector F based on the RGB image, the target point cloud data, and the result of step S2 fusion , obtaining three different perception results one by one; constructing a unified bird's eye view perspective fusion space, projecting the three different perception results to the same coordinate system, adopting a cross-modal adaptive gating fusion mechanism for weighted fusion, and obtaining a fused BEV feature map; The RGB image-based, target point cloud data-based, and depth feature vector F obtained in step S2 are combined to obtain a final result in step S3 fusion corresponding to three different perception results, specifically including: First, a multi-scale semantic-geometric joint detection head is constructed, using a deep residual network with ResNet-50 architecture as the backbone network to extract multi-scale features from the image. The feature maps output from layers 3, 4, and 5 of the backbone network are taken and denoted as C3, C4, and C5, respectively. Their sizes are 1 / 8, 1 / 16, and 1 / 32 of the original RGB image, respectively. Through the top-down path and lateral connections of the Feature Pyramid Network (FPN), feature maps P3, P4, P5, P6, and P7 at five scales are constructed, corresponding to targets of different sizes. Two parallel semantic and geometric branches are set at each pyramid level. The semantic branch is used for target classification, including a 3×3 convolutional layer, global average pooling, and a fully connected layer, to output the target's class probability vector. The geometric branch is used for bounding box regression, including a deformable convolutional layer and a fully connected layer, to output the bounding box coordinates [x, y, w, h] and the confidence score, where confidence ∈ [0, 1]. For bounding boxes with an overlap IoU greater than the threshold of 0.5, the occlusion score of each bounding box is calculated. The confidence score output by the geometric branch is weighted and fused with the occlusion score as the retention priority, and the bounding box with the highest priority is retained. The category, bounding box and confidence score of the visible target are output as the first perception result. An improved DBSCAN clustering method is used based on joint features of distance, velocity, and angle to construct the clustering of each point p in the radar point cloud data. i eigenvectors (d) i , v i θ i ), d i The distance to point i is directly measured by the millimeter-wave radar module; v i Let θ be the radial velocity of point i, obtained from Doppler information from the millimeter-wave radar module; i Let d be the azimuth angle of point i, estimated by the antenna array of the millimeter-wave radar module; let d j、 v j and θ j Let the distance, radial velocity, and azimuth be the distance to point j, respectively. Define a distance metric and perform clustering. Calculate the average velocity for the clustered point cloud clusters. If the average velocity |v... avg If the speed is less than the static speed, it is determined to be a static obstacle; otherwise, it is a dynamic target. A filter is established to track the dynamic target's state x=[d, v, θ, a], where a is the dynamic target's acceleration, and d, v, and θ are the dynamic target's distance, radial velocity, and azimuth angle, respectively. The nearest neighbor data association method is used to match the current frame detection result with the existing trajectory and update the target state as the second perception result. Construct an occluded target detection network, including an input layer, three sequentially arranged fully connected layers, and a multi-task output layer. The input layer receives the depth feature vector F obtained in step S2. fusion The three sequentially arranged fully connected layers are used to further extract discriminative features related to occluded target detection. The multi-task output layer has three parallel branches. The first branch performs motion presence detection, determining whether a moving target exists within the occluded area. It uses the Sigmoid activation function and outputs a probability value P. exist ∈[0,1], when P exist When the value is greater than 0.5, a moving target is identified; the second branch performs a coarse azimuth estimation, estimating the azimuth angle θ of the obscured target. est For the first branch, ∈ [-90°, 90°], a linear activation function is used to output the estimated azimuth angle; the third branch performs motion trend classification to determine the motion trend of the occluded target: moving away, stationary, or approaching; a Softmax activation function is used to output the probability distribution of the three categories; the presence of motion disturbance in the occluded area, the rough azimuth of the occluded target, and the motion trend are used as the third perception result. S4: Perform Kalman filtering on the fused BEV feature map to track the target, maintain the target's tracking state and estimate its motion state; combine the motion trend information of the occluded target with its historical trajectory to predict the potential location area of ​​the target in the future and assess the dynamic collision risk assessment index. S5: When the WiFi imaging sensing module first detects a moving target behind an obstruction, but the visual acquisition module or millimeter-wave radar module has not yet confirmed it, the early warning mechanism is activated, and a graded AEB control strategy is executed based on the dynamic collision risk assessment index.

2. The AEB enhancement method based on WiFi imaging and multimodal sensing fusion according to claim 1, characterized in that, Step S1 involves the WiFi imaging sensing module continuously acquiring Channel State Information (CSI) within the coverage area at a sampling rate of 50Hz. The CSI includes the amplitude and phase information of 32 subcarriers and is stored in a vector C of length 32. raw In the vector, each element represents the amplitude and phase information of the corresponding subcarrier; the visual acquisition module acquires RGB images with a resolution of 1920×1080; the millimeter-wave radar module acquires target point cloud data; and the three types of data, namely channel status information (CSI), RGB images, and target point cloud data, are synchronized through GPS / INS high-precision timestamps to ensure that the multimodal data are accurately aligned in time.

3. The AEB enhancement method based on WiFi imaging and multimodal sensing fusion according to claim 2, characterized in that, Step S2 describes performing multipath-aware adaptive filtering on the Channel State Information (CSI) to obtain filtered CSI data C. clean Specifically, it includes: Using vehicle speed V and signal angle of arrival θ, the Doppler frequency shift is calculated and phase compensation is performed on the channel state information (CSI); the phase-compensated CSI is C comp (t), for C comp (t) First, perform an inverse Fourier transform along the frequency dimension to obtain the time delay response, then perform a Fourier transform along the time dimension to obtain the Doppler response, and construct a two-dimensional response matrix. R ( τ , ν ), τ For time delay, ν Set an adaptive threshold T for the Doppler frequency. multipath For two-dimensional response matrix R ( τ , ν Peak detection is performed, and energy ≥ adaptive threshold T is retained. multipath The path is used as the effective signal path, and the energy < adaptive threshold T multipath The path energy is set to zero, resulting in a clean spectral matrix R. clean For the clean spectral matrix R clean Perform an inverse two-dimensional Fourier transform to recover the time-domain CSI sequence C. filtered ; For time-domain CSI sequences C filtered Set a sliding window of length L. Within each window, at the window center point p... i Calculate the Local Outlier Factor (LOF) k (p i Set the impulse noise threshold B, and when the local outlier factor LRD k When (pi) > B, the center point pi of the window is determined to be an impulse noise point. A cubic spline interpolation method is used, employing the values ​​of the preceding and following normal points to remove and repair the impulse noise point. After traversing all windows, the final filtered CSI data C is obtained. clean .

4. The AEB enhancement method based on WiFi imaging and multimodal sensing fusion according to claim 3, characterized in that, In step S2, the residual perturbation signal ΔC(t) is organized into a three-dimensional tensor, and a spatial-temporal flow spatiotemporal feature extraction network is constructed to extract spatial and temporal feature maps from the three-dimensional tensor. The spatial and temporal feature maps are then fused through a cross-attention mechanism to generate a depth feature vector F. fusion Specifically, it includes: Acquire the residual perturbation signal ΔC(t) of 50-100 consecutive frames and generate the corresponding three-dimensional tensor T; Let the spatial flow network structure include deformable convolutional layers, batch normalization layers, max pooling layers, and ordinary convolutional layers arranged in sequence, and let the spatial slice T of the three-dimensional tensor T be processed. t The input consists of a deformable convolutional layer with a 3×3 kernel, a stride of 1, and 64 output channels, used to focus on key spatial regions where signal energy is concentrated; a batch normalization layer is used to force normalization of the input distribution, and the output of the batch normalization layer introduces nonlinearity through the ReLU activation function; a max pooling layer with a 2×2 kernel and a stride of 2; and a regular convolutional layer with a 3×3 kernel and 128 output channels, used to output the spatial feature vector F. spatial ; Let the time-series flow consist of a first one-dimensional dilated convolutional layer, a gated linear unit (GLU), a second one-dimensional dilated convolutional layer, and a global average pooling layer, arranged sequentially. The three-dimensional tensor T is reorganized into a sequence. A time-series vector is defined for each antenna pair corresponding to a subcarrier location. The sequence-form three-dimensional tensor is input into the first one-dimensional dilated convolutional layer, which expands the receptive field. The first one-dimensional dilated convolutional layer has a kernel size of 5, a dilation rate of 2, and 64 output channels. The GLU selectively transmits time-series information through a gating mechanism. The second one-dimensional dilated convolutional layer also expands the receptive field, with a kernel size of 5, a dilation rate of 4, and 128 output channels. The global average pooling layer aggregates the time-series features of all location points, outputting a 128-dimensional time-series feature vector F. temporal ; The spatial feature vector F spatial and time series eigenvectors F temporal The corresponding query Q is obtained through linear transformation. s With Q t Key K s and K t Value V s and V t Then, the attention weight α of spatial features on temporal features is calculated. s→t Attention weight α of temporal features to spatial features t→s Finally, the two attention weights α s→t and α t→s The data is concatenated and fused through a fully connected (FC) layer to obtain a 256-dimensional deep feature vector F. fusion .

5. The AEB enhancement method based on WiFi imaging and multimodal sensing fusion according to claim 1, characterized in that, The training process of the occluded target detection network adopts an end-to-end training method. The loss function of the occluded target detection network is set to be the sum of three branches: binary cross-entropy loss, mean squared error loss function, and multi-component cross-entropy loss.

6. The AEB enhancement method based on WiFi imaging and multimodal sensing fusion according to claim 1, characterized in that, Step S4 involves performing Kalman filtering tracking on the fused target, maintaining the target ID, and estimating the motion state. The parameters of the Kalman filter are set as follows: state transition matrix A, observation matrix H, process noise covariance matrix Q, and observation noise covariance matrix R. The motion trend information of the occluded target and its historical trajectory are input into the time-series prediction model. This model first extracts local temporal features from the long-term sequence through dilated convolution, then applies multi-head self-attention to capture global long-range dependencies, and finally outputs the Gaussian distribution parameters of the predicted positions at each time point within the next 2-3 seconds—namely, the mean μ and the covariance matrix Σ—to quantify the prediction uncertainty. The dynamic collision risk assessment index R is calculated. risk R risk The calculation formula is: R risk =α0×P collision +β0×V rel +γ0×T time-to-collision , where P collision V represents the collision probability. rel T represents relative velocity. time-to-collision The collision time is represented by α0, β0, and γ0, which are weight terms.

7. The AEB enhancement method based on WiFi imaging and multimodal sensing fusion according to claim 6, characterized in that, The values ​​of α0, β0, and γ0 are 0.5, 0.3, and 0.2, respectively.

8. The AEB enhancement method based on WiFi imaging and multimodal sensing fusion according to claim 6, characterized in that, The implementation of the graded AEB control strategy based on the dynamic collision risk assessment index in step S5 includes the following: When R risk When the value is less than 0.3, the system determines it to be low risk and issues a warning, alerting the driver to the obstructed area ahead via the vehicle's display screen and sound. When 0.3≤R risk When the speed is less than 0.6, the system determines it to be of medium risk and applies partial braking to reduce the vehicle speed to a safe level to reduce the risk of collision. When R risk When the value is ≥0.6, the system determines it as a high risk and performs emergency braking to quickly decelerate the vehicle to a stop in order to avoid a collision.

9. An AEB enhancement system based on WiFi imaging and multimodal sensing fusion, used to implement the method described in any one of claims 1-8, characterized in that, include: The data acquisition unit includes a WiFi imaging sensing module, a visual acquisition module, and a millimeter-wave radar module, which are used to continuously acquire channel state information (CSI), RGB images, and target point cloud data within the coverage area, respectively. The spatiotemporal synchronization submodule communicates with the data acquisition unit and is used to timestamp synchronize the channel state information (CSI), RGB image, and target point cloud data, so that the data acquired from different sources are precisely aligned in time. The data preprocessing submodule, which communicates with the spatiotemporal synchronization submodule, is used to perform multipath-aware adaptive filtering on the Channel State Information (CSI) to obtain filtered CSI data. clean ; Obtain the theoretical background channel H at the current moment bg The filtered CSI data C clean With theoretical background channel H bg Subtraction yields the residual perturbation signal ΔC(t). This residual perturbation signal ΔC(t) is organized into a three-dimensional tensor of "antenna pair-subcarrier-time". A spatial-temporal flow spatiotemporal feature extraction network is constructed to extract spatial and temporal feature maps from the three-dimensional tensor. These spatial and temporal feature maps are then fused using a cross-attention mechanism to generate a deep feature vector F. fusion The RGB image is input into a multi-scale semantic-geometric joint detection head, which outputs the first perception result. The target point cloud data acquired by the millimeter-wave radar is input into a dynamic target-static obstacle joint clustering and association algorithm to obtain the second perception result. The depth feature vector F is then obtained. fusion It inputs the occluded target detection network and outputs the third perception result; The multimodal fusion submodule is used to construct a unified bird's-eye view fusion space, project the three perception results onto the same coordinate system, and use a cross-modal adaptive gating fusion mechanism to perform weighted fusion to obtain the fused BEV feature map; The decision control submodule is used to perform Kalman filtering tracking on the fused BEV feature map, maintain the target's tracking state and estimate its motion state; combine the motion trend information of the occluded target with its historical trajectory to predict the potential location area of ​​the target in the future and evaluate the dynamic collision risk assessment index; and execute a graded AEB control strategy based on the dynamic collision risk assessment index. The execution unit is used to connect with the decision control submodule, the vehicle AEB system, and the ESC vehicle stability system to execute braking or warning commands.

Citation Information

Patent Citations

  • AEB triggering method and device based on traffic conflict, and medium

    CN116620274A

  • Omnidirectional universal AEB method and system based on pure vision

    CN120298985A