Slow target detection and tracking method, system and storage medium based on wireless perception
Through wireless perception-based methods, including 4D data preprocessing, deep neural network training and multi-node collaborative perception, the problem of slow target detection and tracking in urban environments is solved, and high-precision and robust target detection and tracking are achieved.
Patent Information
- Application Number
- CN202411037083.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-31
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-07-31
AI Technical Summary
The prior art is difficult to effectively detect and track slow targets, such as pedestrians and drones, especially in urban environments, which are greatly affected by terrestrial clutter and multipath effects.
Using a wireless perception-based method, the 4D data dynamic model is constructed for pre-processing, the cumulative probability target enhancement is used to design a fully automatic labeling algorithm and deep neural network for supervised training, the domain migration strategy and target detection and tracking network are constructed, and a multi-node collaborative perception and switching method is proposed.
It significantly improves the detection ability of slow targets, reduces the impact of ground objects clutter, enhances the accuracy of detection and tracking, and ensures that the target is always in an effective tracking state and adapts to different scenarios and environments.
Smart Images

Figure CN118962659B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning and wireless signal processing, and in particular to a slow target detection and tracking method, system and storage medium based on wireless perception. Background Art
[0002] With the acceleration of urbanization, traffic congestion, traffic accidents and the increasing demand for urban traffic mobility, slow target detection and tracking based on wireless sensing has become a key task, which is essential for realizing intelligent traffic management, improving traffic efficiency and low-altitude safety. In actual traffic scenarios, slow targets include low-altitude flying targets and ground targets, such as drones, pedestrians, bicycles, etc. Their behavior in urban environments and traffic scenarios is of great significance. Such targets generally have the characteristics of slow radial speed, small Doppler shift, strong maneuverability, and small reflection coefficient, so there is a problem of insufficient detection capability. In addition, there is a serious overlap between the slow target spectrum and the ground object clutter spectrum in the Doppler spectrum, resulting in the detection of slow targets being seriously affected by ground clutter. In addition, the Doppler shift of the echo signal received by the receiving antenna itself is very small, while the surrounding environment, such as the random Doppler components generated by the wind swinging of plants, will cause the clutter spectrum to be broadened, which will seriously affect the detection of slow targets.
[0003] Pedestrians are one of the most common slow-moving targets in urban traffic. Their walking, stopping and crossing behaviors have a direct impact on traffic flow and road safety. Effective identification and tracking of pedestrians can provide important support for traffic signal control, pedestrian safety warning and other aspects, thereby improving the efficiency and safety of urban traffic. At the same time, with the rapid development of drone technology, drones as a new type of slow-moving target are increasingly becoming an important part of urban air traffic. However, the flexibility and diversity of drones bring challenges to their detection and tracking, especially in urban environments, where their flight trajectories are complex and changeable, which puts higher requirements on the accuracy and robustness of detection algorithms. Therefore, studying how to effectively detect and track drones is not only of great significance for urban airspace management and safety monitoring, but also an important application direction of current synaesthesia integration technology.
[0004] Therefore, there is an urgent need in the art for a technical solution that can effectively detect and track slow targets.
[0005] The information disclosed in this background technology section is only intended to enhance the understanding of the overall background of the invention and should not be regarded as an acknowledgment or any form of suggestion that the information constitutes the prior art already known to a person skilled in the art. Summary of the invention
[0006] The object of the present invention is to provide a method, system and storage medium for slow target detection and tracking based on wireless perception.
[0007] To achieve the above object, the present invention provides the following solutions:
[0008] A slow target detection and tracking method based on wireless perception, comprising:
[0009] For the echo signal of the slow target, a 4D data dynamic model is constructed for preprocessing to obtain preliminary 4DFFT cube data;
[0010] Performing target enhancement of cumulative probability on the 4DFFT cube data to improve the signal-to-noise ratio;
[0011] Design different fully automatic labeling algorithms for different target scenarios to achieve supervised training of deep neural networks;
[0012] Construct a domain migration strategy based on the maximum-minimum adversarial method to solve the problem of performance degradation of detection results caused by multipath effects on wireless signals in different scenarios;
[0013] Design a target detection and tracking network based on structured selective scanning spatial state to obtain the position of the target of interest and track it in real time;
[0014] A multi-node collaborative perception and switching method is proposed to achieve seamless switching and continuous tracking of perception nodes, ensuring that the target is always in an effective tracking state.
[0015] Optionally, the fully automatic labeling algorithm is specifically:
[0016] When labeling continuous frames, we first temporarily set the labels of samples that failed to be labeled to -1, then use Minkowski to construct the distance index matrix between all failed samples and all successful samples, and then perform nearest neighbor interpolation matching on the missing frames according to the Hungarian algorithm. Finally, we project the annotation information of the visual space into the radar space to complete the fully automatic labeling of continuous frames.
[0017] Optionally, the domain migration strategy based on the maximum-minimum adversarial method is specifically constructed as follows:
[0018] First, the parameters of the deep learning network are optimized through the maximum-minimum adversarial strategy to migrate the target feature distribution to the source feature distribution; then, the feature extractor parameters are optimized to minimize the decision differences between classifiers, thereby guiding the target sample features to migrate to the source feature distribution.
[0019] Optionally, the target detection and tracking network based on structured selective scanning spatial state is specifically:
[0020] The state-space sequence model is constructed by hidden states Maps an input function or sequence to an output And the matrix As an evolution parameter, and As projection parameters:
[0021] h′(t)=Ah(t))+Bx(t)
[0022] y(t) = Ch(t);
[0023] Define a time scale sampling parameter Δ to convert continuous parameters A and Β into discrete parameters The zero-order hold algorithm is used for conversion, and the conversion formula is as follows:
[0024]
[0025] The state equation suitable for processing discrete data is obtained:
[0026]
[0027] y t =Ch t
[0028] The model calculates the output through a global convolution:
[0029]
[0030] Where M is the length of the input sequence, is a structured convolution kernel;
[0031] The preprocessed 4DFFT radar cube data RF cube features are obtained through Chirp convolution and time downsampling network Then divide them into N patches, each of size Where D is the feature depth, (P,P,P) is the size of each patch; x p Linearly project it onto a vector of size S and add the category code t cls and position encoding Into a vector:
[0032]
[0033] in, is the nth patch, is a learnable projection matrix; then the token sequence T of the previous statel-1 Send it to the lth layer of the perceptual Mamba encoder to get the output T of the current state l ; Finally, the output is normalized and sent to the prediction head to obtain the final prediction confidence map
[0034] Optionally, the multi-node collaborative sensing method is specifically:
[0035] The received signals of multiple base stations are integrated, and the integrated data is processed. The possible correlation between the received signals is used to achieve mutual support between different signal data, thereby improving the perception accuracy.
[0036] Before signal fusion, the received signals of multiple nodes are processed by sliding windows in the same way. When three nodes are used to sense the target, a single node detects echo signals of P periods at a time. The sampling points of each periodic signal are N, and the number of sampling points in each sliding window is n. There is no overlap between the sliding windows, so a periodic signal has a total of Q = N / n windows. By setting n reasonably, the target signals detected by multiple nodes all fall within the same sliding window. Assuming it is the kth window, the signal of the jth window of the i-th periodic echo is:
[0037]
[0038] in are the target echoes of the three base stations, are the clutter interference signals of the three nodes, i∈[1,P], j∈[1,Q];
[0039] The received signal in each sliding window is a random variable. The cross-covariance C between the jth window signals of the i-th cycle echo of the three nodes is calculated. i,j , the formula is:
[0040]
[0041] According to the properties of cross-correlation, the cross-covariance C in the case of target presence and target absence is obtained i,j , whose value satisfies C (i,j) =0,i≠k; C (i,j) ≠0,i=k;
[0042] Then, the mutual covariance of the Q sliding windows of the ith cycle echo of each node signal is calculated and arranged in order to form the correlation sequence of the ith cycle echo:
[0043] C i =[C i,1 ,C i,2 ,…,C i,Q]
[0044] Furthermore, according to the above processing method, the correlation sequence of P periodic received signals is obtained, and the correlation matrix with a dimension of P×Q is obtained:
[0045]
[0046] The elements in the correlation matrix are accumulated by column, that is, accumulated at the same sliding window position, to obtain the cross-covariance sequence R, which is expressed as
[0047]
[0048] For the mutual covariance sequence R, the value of the sliding window mutual covariance corresponding to the presence of a target will be much larger than the value of the sliding window mutual covariance corresponding to the absence of a target. First, preliminarily determine the sliding window where the target is located, and then select a node as the main sensing node. The target sliding window signals of the remaining nodes are superimposed on the received signal of the main sensing node to form a fusion signal.
[0049] Optionally, the multi-node coordinated switching method is specifically:
[0050] Use all currently available nodes to estimate the position of the target. Through multi-node data fusion, the accurate position of the target can be obtained, the environmental conditions around the target position can be analyzed, and whether there are any obstructing objects can be detected. If the target is obstructed or exceeds the detection range of the current node, the best node is determined to locate and track the target based on the location of the target and other characteristics. Usually, the node closest to the target and not obstructed is selected as the target node. By dynamically adjusting the appropriate node to perceive the target, the system's target detection efficiency and accuracy can be improved, and continuous tracking and positioning of the target can be achieved in complex environments.
[0051] A slow target detection and tracking system based on wireless perception, comprising:
[0052] The preprocessing module is used to construct a 4D data dynamic model for preprocessing the echo signal of the slow target and obtain preliminary 4DFFT cube data;
[0053] The target enhancement module is used to perform cumulative probability target enhancement on 4DFFT cube data to improve the signal-to-noise ratio;
[0054] Fully automatic labeling algorithm module, which is used to design different fully automatic labeling algorithms for different target scenarios and realize supervised training of deep neural networks;
[0055] The domain migration strategy building module is used to build a domain migration strategy based on the maximum-minimum adversarial method to solve the problem of performance degradation of detection results caused by multipath effects on wireless signals in different scenarios;
[0056] The target detection and tracking network module is used to design a target detection and tracking network based on structured selective scanning spatial state, obtain the position of the target of interest in real time and track it;
[0057] The multi-node collaborative perception and switching module is used to propose a multi-node collaborative perception and switching method to achieve seamless switching and continuous tracking of perception nodes, ensuring that the target is always in an effective tracking state.
[0058] A storage medium stores a computer program, which runs the above-mentioned slow target detection and tracking method based on wireless perception when executed by a processor.
[0059] Compared with the prior art, the present invention has the following beneficial effects:
[0060] The present invention provides a method, system and storage medium for detecting and tracking slow targets based on wireless perception. By preprocessing the wireless echo signal, establishing a dynamic data model and using the 4DFFT cube data cumulative probability for target enhancement, the influence of ground clutter on target detection is significantly reduced, and the detection capability of slow targets is improved. At the same time, a visual-driven model is used to automatically annotate the target, and a target perception and tracking network of a structured selective scanning space state sequence is used to ensure the accurate acquisition of the target key point set, thereby enhancing the accuracy of detection and tracking. Secondly, in view of the fact that the characteristics of the back reflection signal received by the wireless sensor may be significantly different in different scenarios, directly applying the model in one environment to another environment may cause a sharp decrease in the accuracy of the detection result, a domain migration strategy based on the maximum-minimum adversarial method is constructed, which uses the data and model knowledge in the source domain (an environment with rich labeled data and trained) to migrate it to the target domain (a new environment with relatively scarce data), thereby improving the adaptability and accuracy of the model in the new environment. Finally, the proposed multi-node collaborative perception method avoids the signal blind spot problem of a single node, expands the coverage range, improves the positioning accuracy and the accuracy of target detection, and solves the limitations of single-node perception.
[0061] The present invention deeply mines the rich information contained in wireless signals, combines machine learning and deep learning technologies, and constructs an efficient target detection and tracking model, providing important technical support for the development of intelligent transportation systems. It effectively identifies and tracks pedestrian behavior, supports traffic signal control and pedestrian safety warnings, and copes with the complex flight trajectories of drones, improves urban airspace management and safety monitoring capabilities, and thus improves the overall safety and efficiency of urban transportation. The present invention is not only technologically innovative, but also has significant economic and social benefits. By improving the level of intelligence in traffic management and control, the rate of traffic accidents and congestion can be greatly reduced, traffic management costs can be saved, and the overall operation efficiency of the city can be improved. In addition, ensuring the safety of pedestrians and drones will help build a safe and efficient intelligent city transportation system, provide urban residents with a safer and more convenient travel environment, and promote the sustainable development of smart cities. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0063] Figure 1 A schematic flow chart of a slow target detection and tracking method based on wireless sensing provided in an embodiment of the present invention.
[0064] Figure 2 A schematic diagram of a 4DFFT processing chain based on a cumulative probability target enhancement algorithm provided in an embodiment of the present invention.
[0065] Figure 3 A schematic diagram of a cross-scenario model migration strategy based on a deep adversarial network provided in an embodiment of the present invention.
[0066] Figure 4 A schematic diagram of a target perception and tracking network based on structured selective scanning of spatial state sequences provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0067] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0068] The object of the present invention is to provide a technical solution that can effectively detect and track slow targets. In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0069] Embodiment 1:
[0070] This embodiment provides a slow target detection and tracking method based on wireless perception, including:
[0071] S1: For the echo signal of the slow target, a 4D data dynamic model is constructed for preprocessing to obtain preliminary 4DFFT cube data.
[0072] The traditional two-dimensional range Doppler spectrum cannot effectively distinguish slow targets from background clutter because their signals may be aliased in the spectrum. Therefore, in order to more accurately extract the target and suppress clutter, it is necessary to observe the time-varying motion state of the target, including distance, speed, azimuth and elevation information. To this end, wireless sensors (base stations, WIFI or millimeter wave radars) can be used to obtain the time-varying 4D observation data of slow targets, thereby obtaining the three-dimensional motion trajectory of the target.
[0073] According to the time division multiplexing MIMO strategy, a The MIMO array of transmit antennas and receive elements is combined into a a N e 2D virtual array of elements. Taking the frequency modulated continuous wave (FMCW) radar as an example, assuming that the wireless signal Transmitted to the surveillance domain, the backscattered signal from the target of interest can be approximately expressed as:
[0074]
[0075] Where τ = 2[R 0 +v 0 (t f +t p )] / c is the round trip time, t f ∈[0,T c ] represents fast time, T c represents the pulse duration, t p =n p T c and n p =0,1,...,N p -1 is the number of pulses, N p is the number of transmitted chirps, R 0 is the initial tilt range, v 0is the radial velocity of the slow target, c = 3 × 10 8 m / s is the speed of light.
[0076]
[0077] is a rectangular function, A is the amplitude, κ=B / T c is the chirp rate, B is the chirp bandwidth, d = (λ / 2) is the antenna spacing of the virtual array, λ = (c / f c ) is the wavelength, f c is the carrier frequency, θ 0 and are the initial azimuth and elevation angles, n a =0, 1, ..., N a -1 and n e =0, 1, ..., N e -1 are the column and row numbers of the two-dimensional virtual array, respectively, and the first antenna is taken as the reference element. The 4D Fourier transform (FFT) is applied to (1) and t f ,t m ,n a and n e , we can get
[0078]
[0079] where f r ,f v ,f a and f e are the frequency variables of range, Doppler, azimuth and elevation respectively, sinc(x)=(sinx / x) represents the Sinc function, A 0 is a constant.
[0080] According to the properties of the Sinc function, when f r =(2κR 0 / c),f v =(2v 0 / λ),f a =(dsinθ 0 / λ) and When , a peak point appears, which indicates that there is a target in the corresponding resolution unit. The traditional method uses a threshold to determine whether there is a target in the resolution unit. However, under low signal-to-noise ratio conditions, this method may cause dark targets (such as micro-UAVs) to be missed or information lost, seriously affecting the trajectory tracking performance. To this end, a cumulative probability target enhancement algorithm is proposed to update the 4DFFT cube data to improve the performance of subsequent target detectors, such as Figure 2 shown.
[0081] S2: performing target enhancement of cumulative probability on the 4DFFT cube data to improve the signal-to-noise ratio;
[0082] In order to cope with the situation of weak target signal and complex background clutter, a direct non-threshold 4D observation method is adopted instead of threshold-based measurement to obtain the continuous 3D trajectory of the target. Based on the collected original 4D observation, as shown in (2), it is transformed into a complex dynamic state estimation problem and solved jointly using the recursive Bayesian theory. Specifically, assuming represents the k-th timestamp target's k-th kinematic state vector, where x k ,y k and z k are the coordinates of the base station location as the origin, and are the corresponding kinematic parameters. Then, the dynamic motion in two adjacent timestamps can be described as k =FΠ k-1 +Gω k ,in
[0083]
[0084] represents the state transfer matrix, G represents the transfer matrix, ω k represents the process noise vector, T = N p T c is the time interval between two adjacent timestamps. In addition, in order to characterize the uncertainty of the existence of the target at the kth timestamp, there is a state indicator A k ={0,1} combined with the Markov transfer matrix, that is,
[0085]
[0086] Among them, P b =P(Λ k =1|Λ k-1 =0) and P d =P(Λ k =1|Λ k-1 =0) are the transition probabilities of the target’s birth and death respectively. On this basis, the recursive model with a state indicator between the kth timestamp and the k-1th timestamp can be expressed as
[0087]
[0088] Subsequently, by combining the predefined presence status indicator with the recursive Bayesian theory, the prior probability of the kth timestamp can be predicted:
[0089]
[0090] Where P(Πk |Π k-1 ,Λ k =1,Λ k-1 =1) represents the state transition probability from the k-1th timestamp to the kth timestamp, P 0 (Π k ) is the initial state probability, Θ 1:k =z 1 ,z 2 ,…,z k represents the range-velocity-azimuth-elevation spectrum of k consecutive frames, Contains N r ×N p ×N a ×N e Power observation, power observation in resolution units It can be modeled as
[0091]
[0092] in
[0093]
[0094] is a point spread function, (n r ,n p ,n a ,n e ) is the index of the resolution unit, N r is the total number of samples, n k is the observation noise, r, v, θ and are the distance, speed, azimuth and elevation variables respectively, η r ,η d ,η a and η e are the standard deviation factors of distance, speed, azimuth and elevation, respectively. is the target distance at the kth timestamp From the velocity-azimuth-elevation state, it can be calculated as, θ k =arctan(x k / y k )and Afterwards, relying on the Chapman-Kolmogorov equation, the posterior probability distribution function of the target at the kth timestamp can be updated as
[0095]
[0096] Where “∝” is the proportional operator, P(z k |Π k ,Λ k =1) indicates the likelihood probability:
[0097]
[0098] In theory, the optimal solution of the posterior probability can be obtained by using the above-mentioned Bayesian method. However, in order to optimize the calculation process and prevent limitations in practical applications, the sequential Monte Carlo method is introduced in the preprocessing process to approximate the posterior probability and the cumulative existence probability of the target.
[0099] The basic idea of the Sequential Monte Carlo method is to approximate the density function by propagating a set of independent samples with specific importance weights through the dynamic equation. In the initialization phase, a set of samples are randomly generated and initialized to positions in the monitoring area. The samples are then divided into continuous, newborn, and dead samples by the prior state vector and the corresponding existing probabilities. Subsequently, in the update phase, once a 4D observation is measured at timestamp k, its importance weight can be updated as
[0100]
[0101] in,
[0102]
[0103] represents the global likelihood ratio function, represents the important weight of the qth sample, and Respectively represent the influence area in the distance, speed, azimuth and elevation directions, represents the local likelihood ratio function, which can be estimated as
[0104]
[0105] in, is the likelihood probability of the resolution unit of the existing target, is the probability of the existence of only unit noise. In (14), it is assumed that the observation noise follows a zero-mean Gaussian distribution with a corresponding variance of σ 2 ,have
[0106]
[0107] It should be noted that this assumption is only used to simplify the problem, but the method is not limited to it. After the normalization and resampling process, the cumulative existence probability can be calculated as Timestamp k. If exceeds the preset threshold, the optimal state can be expressed as At this time, it can be determined that the target in the k+1th frame exists, and the sample state expectation can be calculated. Based on the original echo data without threshold processing, the preprocessing algorithm can effectively reduce the interference of clutter, filter out background noise, and then update the 4DFFT cube data. It effectively reduces the missed detection rate and information loss under low signal-to-noise ratio conditions, has good robustness for slow target processing, and provides more significant semantic information for the subsequent deep learning model construction.
[0108] S3: Design different fully automatic labeling algorithms for different target scenarios to achieve supervised training of deep neural networks.
[0109] (1) Fully automatic labeling algorithm:
[0110] For slow pedestrian detection and tracking scenarios, wireless sensors are usually deployed at short distances. Based on such short-distance scenarios, the present invention adopts a solution of data annotation and training in a small scene (source domain), and then uses a domain migration method to migrate the model to any real wireless sensing scene (target domain). This is because data in small scenes are often easy to collect and obtain, while data acquisition in the target domain is generally expensive, and the RF graph generated by the data generated by the target domain after processing has statistical similarity with the source domain. Therefore, an unsupervised domain adaptation method can be used to transfer knowledge of the small scene model, and then data enhancement is performed on the few sample data collected in the target domain. Finally, the data is put into the training set to fine-tune the model to make it suitable for the use of the scene. The scene migration algorithm is shown in (2).
[0111] Unlike vision, wireless sensing uses radio frequency (RF) intensity to describe spatial information instead of RGB images, so it does not have the spatial resolution of optical sensors, which makes it extremely difficult for wireless sensing technology to extract semantic information. In order to achieve supervised training of deep neural networks, a visually driven detection tracker is used to provide the true label of each frame target, and then the target information is mapped to the RF space for model training in step 3. Specifically, when annotating consecutive frames, the labels of those samples that failed to be annotated are first temporarily set to -1, and then the distance index matrix between all failed samples and all successful samples is constructed using Minkowski, and then the nearest neighbor interpolation matching is performed on the missing frames according to the Hungarian algorithm. Finally, the annotation information of the visual space is projected into the radar space to complete the fully automatic labeling of consecutive frames.
[0112] For the slow drone detection and tracking scenario, the perception range of the visual sensor is limited by its line of sight range. Therefore, in this scenario, the present invention uses a fully automatic labeling method based on GPS coordinate positioning, which is used as the Gaussian center point label of the real position. However, the conversion of GPS coordinates to XY coordinates usually needs to consider the geographic projection system, because the earth is a sphere, and XY coordinates are usually used on plane maps, so coordinate transformation is required. In the present invention, the universal transverse Mercator projection system UTM in formula (16) is used for coordinate transformation:
[0113]
[0114] N 1 =I+II·(A 2 / 2)+III·(A 4 / 4)+IV·(A 6 / 6)
[0115] E=E 0 +k 0 ·A+k 0 ·A 3 / 6·(1-η 2 +ν 2 )
[0116] +k 0 ·A 5 / 120·(5-18·η 2 +η 4 +14·ν 2 -58·η 2 ·ν 2 )(16)
[0117] Where N is the offset in the north direction, E is the offset in the east direction, A is the radian value of latitude, B is the radian value of longitude, e is the first eccentricity of the ellipsoid, a is the semi-major axis of the ellipsoid, b is the semi-minor axis of the ellipsoid, v is the radius of curvature of the ellipsoid, ρ is the radius of curvature of the meridian, and η is 2 are reflected polar coordinates, M is the radius of curvature of the meridian multiplied by the semimajor axis of the ellipsoid, I is the semiminor axis of the ellipsoid divided by the radius of curvature of the latitude, II is the semiminor axis of the ellipsoid divided by the radius of curvature of the latitude, III is the semiminor axis of the ellipsoid divided by the radius of curvature of the latitude, IV is the semiminor axis of the ellipsoid divided by the radius of curvature of the latitude, N 1 is the radius of curvature of the meridian divided by the radius of curvature of the latitude, E 0 is the displacement of the region in the east direction, k 0 is the scaling factor of the central meridian. After the coordinate transformation is completed, the above-mentioned fully automatic labeling method is used to complete the fully automatic labeling of continuous frames. The only difference from the previous method is that the positioning information of the camera is replaced by the positioning information of the GPS.
[0118] S4: Construct a domain migration strategy based on the maximum-minimum adversarial method to solve the problem of performance degradation of detection results caused by multipath effects on wireless signals in different scenarios;
[0119] Affected by the multipath effect, the characteristics of the back-reflected signals received by the wireless sensor may vary significantly in different scenarios. Therefore, directly applying the model in one environment to another environment may cause the accuracy of the monitoring results to drop sharply. The domain migration algorithm can utilize the data and model knowledge in the source domain (an environment with rich labeled data and trained) and migrate it to the target domain (a new environment with relatively scarce data), thereby improving the model performance in the target domain. This method can improve the adaptability and accuracy of the model in the new environment while reducing the cost of data annotation. Therefore, in order to achieve effective model migration in different monitoring scenarios, the present invention constructs a domain migration strategy based on the maximum and minimum adversarial method. With the help of the idea of adversarial learning, the deep network is guided to extract deep features with better feature distribution, and then the target features are moved to the source feature distribution. The target features are further transferred to the center of the distribution using the center alignment strategy, thereby improving the recognition ability of the model. The flow chart is shown as follows. Figure 3 shown.
[0120] The network consists of a feature extractor and two classifiers, and is designed to guide the sample features of the source scene to the target scene to eliminate feature offset. This method first optimizes the parameters of the deep learning network through the maximum-minimum adversarial strategy to migrate the target feature distribution to the source feature distribution. Specifically, by optimizing the classifier parameters to maximize the decision difference between the two classifiers, the decision boundary of the classifier closely surrounds the source sample feature distribution; then, by optimizing the feature extractor parameters to minimize the decision difference between the classifiers, the target sample features are guided to migrate to the source feature distribution. Furthermore, this method optimizes the feature generator parameters through the center alignment strategy to migrate the target features to the distribution center of the source features, and optimizes the feature extractor parameters to minimize the distance between the target features and the source class center, and finally realizes the movement of the target sample features to the corresponding source class distribution center. Through the above two strategies, the target features will be successfully embedded in the source feature distribution, so that the classifier suitable for the source scene can also obtain excellent recognition performance in the target scene.
[0121] S5: Design a target detection and tracking network based on structured selective scanning spatial state to obtain the position of the target of interest and track it in real time;
[0122] After the echo signal preprocessing and label automatic generation mentioned above, it is necessary to detect and track the target in the RF map. Inspired by the continuous state system, the present invention designs a target perception and tracking network based on structured selective scan space state sequential (Mamba). Figure 4 shown.
[0123] First, we introduce the State Space Sequential Model (SSM), which consists of a hidden state Maps an input function or sequence to an output And the matrix As an evolution parameter, and As projection parameters:
[0124]
[0125] Since the input of the network is discretized data, it is necessary to define a time scale sampling parameter Δ to transform the continuous parameters A and Β into discrete parameters. The commonly used conversion method is the zero-order hold algorithm, and the conversion formula is as follows:
[0126]
[0127] Therefore, the state equation suitable for processing discrete data is obtained:
[0128]
[0129] In order to parallelize the model, the model calculates the output through a global convolution:
[0130]
[0131] where M is the length of the input sequence, and is a structured convolution kernel. Figure 4 The overview of the designed network model is shown. Since the standard Mamba is designed for one-dimensional sequences, in order to handle wireless sensing tasks, a series of pre-processed 4DFFT radar cube data is first RF cube features are obtained through Chirp convolution and time downsampling network Then divide them into N patches, each of size Where D is the feature depth, (P,P,P) is the size of each patch. Next, x pLinearly project it onto a vector of size S and add the category code t cls and position encoding Into a vector:
[0132]
[0133] in, is the nth patch, is a learnable projection matrix. Then the token sequence T of the previous state l-1 Send it to the lth layer of the perceptual Mamba encoder to get the output T of the current state l Finally, the output is normalized and fed into the prediction head to obtain the final prediction confidence map In fact, the RF representation of the wireless signal and the real heat map are calculated in different coordinate systems, so they do not follow the process of feature selection at the same spatial position. In order to train the network end-to-end, binary cross entropy is used directly on pixels as the feature learning process. The specific implementation method of the network is to first use wireless sensors in a small scene to collect data and train the model, and then migrate the model to the actual scene through the cross-scene transfer learning method of step 2 (2).
[0134] S5: A multi-node collaborative perception and switching method is proposed to achieve seamless switching and continuous tracking of perception nodes, ensuring that the target is always in an effective tracking state.
[0135] (1) Multi-node collaborative perception
[0136] Single base station perception has problems such as limited coverage, low positioning accuracy, insufficient mobility support, and poor anti-interference, making it difficult to accurately and finely detect and track targets. In complex detection environments, the perception requirements of slow-moving targets cannot be met. In response to the above problems, the present invention proposes a multi-node collaborative perception and switching method. Multi-node collaborative perception refers to the use of multiple wireless sensor nodes to jointly complete perception tasks and improve perception performance. First, collaborative perception of multiple nodes can effectively avoid the problem of signal blind spots or insufficient coverage of a single node, thereby providing a wider and more stable coverage range. Secondly, by monitoring the target's signal at the same time through multiple nodes, multi-angle information can be used for positioning, improving positioning accuracy and accuracy, and coping with the challenges of target positioning in complex environments. In addition, when the target moves or crosses different areas, the multi-node perception system can achieve seamless switching and continuous tracking to ensure that the target is always in an effective detection state, avoid signal interruption or loss, and better support the perception of mobile targets.
[0137] Multi-node collaborative perception relies on data fusion between different nodes. The existing data fusion method is mainly distributed fusion, that is, a single node first processes the received signal to obtain characteristic information such as the target's distance, speed, azimuth, etc., and then uploads it to the fusion center, and finally performs data fusion through methods such as maximum likelihood. However, this type of method greatly compresses the received information and has problems such as loss of potential related information. Therefore, the present invention proposes a cross-correlation signal fusion algorithm, which first fuses the received signals of multiple base stations, and then processes the fused data, using the possible correlation between the received signals to achieve mutual support between different signal data, thereby improving perception accuracy.
[0138] Before signal fusion, the received signals of multiple nodes are first processed by sliding windows in the same way. Assuming that three nodes are used to sense the target, the detection time of a single node includes P cycles of echo signals, the sampling points of each cycle signal are N, and the number of sampling points in each sliding window is n. There is no overlap between the sliding windows, so a periodic signal has a total of Q = N / n windows. By setting n reasonably, the target signals detected by multiple nodes all fall within the same sliding window. Assuming it is the kth window, the signal of the jth window of the i-th cycle echo is:
[0139]
[0140] in are the target echoes of the three base stations, are the clutter interference signals of the three nodes, i∈[1,P], j∈[1,Q].
[0141] The received signals in each sliding window can be regarded as random variables, and the cross-covariance C between the jth window signals of the i-th cycle echo of the three nodes is calculated. i,j , the formula is:
[0142]
[0143] The correlation coefficient of the clutter interference signals between different nodes is very small and can be regarded as uncorrelated, while the same target signal detected by multiple nodes has a strong correlation. According to the properties of cross-correlation, the cross-covariance C in the case of target presence and target absence can be obtained: i,j , whose value satisfies C (i,j) =0,i≠k; C (i,j) ≠0,i=k.
[0144] Then, the mutual covariance of the Q sliding windows of the ith cycle echo of each node signal is calculated and arranged in order to form the correlation sequence of the ith cycle echo:
[0145] Ci =[C i,1 ,C i,2 ,...,C i,Q ] (twenty four)
[0146] Furthermore, according to the above processing method, the correlation sequence of P periodic received signals is obtained, and a correlation matrix with a dimension of P×Q can be obtained:
[0147]
[0148] Accumulate the elements in the correlation matrix by column, that is, accumulate them at the same sliding window position, and obtain the cross-covariance sequence R, which can be expressed as
[0149]
[0150] For the cross-covariance sequence R, the value of the sliding window cross-covariance when there is a target is much larger than the value of the sliding window cross-covariance when there is no target, so the sliding window where the target is located can be preliminarily determined. Then, a certain node is selected as the main sensing node, and the target sliding window signals of the remaining nodes are superimposed on the received signal of the main sensing node to form a fusion signal for signal processing in step 1. This method helps to enhance the signal-to-noise ratio of the signal and reduce clutter interference.
[0151] (2) Node switching
[0152] Use all currently available nodes to estimate the position of the target. Through multi-node data fusion, the accurate position of the target can be obtained, the environment around the target position can be analyzed, and whether there are any obstructing objects can be detected. If the target is obscured or exceeds the detection range of the current node, the best node is determined to locate and track the target based on the location of the target and other features. Usually, the node closest to the target and not obscured is selected as the target node. By dynamically adjusting the appropriate node to perceive the target, the system's target detection efficiency and accuracy can be improved, and continuous tracking and positioning of the target can be achieved in complex environments.
[0153] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0154] The principles and implementation methods of the present invention are described in this article using specific examples. The description of the above embodiments is only used to help understand the method and core idea of the present invention. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. A slow target detection and tracking method based on wireless perception, characterized in that: include: For the echo signal of the slow target, a 4D data dynamic model is constructed for preprocessing to obtain preliminary 4DFFT cube data; Performing target enhancement of cumulative probability on the 4DFFT cube data to improve the signal-to-noise ratio; Design different fully automatic labeling algorithms for different target scenarios to achieve supervised training of deep neural networks; Construct a domain migration strategy based on the maximum-minimum adversarial method to solve the problem of performance degradation of detection results caused by multipath effects on wireless signals in different scenarios; Design a target detection and tracking network based on structured selective scanning spatial state to obtain the position of the target of interest and track it in real time; A multi-node collaborative perception and switching method is proposed to achieve seamless switching and continuous tracking of perception nodes, ensuring that the target is always in an effective tracking state.
2. The method for detecting and tracking slow targets based on wireless sensing according to claim 1, characterized in that: The fully automatic annotation algorithm is specifically as follows: When labeling continuous frames, we first temporarily set the labels of samples that failed to be labeled to -1, then use Minkowski to construct the distance index matrix between all failed samples and all successful samples, and then perform nearest neighbor interpolation matching on the missing frames according to the Hungarian algorithm. Finally, we project the annotation information of the visual space into the radar space to complete the fully automatic labeling of continuous frames.
3. The method for detecting and tracking slow targets based on wireless sensing according to claim 1, characterized in that: The domain transfer strategy based on the maximum-minimum adversarial method is specifically as follows: First, the parameters of the deep learning network are optimized through the maximum-minimum adversarial strategy to migrate the target feature distribution to the source feature distribution; Then, the target sample features are guided to migrate towards the source feature distribution by optimizing the feature extractor parameters to minimize the decision differences between classifiers.
4. The method for detecting and tracking slow targets based on wireless sensing according to claim 1, characterized in that: The target detection and tracking network based on structured selective scanning spatial state is specifically: The state-space sequence model is constructed by hidden states Maps an input function or sequence to an output And the matrix As an evolution parameter, and As projection parameters: h′(t)=Ah(t)+Bx(t) y(t) = Ch(r); Define a time scale sampling parameter Δ to convert continuous parameters A and Β into discrete parameters , The zero-order hold algorithm is used for conversion, and the conversion formula is as follows: The state equation suitable for processing discrete data is obtained: The model calculates the output through a global convolution: Where M is the length of the input sequence, is a structured convolution kernel; The preprocessed 4DFFT radar cube data RF cube features are obtained through Chirp convolution and time downsampling network Then divide them into N patches, each of size Where D is the feature depth, (P,P,P) is the size of each patch; xp is linearly projected onto a vector of size S, and the category code tcls and position code are added Into a vector: in, is the nth patch, is a learnable projection matrix; then the token sequence T of the previous state l-1 Send it to the lth layer of the perceptual Mamba encoder to get the output T of the current state l ; Finally, the output is normalized and sent to the prediction head to obtain the final prediction confidence map 5. The method for detecting and tracking slow targets based on wireless sensing according to claim 1, characterized in that: The multi-node collaborative sensing method is specifically as follows: The received signals of multiple base stations are integrated, and the integrated data is processed. The possible correlation between the received signals is used to achieve mutual support between different signal data, thereby improving the perception accuracy. Before signal fusion, the received signals of multiple nodes are processed by sliding windows in the same way. When three nodes are used to sense the target, a single node detects echo signals of P periods at a time. The sampling points of each periodic signal are N, and the number of sampling points in each sliding window is n. There is no overlap between the sliding windows, so a periodic signal has a total of Q = N / n windows. By setting n reasonably, the target signals detected by multiple nodes all fall within the same sliding window. Assuming it is the kth window, the signal of the jth window of the i-th periodic echo is: in are the target echoes of the three base stations, are the clutter interference signals of the three nodes, i∈[1,P], j∈[1,Q]; The received signal in each sliding window is a random variable. The cross-covariance C between the jth window signals of the i-th cycle echo of the three nodes is calculated. i,j , the formula is: According to the properties of cross-correlation, the cross-covariance C in the case of target presence and target absence is obtained i,j , whose value satisfies C (i,j) =0,i≠k; C (i,j) ≠0,i=k; Then, the mutual covariance of the Q sliding windows of the ith cycle echo of each node signal is calculated and arranged in order to form the correlation sequence of the ith cycle echo: C i =[C i,1 ,C i,2 ,...,C i,Q ] Furthermore, according to the above processing method, the correlation sequence of P periodic received signals is obtained, and the correlation matrix with a dimension of P×Q is obtained: The elements in the correlation matrix are accumulated by column, that is, accumulated at the same sliding window position, to obtain the cross-covariance sequence R, which is expressed as For the mutual covariance sequence R, the value of the sliding window mutual covariance corresponding to the presence of a target will be much larger than the value of the sliding window mutual covariance corresponding to the absence of a target. First, preliminarily determine the sliding window where the target is located, and then select a node as the main sensing node. The target sliding window signals of the remaining nodes are superimposed on the received signal of the main sensing node to form a fusion signal.
6. The method for detecting and tracking slow targets based on wireless sensing according to claim 1, characterized in that: The multi-node coordinated switching method is specifically as follows: Use all currently available nodes to estimate the position of the target. Through multi-node data fusion, the accurate position of the target can be obtained, the environmental conditions around the target position can be analyzed, and whether there are any obstructing objects can be detected. If the target is obstructed or exceeds the detection range of the current node, the best node is determined to locate and track the target based on the location of the target and other characteristics. Usually, the node closest to the target and not obstructed is selected as the target node. By dynamically adjusting the appropriate node to perceive the target, the system's target detection efficiency and accuracy can be improved, and continuous tracking and positioning of the target can be achieved in complex environments.
7. A slow target detection and tracking system according to any one of claims 1 to 6, characterized in that: include: The preprocessing module is used to construct a 4D data dynamic model for preprocessing the echo signal of the slow target and obtain preliminary 4DFFT cube data; The target enhancement module is used to perform cumulative probability target enhancement on 4DFFT cube data to improve the signal-to-noise ratio; Fully automatic labeling algorithm module, which is used to design different fully automatic labeling algorithms for different target scenarios and realize supervised training of deep neural networks; The domain migration strategy building module is used to build a domain migration strategy based on the maximum-minimum adversarial method to solve the problem of performance degradation of detection results caused by multipath effects on wireless signals in different scenarios; The target detection and tracking network module is used to design a target detection and tracking network based on structured selective scanning spatial state, obtain the position of the target of interest in real time and track it; The multi-node collaborative perception and switching module is used to propose a multi-node collaborative perception and switching method to achieve seamless switching and continuous tracking of perception nodes, ensuring that the target is always in an effective tracking state.
8. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is executed.
Citation Information
Patent Citations
Small low-flight target visual detection tracking system and method thereof
CN111508002A
Arrival angle estimation method based on adversarial regularization deep neural network
CN111767791A