Arm motion sensing system based on UWB dual-node weight fusion
Through the UWB dual-node weight fusion system, the channel attention mechanism and long short-term memory network are used to process UWB radar signals, achieving high-precision arm motion recognition. This solves the problem of difficulty in identifying when the arm moves perpendicular to the radar axis in existing technologies, and improves the recognition accuracy and range.
Patent Information
- Application Number
- CN202510511448.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-09-09
AI Technical Summary
Existing arm motion recognition methods based on UWB radar technology cannot generate sufficient micro-Doppler signals when the arm moves perpendicular to the radar aiming axis, and it is difficult to achieve accurate recognition between similar arm movements. Traditional machine learning methods have limited accuracy when using small-scale data.
An arm motion perception system based on UWB dual-node weight fusion is adopted. By setting up two UWB nodes with signal transmission and reception functions, the channel attention mechanism and long short-term memory network are used for signal processing and feature extraction, and the convolutional neural network is combined for weight fusion learning and classification to achieve high-precision arm movement recognition.
It improves the accuracy and range of arm movement recognition, overcomes the limitations of single-view radar systems, reduces noise interference in Doppler images, provides more efficient behavior recognition feature expression, and improves the accuracy of arm movement classification.
Smart Images

Figure CN120611264A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of human motion perception, and in particular relates to an arm motion perception system based on UWB dual-node weight fusion. Background Art
[0002] With the development of wireless sensing networks and the Internet of Things (IoT), human motion sensing technology is increasingly being applied in fields such as smart homes, medical rehabilitation, sports training, and virtual reality. Compared to hand gestures, arm motion is more suitable for applications such as long-distance remote control and human-computer interaction. Furthermore, since the arm's radar contact surface is larger than that of hand gestures and less susceptible to jitter, wireless arm motion recognition has attracted increasing attention. In recent years, UWB radar technology has become well-suited for contactless human motion sensing due to its wideband, low power consumption, and high resolution. It can operate in dark environments and effectively distinguish between multiple paths or multiple targets, such as reflective information from the torso and limbs.
[0003] Existing arm motion recognition systems based on UWB radar technology mainly include the following three types: 1. Arm motion sensing systems based on motion trajectory recognition, but this method requires the user to carry a signal receiving device at all times, which is not very convenient and user-friendly. 2. Arm motion sensing systems based on spectral features combined with learning algorithms. Although this method has high accuracy, it requires a large amount of data, has high feature dimensions, and requires complex signal processing. 3. Arm motion sensing systems based on micro-Doppler. The frequency shift and tiny vibrations caused by arm motion can be recognized through the micro-Doppler effect, and the user does not need to carry a device. Compared with the above two methods, this method can provide a more natural and simple human-computer interaction method for smart devices.
[0004] However, a major issue with this approach is that when the arm moves perpendicular to the radar's aiming axis, it doesn't generate sufficient micro-Doppler signals. Furthermore, many existing methods struggle to accurately identify similar arm movements. Current research focuses on integrating this with image recognition and deep learning. However, deep learning models are complex and expensive, and traditional machine learning methods still have limited accuracy when using small datasets. Therefore, developing reliable recognition algorithms is crucial for achieving efficient arm motion sensing. Summary of the Invention
[0005] In response to the above problems, the present invention provides an arm motion perception system based on UWB dual-node weight fusion, which aims to collect motion information through dual-node radar and use learning algorithms such as channel attention mechanism for weight fusion to achieve high-precision arm motion classification and recognition.
[0006] According to an embodiment of the present disclosure, an arm motion perception system based on UWB dual-node weight fusion is provided, wherein the system includes a data acquisition unit, a signal processing unit and a weight fusion learning classification unit, wherein:
[0007] The data acquisition unit is used to set up two UWB nodes with signal transmission and reception functions in a scene with multiple passive targets. The signals transmitted by the two nodes are reflected by the target to form a multipath channel. The movement of the arm causes the radar signal to produce different Doppler effects when it is received, and the received signal is input into the signal processing unit;
[0008] A signal processing unit, configured to perform denoising and morphological filtering on the signal received by the data acquisition unit, and extract data features related to the action in the signal;
[0009] The weighted fusion learning classification unit is used to perform weighted fusion learning classification on data features through convolutional neural networks, long short-term memory networks, and channel attention mechanisms to obtain a variety of arm movement perceptions.
[0010] A further technical solution of the present invention is: in the data acquisition unit:
[0011] The transmitted signal is a periodic Gaussian pulse signal. The transmitted signal T(t) of the nth period is defined as:
[0012] T(t)=G(t-nT)exp(j2πf c t)
[0013]
[0014] Where t represents fast time, f c Indicates the carrier frequency;
[0015] The received signal R(n,t) after down-conversion after the signal reflected by the static target and the dynamic target is expressed as:
[0016]
[0017] Where Δt is the initial clock deviation between the transmitter and the receiver, is the signal transmission delay caused by the initial position between the dynamic target and the UWB radar, is the signal transmission delay between the dynamic target and the ultra-wideband radar in the nth signal cycle, represents the propagation attenuation coefficient of the signal in the nth signal cycle, z(t) is the signal noise, f d is the Doppler frequency of the dynamic target relative to the UWB radar.
[0018] A further technical solution of the present invention is: when the signal processing unit performs denoising and morphological filtering on the signal:
[0019] Use loopback filter to filter static clutter:
[0020] c(n,m)=αc(n-1,m)+(1-α)R(n,m)
[0021] y(n,m)=R(n,m)-c(n,m)
[0022] Where c(m,n) is the background clutter signal, y(m,n) represents the signal after static target removal, α is the update coefficient of the filter, and R(n,m) represents the received signal.
[0023] A further technical solution of the present invention is: the signal after removing clutter and static targets is converted into a frequency domain signal S(n,k) by short-time Fourier transform:
[0024]
[0025] Where k is the Doppler frequency, l is the adjustable width of the short-time Fourier transform window, w() is the window function, and the superscript H represents the number of multi-channels of the signal.
[0026] A further technical solution of the present invention is: when the signal processing unit extracts data features related to the action in the signal:
[0027] By calculating the amplitude sum within a fixed time window and segmenting individual movements based on a threshold;
[0028]
[0029] S(n,k) represents the signal after denoising and morphological filtering in the nth cycle, k is the Doppler frequency, and the time difference between N1 and N2 is the maximum motion duration value L, that is, the fixed width of each time window;
[0030] Using a sliding time window, the amplitude sum of different time windows W(m) is calculated, and the starting point of each motion signal is determined by the amplitude and peak value.
[0031] A further technical solution of the present invention is: selecting the amplitude change curve as the extracted feature, that is, selecting the amplitude peak at a single time point as the feature point, and using the distribution curve of the feature point on the time axis as the amplitude peak curve. This curve reduces the interference of low amplitude values on the classification and recognition of arm movements by focusing on the time-frequency distribution of the peak amplitude point;
[0032]
[0033] K threshold=δmaxK(n)
[0034] K(n)=0,if K(n)<K threshold
[0035] K(n) represents the velocity index of the maximum amplitude, δ represents the threshold penalty factor, and in order to ignore the smaller part of K(n), K(n) is multiplied by δ to obtain the threshold K threshold , used to decide whether to set to zero;
[0036] Peak amplitude curve PAC(n) for each movement: PAC(n) = K(n), n = n i ,n i +1,...,n f .
[0037] A further technical solution of the present invention is: extracting data features related to the action from the extracted signal based on morphological filtering processing, specifically comprising the following steps:
[0038] Determine the threshold, let P represent the energy distribution of the Doppler spectrum, and calculate the threshold as:
[0039] Determine the high-density energy area, that is, greater than the threshold P th The energy region is retained and the boundaries of the high energy density region are determined;
[0040] Use Hampel filter to filter the boundary;
[0041] Perform morphological processing on high power density areas;
[0042] The peak amplitude curve PAC(n) is smoothed in multiple steps: first, the Hampel method is used to remove outliers, then it is further smoothed using a moving average, and finally a smoother curve is generated by spline interpolation.
[0043] A further technical solution of the present invention is: the weight fusion learning classification unit uses a convolutional neural network to extract spatial features, uses a channel attention mechanism to calculate the weight of each time series, takes the weighted sum of the vectors of each time series as a feature vector, uses a long short-term memory network to learn the time information of the feature vector, and then performs Softmax classification of the fully connected layer to obtain the classification result.
[0044] A further technical solution of the present invention is: the weight fusion learning classification unit specifically includes the following steps:
[0045] Obtain feature vector through convolutional neural network: F f =CNN(P f ),F s =CNN(Ps ), P f 、P s They represent the peak amplitude curves of the two signal nodes respectively, and T is the length of the time series;
[0046] Merge the feature vectors into
[0047] Calculate global information of each channel i represents the number of channels;
[0048] Calculate channel attention weight: w i =softmax(f c2 (ReLU(f c1 (F i )))), f c1 , f c2 is a fully connected layer;
[0049] The fused weighted feature vector: F fused =w1⊙P f +w2⊙P s ;
[0050] After long short-term memory network: F LSTM =LSTM(F fused );
[0051] Output action category y out =softmax(W class ·F LSTM +b class ), W class represents the weight of the fully connected layer, b class A further technical solution of the present invention is: during the training process, a cross entropy loss function is used to optimize the model: Among them, y i is the true label, is the probability distribution predicted by the model. Through the back-propagation algorithm, the weights and biases in the long short-term memory network are adjusted to minimize the loss function.
[0052] The embodiment of the present disclosure provides an arm motion perception system based on UWB dual-node weight fusion, which has the following beneficial effects:
[0053] The proposed system is based on UWB dual-node weighted fusion. Unlike single-radar systems, it can capture Doppler information from the front view and obtain additional motion information from the side view. This system improves activity recognition accuracy and expands the range of activity recognition, overcoming the limitations of single-view radar systems in identifying diverse activities.
[0054] The system proposed in this paper improves the Doppler image processing method of dual-view radar: a morphologically constrained noise reduction method effectively reduces noise interference in the Doppler image. This makes the extracted peak amplitude curve features more stable and accurate, providing a more representative feature expression for subsequent behavior recognition.
[0055] The system of the present invention uses a CNN-Channel-Attention-LSTM method to perform dual-node weighted fusion learning and classification. A CNN captures all subtle changes in each time step and extracts spatial features; an LSTM network extracts temporal information from feature vectors; and Channel-Attention focuses on important information, calculating the weight of each time series. The weighted sum of all time series vectors is then taken as the feature vector. This dual-node weighted fusion learning and classification combines temporal and spatial information.
[0056] In summary, the arm motion perception system based on UWB dual-node weight fusion in the present invention collects motion information through dual-node radar, uses learning algorithms such as channel attention mechanism for weight fusion, and realizes high-precision arm motion classification and recognition.
[0057] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present invention and, together with the description, serve to explain the principles of the present invention.
[0059] Figure 1 Schematic diagram of the arm motion perception system based on UWB dual-node weight fusion in an embodiment of the present invention;
[0060] Figure 2 This is a principle block diagram of an arm motion perception system based on UWB dual-node weight fusion in an embodiment of the present invention;
[0061] Figure 3 Schematic diagram of the network structure in an embodiment of the present invention;
[0062] Figure 4 is the original Range-Time (RT) diagram in an embodiment of the present invention;
[0063] Figure 5 is a time-frequency diagram after STFT processing in an embodiment of the present invention;
[0064] Figure 6 This is a single motion effect diagram segmented based on a sliding time window in an embodiment of the present invention;
[0065] Figure 7 This is a flow chart of the PAC process based on morphological filtering in an embodiment of the present invention;
[0066] Figure 8 This is a flow chart of a dual-node weighted fusion learning classification algorithm based on a channel attention mechanism in an embodiment of the present invention;
[0067] Figure 9 This is a test experiment scene diagram in an embodiment of the present invention;
[0068] Figure 10 This is a test action analysis diagram in an embodiment of the present invention;
[0069] Figure 11 This is a comparison chart of the dual-node and single-node precision confusion matrices in the embodiment of the present invention. Figure 11 (a) is the confusion matrix diagram of the double-node weight fusion accuracy; Figure 11 (b) is the confusion matrix diagram of the positive single-point weight fusion accuracy; Figure 11 (c) is the side single node weight fusion accuracy confusion matrix diagram;
[0070] Figure 12 : is an accuracy confusion matrix diagram without morphological filtering in an embodiment of the present invention;
[0071] Figure 13 This is a precision confusion matrix diagram of an embodiment of the present invention in which the weights of two nodes are each 0.5. DETAILED DESCRIPTION
[0072] The present invention will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are intended only to illustrate the present invention and are not intended to limit the present invention. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions relevant to the present invention, not all structures.
[0073] Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flow charts. Although the flow charts describe the steps as sequential processes, many of the steps can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operation is completed, but can also have additional steps not included in the accompanying drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.
[0074] The embodiments of the present invention relate to the field of human motion perception and can be applied to medical rehabilitation systems, sports training systems and virtual reality systems, and more particularly to an arm motion perception method based on UWB (Ultra-Wideband) radar technology.
[0075] The embodiment provides the following content for an arm motion perception system based on UWB dual-node weight fusion:
[0076] like Figure 1 As shown, the system 100 includes a data acquisition unit 110, a signal processing unit 120, and a weight fusion learning classification unit 130, wherein:
[0077] The data acquisition unit 110 is configured to set up two UWB nodes with combined signal transmission and reception functions in a scenario containing multiple passive targets. The signals transmitted by the two nodes are reflected by the targets to form a multipath channel. The movement of the arm causes different Doppler effects on the radar signal when it is received. The received signal is then input into the signal processing unit.
[0078] The signal processing unit 120 is used to perform denoising and morphological filtering on the signal received by the data acquisition unit, and extract data features related to the action in the signal;
[0079] The weight fusion learning classification unit 130 is used to perform weight fusion learning classification on data features through a convolutional neural network, a long short-term memory network, and a channel attention mechanism to obtain a variety of arm movement perceptions.
[0080] Specifically, if Figure 2 As shown, the arm motion perception system based on UWB dual-node weighted fusion consists of three main components: data acquisition and signal processing, a weighted fusion learning and classification algorithm based on a CNN-Attention-LSTM, and the final output. During the data acquisition phase, two UWB nodes with combined signal transmission and reception functions are located in a scene with multiple passive static targets, one in front of and one to the side of the human target at 90 degrees. The signals transmitted by the two nodes are reflected by the target, forming a multipath channel. The arm motion causes different Doppler effects on the received radar signal. Based on traditional denoising and morphological filtering methods, motion data features are extracted. Then, a weighted fusion learning and classification method using a convolutional neural network (CNN), a channel attention mechanism, and a long short-term memory (LSTM) network is used to achieve diverse arm motion perception.
[0081] In one embodiment, Figure 3 As shown in the figure, the front and side nodes broadcast sensing signals and simultaneously receive echo signals reflected by multiple targets. The change in signal phase caused by human target motion produces a Doppler shift at the node, while the signal reflected by a static target has no Doppler frequency shift at the node.
[0082] Specifically, the radar acquisition device used in the embodiment is the XeThru X4 UWB sensor of Novelda. The transmitted signal is a periodic Gaussian pulse signal. The transmitted signal T(t) of the nth period can be defined as:
[0083] T(t)=G(t-nT)exp(j2πf c t)
[0084]
[0085] Where t represents fast time, f c represents the carrier frequency, A represents the amplitude, σ represents the amplitude variance, and T represents the period. The received signal after down-conversion after the signal reflected by the static target and the dynamic target can be expressed as follows:
[0086]
[0087] Where Δt is the initial clock deviation between the transmitter and the receiver, is the signal transmission delay caused by the initial position between the dynamic target and the UWB radar, is the signal transmission delay between the dynamic target and the ultra-wideband radar in the nth signal cycle, represents the propagation attenuation coefficient of the signal in the nth signal cycle, f d is the Doppler frequency of the dynamic target relative to the ultra-wideband radar, L represents the number of multipaths, z(t) is the signal noise, α l represents the attenuation coefficient of the lth signal propagation, τ l represents the signal propagation delay between the dynamic target and the UWB radar on the lth path. The sampled reflected signal can be defined by the matrix R(m,n), where n is the frame index and m is the number of fast time sampling points. Figure 4 As shown, this is the Rnage-Time (RT) diagram of the original data. As the slow time on the horizontal axis increases, the Rnage changes of the five actions can be clearly seen.
[0088] Taking the front node as an example, it is necessary to perform denoising and frequency domain processing on the received signal. First, use a loopback filter to filter out static clutter:
[0089] c(n,m)=αc(n-1,m)+(1-α)R(n,m)
[0090] y(n,m)=R(n,m)-c(n,m)
[0091] Where c(m,n) is the background clutter signal, y(m,n) represents the signal after static target removal, α is the filter update coefficient, R(n,m) represents the received signal, n is the frame index, and m is the number of fast-time sampling points.
[0092] Taking into account the breathing and heartbeat of the dynamic target, as well as the jitter caused by other parts of the body, it can be converted to the frequency domain through Short Time Fourier Transform (STFT) for further processing:
[0093]
[0094] Where k is the Doppler frequency, l is the adjustable width of the short-time Fourier transform window, w() is the window function, and the superscript H represents the number of multi-channel signals. Figure 5 The figure below shows the time-frequency diagram obtained after STFT transformation, which removes the influence of clutter and static targets. The Doppler changes over time for the five actions can be seen.
[0095] Determining the start time of a movement is key to segmenting multiple movements into individual movements within a period of time. In this embodiment, the sum of amplitudes within a fixed time window is calculated and segmented into individual movements based on a threshold.
[0096]
[0097] W(I) represents the time window sequence, I represents the index of the time window (i.e. the I-th sliding window), and the index range of negative speed is: K N1 To K N2 , K N1 , K N2 It is the range of negative Doppler velocity generated by arm movement relative to the radar receiver, from K N1 Speed to K N2 Speed; Index range for positive speed: K P1 To K P2 , K P1、 K P2 It is the range of Doppler velocity generated by arm movement relative to the radar receiver, from K P1 Speed to K P2 Speed; N1 and N2 represent the start and end indexes of the time window on the time axis, respectively. S(n,k) represents the signal after denoising and morphological filtering in the nth cycle. k is the Doppler frequency. The time difference between N1 and N2 is the maximum motion duration value L, that is, the fixed width of each time window.
[0098] Using the sliding time window, calculate the amplitude and peak value of different time windows W(I), and determine the starting point of each motion signal by the amplitude and peak value. Figure 6 Figure 2 shows three arm circle movements captured during a single experimental recording. The segmentation algorithm accurately identified the locations where the movements began, as shown in the window.
[0099] The velocity spectrum at each time point has a different amplitude distribution. Different movements have different amplitude distributions on the time and velocity axes. Therefore, arm movements can be identified based on the trend of velocity amplitude changes. The velocity index K(n) corresponding to the maximum amplitude can effectively reflect the arm's velocity. Conversely, due to various reasons for the body's twitching, the amplitude value is small and unrelated to the arm. Therefore, the amplitude change curve is selected as the extracted feature. That is, the peak amplitude of a single time point is selected as the feature point, and the distribution curve of the feature point on the time axis is used as the peak amplitude curve (Peak Amplitude Curve, PAC). This curve only focuses on the time-frequency distribution of the peak amplitude point, which can effectively reduce the interference of low amplitude values on the classification and recognition of arm movements.
[0100]
[0101] K threshold =δmaxK(n)
[0102] K(n)=0,if K(n)<K threshold
[0103] K(n) represents the velocity index of the maximum amplitude. In order to ignore the smaller part of K(n), the K(n) peak value is multiplied by δ to obtain the threshold value (denoted as K threshold ) to decide whether to set it to zero, and δ represents the threshold penalty factor. Therefore, the peak amplitude curve data PAC(n) of each movement can be obtained:
[0104] PAC(n)=K(n), where n=n i ,n i +1,...,n f Indicates the range of frame index n i to n f .
[0105] Morphological filtering is used to extract action-related data features from the signal. Specifically, this method effectively reduces noise interference in the Doppler image, making the extracted PAC features more stable and accurate, providing a more representative feature expression for subsequent action recognition.
[0106] Each column of a Doppler spectrum represents a frame, but the energy distribution varies between frames. When multiple frames are stitched together, one column of the spectrum becomes particularly dark, while the colors elsewhere are suppressed, resulting in poor learning results. Therefore, it is necessary to normalize the data in each column to ensure that the energy distribution of all frames is uniform.
[0107] Step 1: First, determine the threshold. Assuming P represents the energy distribution of the entire image, the threshold can be calculated as:
[0108] α represents the penalty factor, P represents the mean of P;
[0109] Step 2: Determine the high-density energy region, that is, the energy region greater than the threshold will be retained.
[0110] Step 3: Based on the high-density energy region, determine the boundary of the high-energy density region.
[0111] Step 4: Use the Hampel filter to filter the boundaries. Because the false target itself has relatively low energy, after the threshold setting and comparison in steps 1 and 2, the false target area will be divided into small blocks. The Hampel filter can then filter out these small blocks, thereby removing the false target.
[0112] Step 5: Perform morphological processing on the high power density area (opening operation: can remove small noise points, usually suitable for cleaning small isolated areas; closing operation: can fill small holes to make the boundary more coherent and smooth), improve and optimize the obtained boundary, and thus remove some small points or parts with high discreteness.
[0113] Step 6: Perform multi-step smoothing on the PCA curve: first use the Hampel method to remove outliers, then use moving average to further smooth it, and finally generate a smoother curve through spline interpolation to make the PAC curve smoother and more unified in features.
[0114] like Figure 7 As shown in the figure, the first row shows four PAC features extracted from the original Velocity-Time graph for the same action. It can be observed that when the micro-Doppler signal is weak or chaotic, the PAC features obtained for the same action are more different. The second row shows the PAC curve extracted from the Doppler graph after filtering. At this time, the PAC features become more distinct and consistent, but there are still some outliers and large differences in turning points, resulting in a lack of uniqueness and uniformity in the PAC curve features of the same action at different times. After further smoothing, the results in the last row show that the four obtained PAC curve features are more distinct and unified.
[0115] The weighted fusion learning classification unit uses a convolutional neural network to extract spatial features, uses the channel attention mechanism to calculate the weight of each time series, takes the weighted sum of the vectors of each time series as the feature vector, uses the long short-term memory network to learn the temporal information of the feature vector, and then performs Softmax classification on the fully connected layer to obtain the classification result.
[0116] Specifically, if Figure 8 The figure below shows the framework of the entire learning and classification algorithm. Using a one-dimensional CNN, all subtle changes at each time step can be captured. These changes are important for sequential learning in LSTM activity recognition. CNN and LSTM are combined for activity recognition, with CNN used for spatial feature extraction and LSTM networks used for temporal information learning. Channel-Attention can selectively filter out a small amount of important information from a large amount of information and focus on this important information. Using Channel-Attention to calculate the weight of each time series, then taking the weighted sum of all time series vectors as the feature vector, and then performing Softmax classification in the fully connected layer, this improves the classification results. The specific implementation steps are as follows:
[0117] 1) PAC curve acquisition
[0118] Let each PAC curve be P f (front) and P s (side), T is the length of the time series. For each action i (i = 1, 2, ..., 10), it can be expressed as:
[0119] 2) CNN feature extraction
[0120] Design a convolutional neural network with three convolutional layers and input each PAC curve into the CNN. Assume that the output of the CNN is a feature vector. Set the CNN operation as follows:
[0121] The convolution operation can be expressed as: X = W*PAC+b, where W is the convolution kernel, * represents the convolution operation, and b is the bias term.
[0122] Use nonlinear activation function (ReLU): X act =f(X)=max(0,X), after 2-3 layers of convolution processing, the output feature vector of each PAC curve is obtained: F f =CNN(P f ),F s =CNN(P s ),in,
[0123] Each convolutional layer is followed by a max pooling layer (nn.MaxPool1d), which reduces the size of the feature map while extracting important features. The formula for max pooling is: MaxPool(x) = max(x[i:i+k]), where k is the size of the pooling window. Pooling helps reduce computational effort and prevent overfitting.
[0124] Dropout layer: Through nn.Dropout(0.5), a part of neurons are randomly discarded during the training process. This can reduce the model's excessive dependence on certain features and effectively prevent overfitting.
[0125] 3) Feature vector merging
[0126] For each action i, two feature vectors are obtained: The two feature vectors are combined and input into the Softmax layer. The feature vectors are combined into:
[0127] 4) Channel-Attention Mechanism
[0128] After the feature vectors are merged, the channel attention mechanism is used to perform weighted fusion on the PACs of the two nodes processed by CNN in order to filter out important information. This is achieved by the following steps:
[0129] Calculate the global information for each channel: Among them F i (t) is the feature at time step t, indicating the number of channels.
[0130] Calculate channel attention weight: attn = softmax(f c2 (ReLU(f c1 (avg_pool)))), where f c1 , f c2 It is a fully connected layer.
[0131] 5) Softmax calculates channel weight vector
[0132] The output of the Softmax layer can be expressed as the weight of each time step under each channel, denoted as w: For each time step t, a weight vector between 0 and 1 is generated:
[0133] 6) Weight multiplication of weight and characteristic curve
[0134] Use the generated weight vector W to multiply the two original PAC curves to obtain the weighted feature curve: Where ⊙ represents element-wise multiplication.
[0135] 7) Weighted feature curve fusion
[0136] Add the two weighted feature curves to get the final fused feature vector:
[0137] 8) Input to LSTM for time dimension processing
[0138] The feature map processed by CNN and channel attention mechanism is then input into LSTM layer. The purpose of this step is to capture the temporal dependencies in the sequence data. The fused feature vector is input into LSTM for further processing. The basic unit of LSTM is to input F fused The processing can be expressed as:
[0139] Input gate i t :i t =σ(W i ·[F fused,t ,h t-1 ]+b i ).
[0140] Forget Gate f t :f t =σ(W f ·[F fused,t ,h t-1 ]+b f ).
[0141] Output gate o t :o t =σ(W0·[F fused,t ,h t-1 ]+b0).
[0142] Cell status update
[0143] Hide status update:h t =o t *tanh(C t ).
[0144] 9) Action Classification
[0145] In the last layer of LSTM, the hidden state h of the last time step is extracted T As the feature representation of the entire sequence. This feature vector is input into a fully connected layer and activation function for action classification: out =softmax(W class ·h T +b class ), where W class is the weight of the fully connected layer, b classis a bias term, and the final output y is a probability distribution with a length equal to the number of action categories, indicating the classification probability of each action.
[0146] 10) Training process
[0147] During training, the cross entropy loss function is used to optimize the model: Among them, y i is the true label, is the probability distribution predicted by the model. Through the back-propagation algorithm, the weights and biases in the LSTM are adjusted to minimize the loss function.
[0148] Table 1 is the pseudo code of the algorithm flow of the entire learning part.
[0149]
[0150]
[0151] In order to prove the effectiveness of the system of the present invention, a test experiment was conducted and applied to typical radar signal waveforms, including but not limited to linear frequency modulation continuous wave, pulse ultra-wideband signal, etc. The ultra-wideband signal module was used to verify the effect. In the experiment, two anchor points with ultra-wideband signal transceiver functions and one dynamic target were set up, such as Figure 9 As shown, the distance between the person and the front and side nodes is 1.2 meters. The action data is set as follows Figure 10 shown.
[0152] Figure 11 The effect of arm movement recognition using the algorithm of the present invention is given. Figure 11 (a) and Figure 11 (b) shows the results of arm movement recognition with only a single node on the front or side. Figure 11 (c) shows the effect of dual-node weight fusion. It can be seen that the dual-node fusion process achieves an accuracy of 92.81%, while the recognition accuracy using only a single node from the front is 81.61%, and the recognition accuracy using only a single node from the side is 52.14%. These results show that dual-nodes improve the accuracy of arm movement recognition by at least 10% compared to single-node methods.
[0153] In order to illustrate that morphological filtering can effectively improve the accuracy of arm movement recognition, the accuracy confusion matrix diagram before the improvement of PAC based on morphological filtering is given. Figure 12 As shown in the figure, it can be seen that before the morphological filtering method is used to process PAC, the classification accuracy after fusing the two nodes based on the channel attention mechanism weight is only 61.78%, which greatly reduces the recognition accuracy.
[0154] Finally, in order to illustrate the effectiveness of the dual-node weight fusion classification algorithm based on the channel attention mechanism proposed in this invention, the front and side nodes are directly set to a weight of 0.5 for each time step, and the final recognition results are as follows: Figure 13 As shown in Figure 3, the overall classification accuracy is 88.2%, while the weight fusion method based on the channel attention mechanism improves the accuracy by about 4.61%.
[0155] In addition, in addition to the above-mentioned units and modules, the system may also include other components. However, since these components are irrelevant to the content of the embodiments of the present disclosure, their illustration and description are omitted here.
[0156] Based on the technical solutions provided by the above embodiments, an arm motion perception system based on UWB dual-node weight fusion has the following beneficial effects:
[0157] The proposed system is based on UWB dual-node weighted fusion. Unlike single-radar systems, it can capture Doppler information from the front view and obtain additional motion information from the side view. This system improves activity recognition accuracy and expands the range of activity recognition, overcoming the limitations of single-view radar systems in identifying diverse activities.
[0158] The system proposed in this paper improves the Doppler image processing method of dual-view radar: a morphologically constrained noise reduction method effectively reduces noise interference in the Doppler image. This makes the extracted peak amplitude curve features more stable and accurate, providing a more representative feature expression for subsequent behavior recognition.
[0159] The system of the present invention uses a CNN-Channel-Attention-LSTM method to perform dual-node weighted fusion learning and classification. A CNN captures all subtle changes in each time step and extracts spatial features; an LSTM network extracts temporal information from feature vectors; and Channel-Attention focuses on important information, calculating the weight of each time series. The weighted sum of all time series vectors is then taken as the feature vector. This dual-node weighted fusion learning and classification combines temporal and spatial information.
[0160] In summary, the arm motion perception system based on UWB dual-node weight fusion in the present invention collects motion information through dual-node radar, uses learning algorithms such as channel attention mechanism for weight fusion, and realizes high-precision arm motion classification and recognition.
[0161] In this document, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a step or method that comprises a series of elements includes not only those elements, but also includes other elements not expressly listed, or also includes elements inherent to such step or method.
[0162] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should be considered to fall within the scope of protection of the present invention.
Claims
1. An arm motion perception system based on UWB dual-node weight fusion, characterized in that: The system includes a data acquisition unit, a signal processing unit and a weight fusion learning classification unit, wherein: The data acquisition unit is used to set up two UWB nodes with signal transmission and reception functions in a scene with multiple passive targets. The signals transmitted by the two nodes are reflected by the target to form a multipath channel. The movement of the arm causes the radar signal to produce different Doppler effects when it is received, and the received signal is input into the signal processing unit; A signal processing unit, configured to perform denoising and morphological filtering on the signal received by the data acquisition unit, and extract data features related to the action in the signal; The weighted fusion learning classification unit is used to perform weighted fusion learning classification on data features through convolutional neural networks, long short-term memory networks, and channel attention mechanisms to obtain a variety of arm movement perceptions.
2. The arm motion perception system based on UWB dual-node weight fusion according to claim 1 is characterized in that: In the data acquisition unit: The transmitted signal is a periodic Gaussian pulse signal. The transmitted signal T(t) of the nth period is defined as: T(t)=G(t-nT)exp(j2πf c t) Where t represents fast time, f c represents the carrier frequency, A represents the amplitude, σ represents the amplitude variance, and T represents the period; The received signal R(n,t) after down-conversion after the signal reflected by the static target and the dynamic target is expressed as: Where Δt is the initial clock deviation between the transmitter and the receiver, is the signal transmission delay caused by the initial position between the dynamic target and the UWB radar, is the signal transmission delay between the dynamic target and the ultra-wideband radar in the nth signal cycle, represents the propagation attenuation coefficient of the signal in the nth signal cycle, z(t) is the signal noise, f d is the Doppler frequency of the dynamic target relative to the ultra-wideband radar, L represents the number of multipaths, α l represents the attenuation coefficient of the lth signal propagation, τ l represents the signal propagation delay between the dynamic target and the UWB radar on the lth path.
3. The arm motion perception system based on UWB dual-node weight fusion according to claim 1 is characterized in that: When the signal processing unit performs denoising and morphological filtering on the signal: Use loopback filter to filter static clutter: c(n,m)=αc(n-1,m)+(1-α)R(n,m) y(n,m)=R(n,m)-c(n,m) Where c(m,n) is the background clutter signal, y(m,n) represents the signal after static target removal, α is the filter update coefficient, R(n,m) represents the received signal, n is the frame index, and m is the number of fast-time sampling points.
4. The arm motion perception system based on UWB dual-node weight fusion according to claim 3 is characterized in that: The signal after removing clutter and static targets is converted into a frequency domain signal S(n,k) through short-time Fourier transform: Where k is the Doppler frequency, l is the adjustable width of the short-time Fourier transform window, w() is the window function, and the superscript H represents the number of multi-channels of the signal.
5. The arm motion perception system based on UWB dual-node weight fusion according to claim 1 is characterized in that: When the signal processing unit extracts data features related to the action in the signal: By calculating the amplitude sum within a fixed time window and segmenting individual movements based on a threshold; W(I) represents the time window sequence, I represents the index of the time window, and the index range of negative speed is: K N1 To K N2 , K N1 , K N2 It is the range of negative Doppler velocity generated by arm movement relative to the radar receiver, from K N1 Speed to K N2 Speed; Index range for positive speed: K P1 To K P2 , K P1、 K P2 It is the range of Doppler velocity generated by arm movement relative to the radar receiver, from K P1 Speed to K P2 Speed; N1 and N2 represent the start and end indexes of the time window on the time axis, respectively. S(n,k) represents the signal after denoising and morphological filtering in the nth cycle. k is the Doppler frequency. The time difference between N1 and N2 is the maximum motion duration value L, that is, the fixed width of each time window. Using a sliding time window, the amplitude sum of different time windows W(m) is calculated, and the starting point of each motion signal is determined by the amplitude and peak value.
6. The arm motion perception system based on UWB dual-node weight fusion according to claim 5 is characterized in that: The amplitude change curve is selected as the extracted feature, that is, the amplitude peak at a single time point is selected as the feature point, and the distribution curve of the feature point on the time axis is used as the amplitude peak curve. This curve reduces the interference of low amplitude values on the classification and recognition of arm movements by focusing on the time-frequency distribution of the peak amplitude point; K threshold =δmaxK(n) K(n)=0,if K(n)<K threshold K(n) represents the velocity index of the maximum amplitude, δ represents the threshold penalty factor, and in order to ignore the smaller part of K(n), K(n) is multiplied by δ to obtain the threshold K threshold , used to decide whether to set to zero; The peak amplitude curve PAC(n) of each movement: PAC(n) = K(n), where n = n i ,n i +1,...,n f Indicates the range of frame index n i to n f .
7. The arm motion perception system based on UWB dual-node weight fusion according to claim 6 is characterized in that: The extraction signal is extracted based on morphological filtering processing and data features related to the action, specifically comprising the following steps: Determine the threshold, let P represent the energy distribution of the Doppler spectrum, and calculate the threshold as: α represents the penalty factor, represents the mean of P; Determine the high-density energy area, that is, greater than the threshold P th The energy region is retained and the boundaries of the high energy density region are determined; Use Hampel filter to filter the boundary; Perform morphological processing on high power density areas; The peak amplitude curve PAC(n) is smoothed in multiple steps: first, the Hampel method is used to remove outliers, then it is further smoothed using a moving average, and finally a smoother curve is generated by spline interpolation.
8. The arm motion perception system based on UWB dual-node weight fusion according to claim 1 is characterized in that: The weight fusion learning classification unit uses a convolutional neural network to extract spatial features, uses a channel attention mechanism to calculate the weight of each time series, takes the weighted sum of the vectors of each time series as a feature vector, uses a long short-term memory network to learn the time information of the feature vector, and then performs Softmax classification on the fully connected layer to obtain the classification result.
9. The arm motion perception system based on UWB dual-node weight fusion according to claim 8, characterized in that: The weight fusion learning classification unit specifically includes the following steps: Obtain feature vector through convolutional neural network: F f =CNN(P f ),F s =CNN(P s ), P f 、P s They represent the peak amplitude curves of the two signal nodes respectively, and T is the length of the time series; Merge the feature vectors into Calculate global information of each channel i represents the number of channels; Calculate channel attention weight: w i =softmax(f c2 (ReLU(f c1 (F i )))), fc1, fc2 are fully connected layers; The fused weighted feature vector: F fused =w1⊙P f +w2⊙P s , ⊙ represents element-by-element multiplication; After long short-term memory network: F LSTM =LSTM(F fused ); Output action category y out =softmax(W class ·F LSTM +b class ), W class represents the weight of the fully connected layer, and bclass is the bias term.
10. The arm motion perception system based on UWB dual-node weight fusion according to claim 9, characterized in that: During training, the cross entropy loss function is used to optimize the model: Among them, y i is the true label, is the probability distribution predicted by the model. Through the back-propagation algorithm, the weights and biases in the long short-term memory network are adjusted to minimize the loss function.