Human body behavior real-time identification system based on FPGA and millimeter wave radar
By deploying radar signal preprocessing and neural network identification modules in FPGAs, the millimeter wave radar data acquisition card is directly controlled to collect data and transmit it over Ethernet, which solves the problem that human behavior recognition system cannot be offline recognized in the prior art, and achieves high real-time and high accuracy real-time recognition.
Patent Information
- Application Number
- CN202510454034.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-09-02
AI Technical Summary
The existing human behavior recognition system based on millimeter wave radar and FPGA cannot achieve offline recognition, and the real-time performance is insufficient and the data transmission accuracy is low.
By deploying radar signal preprocessing module and neural network identification module in FPGA, the millimeter-wave radar data acquisition card is directly controlled to collect data, and transmit it over Ethernet to realize data encapsulation and transmission. Deploy it on FPGA in combination with the neural network model to form a offline real-time identification system.
Offline recognition is realized, real-time and data transmission accuracy are improved, data loss and damage are reduced, and a complete system that can be offline is formed.
Smart Images

Figure CN120579010A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the technical field of human behavior recognition, and more specifically, relates to a real-time human behavior recognition system based on FPGA and millimeter-wave radar. Background Art
[0002] With the promotion of deep learning technology, researchers have combined deep learning algorithms with millimeter-wave radar signal processing algorithms to perform human behavior recognition. However, current research on human behavior recognition mostly stays at the stage of designing algorithms in the host computer. In this case, the sensor itself, data transmission chain and host computer will make the overall system cost high and bulky. Moreover, the system composed of the above methods is only in the laboratory stage and cannot be separated from the computer to form a real-time human behavior recognition system based on millimeter-wave radar that can be offline.
[0003] The existing fall detection method and device based on millimeter-wave radar and FPGA (application number: 202210969297.7) uses radar to collect human posture data of falls and non-falls, preprocesses it to obtain feature maps, trains a convolutional neural network to obtain a convolutional neural network model, and then deploys the entire data processing flow and convolutional neural network model inside the FPGA for fall detection, thereby determining fall behavior.
[0004] While this approach can partially achieve detection outside the laboratory, it has significant limitations in practical applications. For details, see the corresponding paper for this patent, "FPGA Implementation of a Millimeter-Wave Radar-Based Fall Detection System." The paper clearly states: "The AWR1642 millimeter-wave radar and DCA1000 data acquisition card collect human activity data within a space and transmit the data to a miniPC via Ethernet. The miniPC then forwards the data to the FPGA via Ethernet." This demonstrates that this solution is not completely offline and PC-free. In practical applications, data transmission through the miniPC can lead to delays and data corruption. Therefore, a real-time human behavior recognition system based on FPGAs and millimeter-wave radar is urgently needed to address these issues. Summary of the Invention
[0005] The purpose of the embodiments of the present application is to provide a real-time human behavior recognition system based on FPGA and millimeter-wave radar to solve the technical problems in the prior art that the human behavior recognition system cannot achieve offline recognition, lacks real-time performance, and has low data transmission accuracy.
[0006] To achieve the above objectives, the embodiment of the present application provides a real-time human behavior recognition system based on FPGA and millimeter-wave radar, including a millimeter-wave radar, and also includes an electrically connected radar signal preprocessing module and a neural network recognition module; The radar signal preprocessing module is used to enable the FPGA to directly connect to and control the data acquisition card in the millimeter-wave radar to collect the radar's raw echo data, extract specific sampling points, and compress each frame of the chirp signal to obtain a one-dimensional vector. The specific sampling points and the one-dimensional vector are windowed to obtain windowed data, and then FFT is performed to obtain the FFT spectrum data. After normalization and conversion, the binarized distance-time graph and velocity-time graph are obtained, and the fusion generates a feature map. The neural network recognition module is used to store the parameters obtained from feature map training into the neural network model and deploy it to the FPGA to achieve offline recognition; FPGA is used to implement data encapsulation and transmission in Ethernet communication.
[0007] Preferably, the radar signal preprocessing module includes an interface module; The interface module is used to enable the FPGA to directly control the data acquisition card to collect human motion data and splice it to obtain complete radar raw echo data. At the same time, the data acquisition card triggers the next acquisition after completing one acquisition, thus capturing human motion data without interruption. The interface module includes a conversion module, a sending module and a receiving module; The conversion module is used to convert the GMII and RGMII of the Ethernet interface into each other; The sending module is used to implement sending preparation commands to the data acquisition card on the FPGA; The receiving module is used to receive the preparation completion signal, radar original echo data and sending completion signal sent by the data acquisition card.
[0008] Preferably, the radar signal preprocessing module further includes a cache module, a sampling module, a range compression module, a windowing module, an STFT module, a normalization module, a threshold processing module and a feature fusion module; The cache module is used to align the data of the interface module across clock domains; The sampling module is used to extract the specific sampling point of each chirp in the radar raw echo data; The distance compression module is used to add each frame chirp in the slow time dimension and then average it to obtain a one-dimensional vector.
[0009] Preferably, the windowing module is used to store the window function data through the ROM core, and the reading of the ROM data and the input of the one-dimensional vector and the specific sampling point are controlled by the input enable signal of the one-dimensional vector and the specific sampling point to be synchronized, and the one-dimensional vector and the specific sampling point are divided into real part and imaginary part, which are respectively multiplied with the window function data into the multiplier IP core, and the output result is then spliced with the real part and the imaginary part to obtain the windowed data.
[0010] Preferably, the STFT module is used to perform FFT on each frame of windowed data to obtain frequency spectrum data after FFT.
[0011] Preferably, the normalization module compensates the amplitude of the spectrum data after FFT by dividing by the length and average value of the window function data; The threshold processing module is used to convert the amplitude into decibel scale, and convert the speed time graph and distance time graph drawn by the spectrum data into binary distance time graph and speed time graph according to the threshold.
[0012] Preferably, the feature fusion module is used to perform transposition transformation and corresponding element fusion on the binarized distance-time graph and speed-time graph to generate a feature graph.
[0013] Preferably, the neural network recognition module includes a neural network module and a software and hardware combination module; The neural network module is used to store the weights and biases obtained after feature map training into the neural network model; The hardware and software combination module is used to deploy the neural network model into FPGA.
[0014] Preferably, the training preprocessing module is used to obtain feature maps, divide the training sets and test sets of different actions according to the actions and thresholds, build a neural network model after encapsulation and test to obtain weights and biases.
[0015] Preferably, the software and hardware combination module includes a data_ram_1 module, a conv_1 module, a pool_layer_1 module, a data_ram_2 module, a conv_2 module, a pool_layer_2 module, an FC1 module, an FC2 module, and an out_label module, for deploying the neural network model into the FPGA; The data_ram_1 module is used to cache feature maps; The conv_1 module is used to perform convolution operation on the feature map to obtain the convolved feature map; The pool_layer_1 module is used to perform pooling operation on the convolution feature map to obtain the pooled feature map; The data_ram_2 module is used to cache the pooled feature map. When the pooled feature map reaches the threshold, the next convolution begins. The conv_2 module is used to perform a secondary convolution operation on the pooled feature map to obtain the secondary convolution feature map; The pool_layer_2 module performs a secondary pooling operation on the feature map after the secondary convolution to obtain the feature map after the secondary pooling; The FC1 module is used to flatten the feature map after secondary pooling and perform multiplication and accumulation to obtain 100 nodes; The FC2 module is used to multiply and accumulate 100 nodes to obtain 6 nodes; The out_label module is used to output the maximum value among the 6 nodes and output the human action label corresponding to the maximum value to the display module for display.
[0016] The beneficial effects of the present application are as follows: the present application proposes a real-time human behavior recognition system based on FPGA and millimeter-wave radar. By deploying a state machine in the FPGA that can simulate the data encapsulation and transmission process of the PC in Ethernet communication, the FPGA can simulate the PC and directly control the start and stop state of the data acquisition card to collect data. This direct acquisition method can reduce the time for traditional PCs to repeatedly transmit data, and has good real-time performance. By directly transmitting information to the FPGA via Ethernet, the loss or damage of data packets caused by line interference is reduced, and the accuracy is higher and the robustness is better. The trained neural network model is deployed using FPGA, and combined with the above-mentioned direct control of the data acquisition card to acquire data and transmit via Ethernet, a complete system that can be offline and separated from the PC is formed, which solves the problem that the human behavior recognition system in the prior art cannot achieve offline recognition. In summary, the present application realizes offline recognition, and uses FPGA to simulate PC to directly transmit data via Ethernet with good real-time performance, good robustness and high accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0018] Figure 1 A schematic diagram of a real-time human behavior recognition system based on FPGA and millimeter-wave radar provided in one embodiment of the present application; Figure 2 A preparation command for a data acquisition card provided in one embodiment of the present application; Figure 3 A state diagram of a sending module provided in one embodiment of the present application; Figure 4 A data packet of the data acquisition card provided in one embodiment of the present application; Figure 5 A data packet of raw radar echo data from a data acquisition card provided in one embodiment of the present application; Figure 6 The data acquisition card provided in one embodiment of the present application has completed sending a data packet; Figure 7 A state diagram of a receiving module provided in one embodiment of the present application; Figure 8A structural diagram of a cache module provided in one embodiment of the present application; Figure 9 A one-dimensional vector output by a distance compression module provided in an embodiment of the present application; Figure 10 A structural diagram of a windowing module provided in one embodiment of the present application; Figure 11 A normalized speed-time graph outputted according to an embodiment of the present application; Figure 12 A speed-time graph output after threshold processing provided by an embodiment of the present application; Figure 13 Modules required for the neural network hardware deployment provided in one embodiment of the present application; Figure 14 A structural diagram of the data_ram_1 module provided in one embodiment of the present application; Figure 15 A ROM diagram of the conv_1 module provided in one embodiment of the present application; Figure 16 A schematic diagram of the pool_layer_1 module provided in an embodiment of the present application; Figure 17 A schematic diagram of the conv_2 module provided in one embodiment of the present application; Figure 18 A schematic diagram of a RAM cache provided in one embodiment of the present application; Figure 19 The weight parameter arrangement method provided in one embodiment of the present application; Figure 20 A simulation diagram of an interface module provided in an embodiment of the present application; Figure 21 The radar raw echo data provided in one embodiment of the present application; Figure 22 The first group of 256 data of the radar signal preprocessing module provided in one embodiment of the present application; Figure 23 Data output by the signal processing algorithm provided in one embodiment of the present application; Figure 24 The label finally output by the neural network recognition module provided in one embodiment of the present application; Figure 25 The specific values of each node output by the online test of the neural network provided in one embodiment of the present application. DETAILED DESCRIPTION
[0019] In order to make the technical problems, technical solutions and beneficial effects to be solved by this application more clearly understood, this application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0020] This application provides a real-time human behavior recognition system based on an FPGA and millimeter-wave radar. This system abandons the traditional method of using a miniPC in place of a PC to control data acquisition and forwarding. Instead, it incorporates a complete, offline, PC-independent module: a radar signal preprocessing module. This module connects directly to the millimeter-wave radar's data acquisition card, eliminating the traditional miniPC data forwarding mechanism and truly achieving offline recognition without PC control. This approach reduces the time required for the miniPC to repeatedly transmit data, resulting in improved real-time performance. Directly transmitting data to the FPGA via Ethernet reduces packet loss or corruption due to line interference, resulting in higher accuracy and greater robustness.
[0021] See also Figure 1 , which is a schematic diagram of a real-time human behavior recognition system based on FPGA and millimeter-wave radar provided in one embodiment of the present application, including: an electrically connected millimeter-wave radar, a radar signal preprocessing module and a neural network recognition module.
[0022] The millimeter-wave radar consists of a radar sensor chip, an evaluation board, and a data acquisition card for receiving raw radar echo data. The radar sensor chip uses the AWR1642 millimeter-wave radar sensor, with two transmitting antennas and four receiving antennas arranged in a planar array. The evaluation board uses the AWR1642BOOST 77GHz millimeter-wave sensor evaluation board, which has two transmitters and four receivers arranged in a planar array. The data acquisition card uses the DCA1000 high-speed data acquisition card, which streams data in real time via 1Gbps Ethernet. There are no restrictions on the models of the radar sensor chip, evaluation board, and data acquisition card; you can customize them based on your needs.
[0023] The parameters of the millimeter-wave radar are set as follows: 128 frames are set for one data acquisition, 128 chirps are set for each frame, and 64 sampling points are set for each chirp.
[0024] The radar signal preprocessing module includes an interface module, a cache module, a sampling module, a range compression module, a windowing module, a STFT module, a normalization module, a threshold processing module and a feature fusion module.
[0025] The radar signal preprocessing module is used to enable the FPGA to directly connect to and control the data acquisition card in the millimeter-wave radar to collect the radar's raw echo data, extract specific sampling points, and compress each frame of the chirp signal to obtain a one-dimensional vector. The specific sampling points and the one-dimensional vector are windowed to obtain windowed data, and FFT is performed to obtain the FFT spectrum data. After normalization and conversion, the binarized distance-time diagram and velocity-time diagram are obtained, which are fused to generate a feature map.
[0026] The interface module includes a conversion module, a sending module and a receiving module. It is used to enable the FPGA to directly control the data acquisition card to collect human motion data and splice it to obtain complete radar raw echo data. At the same time, after the data acquisition card completes one acquisition, it triggers the next acquisition and captures human motion data uninterruptedly.
[0027] The conversion module is used to convert the GMII and RGMII of the Ethernet interface.
[0028] This conversion is designed to adapt to the communication requirements between different physical layer (PHY) chips and media access control layers (MAC), while reducing the number of pins and lowering power consumption. The RGMII interface is suitable for communication rates of 10M / 100M / 1000Mbps and has strong compatibility.
[0029] In an optional embodiment, RGMII uses a 4-bit data interface. At a 1000Mbps communication rate, the transmit clock line ETH_TXC and the receive clock line ETH_RXC operate at a 125MHz clock frequency. DDR (Double Data Rate) is used to transmit 8-bit data signals within a single clock cycle. The rising edge transmits or receives the lower 4 bits of data, while the falling edge transmits or receives the upper 4 bits of data. The control signals ETH_TXCTL and ETH_RXCTL also use DDR, transmitting two control signals within a single clock cycle. The rising edge transmits or receives the data enable (TX_EN / RX_DV) signal, while the falling edge transmits or receives the exclusive-OR value of the enable signal and the error signal (TX_ERR xor TX_EN, RX_ERR xor RX_DV). When the valid signal RX_DV is high (indicating valid data) and the error signal RX_ERR is low (indicating no data errors), the XOR result is high. Therefore, only when the rising and falling edges of the control signals ETH_RXCTL and ETH_TXCTL are high at the same time, the sent and received data can be valid and correct.
[0030] The sending module is used to enable the FPGA to simulate the PC and send a preparation command to the data acquisition card.
[0031] See also Figure 2The sending module is the sending module of UDP data packets. By using wireshark (network packet analysis software) to capture the data packets between the data acquisition card and its corresponding host computer, it can be seen that its starting valid data packet is 5a_a5_05_00_00_00_aa_ee. For details, see Figure 2 .
[0032] FPGA simulation of PC implementation: Deploy a state machine in the FPGA that can simulate the data encapsulation and transmission process of the PC in Ethernet communication. Specifically, the network communication process of the PC can be equivalent to the jump process of the state machine, and different states correspond to different stages of the communication process. By implementing this state machine jump process on the FPGA, the behavior of the PC in network communication can be simulated. For details, see Figure 3 Therefore, the sending module can make the FPGA board simulate the PC, and then send the preparation command directly to the data acquisition card without the host computer.
[0033] See also Figure 3 The following is a state diagram of the state machine implementation of the sending module, including the state jump and trigger signal of each step. From the diagram, we can intuitively see the function implemented by each state and the conditions for jumping to the next state.
[0034] In the initial state st-idle, the system waits for a trigger to send a command. When a command is sent, the system enters the IP header checksum state st-check-sum to calculate the IP header checksum. After the checksum is complete, the system enters the preamble and start-of-frame delimiter (SFD) state st-preamble. After completing the preamble and SFD, the system enters the Ethernet frame header state st-eth-head. Once the Ethernet frame header is sent, the system enters the IP and UDP header states st_ip_head. Once the IP and UDP headers are sent, the system enters the data state st_tx_data to send the actual start-up packet. Once the data is sent, the system enters the CRC state st_crc to send the cyclic redundancy check (CRC). Upon completion, the system returns to the initial state st-idle, completing the loop. Each state has a corresponding transition condition and a skip condition (skip_en). These conditions determine whether a state can skip certain steps. For example, if a required operation in a state has been completed, the skip_en condition allows the system to jump directly to the next state without going through intermediate states.
[0035] In summary, by implementing IP header verification, sending preamble and SFD, Ethernet frame header, IP header, UDP header, data payload, and CRC, FPGA can simulate the data encapsulation and transmission process of PC in Ethernet communication.
[0036] The receiving module is used to receive the preparation completion signal, radar raw data and sending completion signal sent by the data acquisition card.
[0037] See also Figure 4 , which is a preparation completion data packet obtained by the receiving module according to the preparation completion signal sent by the data acquisition card in one embodiment of the present application. It can be seen from the wireshark packet capture that the valid data of the preparation completion data packet of the present application is 5a_a5_05_00_00_00_aa_ee.
[0038] It is worth noting that in the radar raw data packet, the first 10 bytes of valid data in each radar raw data packet are not the radar echo data. The real echo data starts from the 11th byte of valid data in the radar raw data packet, such as Figure 5 Therefore, a state is set when receiving the 10 bytes, named st_radar_head state. Figure 7 It will be shown in the.
[0039] After the acquisition is completed, the data acquisition card will send a data packet with valid data of 5a_a5_0a_00_00_01_aa_ee, indicating that the sending is completed. Figure 6 shown.
[0040] See also Figure 7 , is the state diagram of the receiving module of this application, and the specific steps are as follows: In the initial state st_idle, it is ready to receive data. After receiving the first byte of the preamble, it jumps to the state st_preamble for receiving the preamble and the start of the frame. After receiving the correct preamble and SFD (start of frame), it jumps to the state st_eth_head for receiving the Ethernet frame header. In the state st_eth_head for receiving the Ethernet frame header, if the MAC error or the protocol error occurs, it jumps to the state st_rx_end for receiving the Ethernet frame header and ends the reception. Otherwise, after receiving the Ethernet frame header, it jumps to the state st_ip_head for receiving the IP header. If the protocol error or the IP address error occurs, it jumps to the state st_rx_end for receiving the IP header and ends the reception. Otherwise, it jumps to the state st_rx_end for receiving the IP header and ends the reception. Receive the IP header. When the IP header is completely received, jump to the UDP header receiving state st_udp_head, receive the UDP header. When the UDP header is received, jump to the radar data header receiving state st_radar_head, receive the radar data header. If a ready signal is received or a completion signal is sent, jump to the reception completion state st_rx_end, and the LED state flips; otherwise, enter the radar valid data receiving state st_rx_data, receive the radar's original echo data, after receiving the radar's original echo data, jump to the reception completion state st_rx_end, end the reception, jump to the initial state st_idle, wait for the next reception, and repeat this cycle.
[0041] The cache module is used to solve the problem of data alignment across clock domains and ensure the correct transmission and processing of data at different clock frequencies.
[0042] Specifically, since the clock frequency of the interface module's transmit clock line ETH_TXC and receive clock line ETH_RXC is 125Mhz, and the standard clock is 50M, a data cache module is needed to ensure data alignment to solve the above problem. Here, asynchronous FIFO is used as the main component of the cache module.
[0043] The main implementation process of the cache module is: write the radar raw echo data sent by the interface module into the FIFO, count while writing, and enable reading the FIFO when 64 data are written. During the process of reading the FIFO, writing the FIFO is also carried out at the same time, and the cycle repeats. The structure of the cache module is as follows: Figure 8 shown.
[0044] In the figure, the input signals include rst_n, wr_clk, data_en, and data_in. Specifically, rst_n is a reset signal, which is valid at a low level. When this signal is at a low level, the cache module is reset. wr_clk is a write clock signal, which is used to control the data write operation. In this embodiment, the frequency of the write clock signal is 125MHz. data_en is a data enable signal, which is valid at a high level. When the data enable signal is at a high level, data is allowed to be written to the cache module. data_in is the data written to the cache module.
[0045] The output signals include rd_clk, data_out_en, and data_out. Specifically, rd_clk is the read clock signal used to control data read operations. In this embodiment, the read clock signal frequency is 50 MHz. data_out_en is the data output enable signal, active high. When the data output enable signal is high, the radar raw echo data can be read from the cache module. data_out is the data read from the cache module.
[0046] The cache module uses an asynchronous FIFO (first-in, first-out) memory to buffer data. The asynchronous FIFO can safely transmit radar raw echo data between different clock domains, avoiding data loss or misalignment caused by clock frequency differences.
[0047] The sampling module is used to extract the specific sampling point of each chirp in the original radar echo data.
[0048] Since each radar raw echo data acquisition consists of 128 frames, each frame consists of 128 chirps, and each chirp has 64 sampling points, this application sets a setting in the sampling module to select the 32nd sampling point of each chirp as the specific sampling point and uses FIFO (first-in-first-out) for data storage. After each frame of data, that is, 128 sampling points, is extracted, the data is transferred to the next processing module. During the entire data acquisition process, a total of 128 such transmissions are performed, so that the extracted data is more representative.
[0049] The distance compression module is used to compress each chirp frame in the slow time dimension to obtain a one-dimensional vector.
[0050] Specifically, in an optional embodiment, the distance compression module compresses the 128 chirps in each frame in the slow time dimension, compressing 64 of the 128 chirps into a one-dimensional vector form. Thus, one frame is compressed into two one-dimensional vectors. The compression method includes adding and averaging the sampling points of each 64 chirps in each frame at fixed sampling positions along the slow time dimension to obtain a one-dimensional vector of scale 64x1.
[0051] Assumptions Indicates the Chirps (i=0,1,2,...,63) at the jth sampling point (j=0,1,2...,N-1) in the fast time dimension, where N is the number of sampling points in each chirp), then for each sampling point j, we calculate the average of all 64 chirps as follows: ; This results in a one-dimensional vector of length N. , which consists of the average values corresponding to all sampling points j, and the formula is as follows: ; See also Figure 9 Figure 2 shows the distance compression module outputting a one-dimensional vector. The distance compression module can be implemented using two FIFOs, each containing 64 one-dimensional vectors (numbered 0 to 63 in the figure). The one-dimensional vector numbered 0 is written to the first FIFO. Upon the arrival of the one-dimensional vector numbered 1, a read operation is performed on the first FIFO to ensure data alignment. The added vectors are then written to the second FIFO. Upon the arrival of the one-dimensional vector numbered 2, a read operation is performed on the second FIFO. The added vectors are then written to the first FIFO. This cycle repeats until the arrival of the one-dimensional vector numbered 63. The summation and averaging operations are performed to obtain the distance-compressed one-dimensional vector, which is then passed to the windowing module. In this embodiment, each data acquisition cycle transmits 2x128 64x1 one-dimensional vectors.
[0052] The windowing module is used to perform windowing processing on one-dimensional vectors and specific sampling points.
[0053] The windowing module applies windowing to each frame of one-dimensional vectors or specific sampling points sent by the range compression module or sampling module. In the velocity-time graph branch, a Hamming window function is used to window the specific sampling points obtained by the sampling module. The window length is 256, with no overlap. In the range-time graph branch, a Hamming window is used to window the one-dimensional vectors obtained by the range compression module. The window length is 64, with no overlap. This approach effectively reduces spectral leakage without excessively sacrificing frequency resolution, significantly improving the performance of the downstream STFT module.
[0054] Specifically, the input signals of the windowing module are the system clock signal sys_clk, the system reset signal sys_rst, the upper layer data data_in, the input enable signal data_en of the upper layer data, the output enable signal win_en and the output windowed data win_data. For a specific block diagram, see Figure 10 The specific workflow is as follows: the window function data is stored in the ROM core in advance, the upper-layer data input enable signal data_en is used to control the synchronization of ROM data reading and upper-layer data data_in input, the upper-layer data data_in is divided into real and imaginary parts, and the real and imaginary parts are respectively multiplied with the window function data in the multiplier IP core. The output result is then spliced with the real and imaginary parts to obtain the windowed data win_data, and the output enable signal win_en is used as the input signal of the FFT core.
[0055] The STFT module is used to perform fast Fourier transform on each frame of windowed data output by the windowing module to obtain the spectrum data after FFT.
[0056] Because millimeter-wave radar signals are typically non-stationary and may contain Doppler shift and other transient features, the use of STFT can provide more detailed time-frequency analysis, thereby better understanding and processing the original echo data, and thus obtaining the original time Doppler feature map. To extract key features and facilitate neural network training, the feature map output by the STFT module is not the final feature map. Instead, it needs to go through subsequent normalization and threshold processing modules to obtain the processed velocity time map and distance time map.
[0057] The normalization and threshold processing module is used to normalize and transform the spectrum data after FFT to obtain the distance time graph and speed time graph.
[0058] Specifically, the amplitude loss caused by the windowing module is first compensated by dividing the data by the length and average value of the window function, making the spectrum data after FFT closer to the true amplitude of the original signal. The normalized amplitude is then converted to a decibel scale to better represent and visualize data with a wide dynamic range. When analyzing Doppler maps, this not only provides a more intuitive representation of signal strength but also effectively highlights the relative differences between different frequency components.
[0059] Specifically, the average value C of the window function is first calculated to compensate for the amplitude loss caused by windowing, as shown in the following formula: ; Where C is the average value of the window function, w is the coefficient of the window function, wlen is the window length, and m is the coefficient number of the window function.
[0060] Then the amplitude is normalized. Here, the window function length wlen and the average value C are used to compensate for the effect of windowing, making the normalized amplitude closer to the true amplitude of the original signal, as shown in the following formula: ; Where S is the amplitude after normalization, wlen is the window length, and C is the average value of the window function. a is the real part of the complex data, and b is the imaginary part of the complex data.
[0061] Finally, the normalized amplitude is converted to a logarithmic scale, which can better represent the data with a wide dynamic range, making both small and large amplitude changes clearly visible, as shown in the following formula: ; Where B is the normalized amplitude converted to a logarithmic scale, with the unit being decibel, and S is the amplitude value after amplitude normalization.
[0062] See also Figure 11 , which is a speed-time graph drawn according to the spectrum data after normalization. It can be seen that there is a lot of noise in the feature graph, which will cause certain interference to the recognition accuracy of the subsequent neural network. Therefore, a suitable threshold should be selected to filter out the noise. The speed-time graph and the distance-time graph are respectively converted into binary images by taking appropriate thresholds. This can not only filter out the noise, but also highlight the required features of the feature graph, thereby increasing the recognition accuracy of the entire system. In an optional embodiment, 22 is selected as the threshold. Since performing logarithmic operations and square roots in FPGA is very cumbersome and has certain errors, it is chosen to mathematically simplify the mathematical expression of the binary comparison. Specifically as follows: ; After simplification, we can get: ; Finally, shift both sides right by 14 bits and compare both sides of the equal sign.
[0063] In summary, we get a speed-time graph with a size of 256x64 and a distance-time graph with a size of 64x256. The speed-time graph output after threshold processing is as follows: Figure 12 It is worth noting that the distance-time diagram is processed in the same way as the speed-time diagram.
[0064] The feature fusion module is used to perform transposition transformation and corresponding element fusion on the binarized distance-time graph and speed-time graph to generate a feature graph.
[0065] The distance-time graph provides information on the target's distance relative to the millimeter-wave radar over time, while the velocity-time graph describes the target's velocity over time. By integrating these two, the target's position and motion state can be more accurately determined, thereby improving the accuracy of classifying different behaviors. Furthermore, integrating features from multiple layers enables more generalizable models, helping to reduce overfitting of the training set and improving generalization capabilities to unknown data.
[0066] Through transposition transformation, the speed time map and the distance time map are transformed into the same size and number of channels. Then, the elements at corresponding positions of different feature maps are fused here to finally obtain the fused feature map of size 64x256 required by the next module.
[0067] The training data preprocessing module is used to obtain feature maps, divide the training and test sets of different actions according to the actions and thresholds, build the neural network model after packaging, and test to obtain weights and biases. The weights and biases are used for the hardware deployment of the subsequent neural network model.
[0068] The acquired feature map of size 64x256 is grayscale binarized and rotated 90° counterclockwise to provide a foundation for the subsequent software and hardware integration, namely the FPGA deployment module.
[0069] The radar signal preprocessing module is used to obtain feature maps of different actions. This application obtains six actions, including falling backwards, approaching the radar, moving horizontally, moving away from the radar, falling sideways, and standing still. It is worth noting that there is no restriction on the actions obtained here and they can be set according to the actual situation.
[0070] We obtained 1,280 feature maps for each action, for a total of 7,680 feature maps, which were used as the dataset. We randomly extracted 80% of the feature maps from each action dataset as the training set, and the remaining 20% as the test set for training the neural network.
[0071] Create a class named CustomData in Python to encapsulate the dataset. CustomData's functionality includes loading the dataset and calculating the mean and standard deviation of the feature map pixels. In one alternative embodiment, calculating the mean across the entire dataset yields a mean of 0.044033 and a standard deviation of 0.205168. The training and test sets are then processed sequentially in the same manner.
[0072] The neural network module is used to build neural networks, train data sets, and perform verification.
[0073] The neural network building module of this application is built on the basis of LeNet, including input layer, convolution layer 1, maximum pooling layer 1, convolution layer 2, maximum pooling layer 2, flattening layer, fully connected layer 1 and fully connected layer 2. Specifically, the parameter configuration of each layer of the neural network is shown in Table 1:
[0074] In the table, the first layer is the input layer, of type Input, with an input size of [1, 64, 256] and an output size of [1, 64, 256], where 1 is the number of samples input to the network at a time, 64 is the height of the feature map, and 256 is the width of the feature map. The second layer is Convolutional Layer 1, the third layer is Max Pooling Layer 1, the fourth layer is Convolutional Layer 2, the fifth layer is Max Pooling Layer 2, the sixth layer is Flattening Layer, the seventh layer is Fully Connected Layer 1, and the eighth layer is Fully Connected Layer 2.
[0075] The dataset was used to train a neural network model, achieving an accuracy of 98.96% on the validation set. After training, the trained neural network model was stored, along with the parameters for each component of the model (weights W and biases b).
[0076] To verify the trained neural network model, its forward propagation process is as follows: The input feature map first passes through the cov1 layer, where it undergoes convolution and ReLU activation, and then through the max pooling layer 1. The output passes through the cov2 layer, where it undergoes convolution and ReLU activation, and then again through the max pooling layer 2. After two convolutions and pooling cycles, the feature map is flattened into a one-dimensional vector. This one-dimensional vector passes through the fc1 layer, where it undergoes a linear transformation and ReLU activation. Finally, the output passes through the fc2 layer, where it undergoes a linear transformation, resulting in the final classification result.
[0077] After network verification, the neural network layout was obtained. The neural network was trained to produce the trained model.pth file. Based on the trained model.pth file, a validation set module was constructed. Six actions were taken, each with 128 feature maps, to validate the neural network model. At the software level, the final model achieved good accuracy on the validation set (96.63% in this example).
[0078] The hardware and software combination module is used to deploy the neural network model into FPGA.
[0079] See also Figure 13 , which is a module distribution diagram for the hardware implementation of the neural network model. The hardware and software combined modules include the data_ram_1 module, conv_1 module, pool_layer_1 module, data_ram_2 module, conv_2 module, pool_layer_2 module, FC1 module, FC2 module, and out_label module.
[0080] Specifically, the data_ram_1 module is used to cache feature maps.
[0081] In the data_ram_1 module, 64 RAMs with a depth of 256 are used to store the 64x256 feature map, and the corresponding rows are enabled for the corresponding RAM, such as Figure 14 As shown in the figure, after all feature maps are cached, the conv_start signal is pulled high.
[0082] The conv_1 module is used to perform convolution operations on the feature map to obtain the convolved feature map.
[0083] The conv_1 module is the first convolution layer of the convolutional neural network. It contains 8 3x3 convolution kernels with a stride value of 1. Therefore, if a 64x256 feature map is input, 8 62x254 feature maps will be output. In the conv_1 module, the weight parameters in the convolution kernel are amplified by 8192 times and rounded to facilitate subsequent multiplication and addition operations. Among them, the first row of weights of the 8 convolution kernels are placed in the first ROM 0, the second row is placed in ROM 1, and the third row is placed in ROM 2. In addition, 8 biases need to be placed in ROM_b. Therefore, this module requires a total of 3 ROM kernels with a depth of 24 and 1 ROM kernel with a depth of 8. The ROM diagram is as follows: Figure 15 As shown, it is finally subjected to sliding window calculation with the cached feature map and finally output to the pool_layer_1 module.
[0084] The pool_layer_1 module is used to perform pooling operation on the convolution feature map to obtain the pooled feature map.
[0085] The pool_layer_1 module is the first pooling layer. It performs maximum pooling on the 8 62x254 feature maps output by the conv_1 module and outputs 8 31x127 feature maps. It compares the data of the even rows (starting from 0) in pairs and stores the larger number in the FIFO. Then, it reads the FIFO at the same time as the odd row data comes in to obtain the maximum pooling value. Therefore, a FIFO with a depth of 127 is required in the pool_layer_1 module. The schematic diagram of the pool_layer_1 module is as follows: Figure 16 shown.
[0086] The data_ram_2 module is used to cache the pooled feature maps, and the next convolution begins after all feature maps are cached.
[0087] In an optional embodiment, the data_ram_2 module is used to cache data between the pool_layer_1 module and the conv2 module. After caching eight 31x127 feature maps, the conv2_start signal is pulled high to start the second layer of convolution. This implementation is similar to the data_ram_1 module and will not be repeated here.
[0088] The conv_2 module is used to perform a secondary convolution operation on the pooled feature map to obtain the secondary convolution feature map.
[0089] The conv_2 module is the second convolution layer, which contains 8 3x3 convolution kernels with a step value of 1, and outputs 8 29x125 feature maps. The implementation of the conv_2 module fully reflects the parallelism of the FPGA. In the conv_2 module, the 8 feature maps cached in the data_ram_2 module are simultaneously subjected to the same convolution operation as the conv_1 module. That is, this application can use the time consumed by the traditional neural network to convolve one feature map to complete the convolution of 8 feature maps. The schematic diagram of the conv_2 module is as follows Figure 17 shown.
[0090] The pool_layer_2 module performs a secondary pooling operation on the feature map after the secondary convolution to obtain the feature map after the secondary pooling.
[0091] The pool_layer_2 module is the second pooling layer. It performs maximum pooling on the 8 29x125 feature maps output by the conv_2 module (after being reduced by 8192 times). Note that the last row and the last column of the feature map are cropped to obtain 8 28x124 feature maps. Finally, 8 14x62 feature maps are output. Its implementation is similar to that of the pool_layer_1 module and will not be repeated here.
[0092] The FC1 module is used to flatten the feature map after secondary pooling and perform multiplication and accumulation to obtain 100 nodes.
[0093] The FC1 module is the first fully connected layer. It flattens the 8 14x62 feature maps output by the pool_layer_2 module, multiplies them with 100 sets of coefficients, and then outputs 100 nodes. First, use 8 RAMs with a depth of 868 to cache the 8 feature maps of the upper module. The schematic diagram is as follows Figure 18 As shown, then 8 ROMs with a depth of 86800 are instantiated to store the corresponding weight coefficients, and finally 1 ROM with a depth of 100 is instantiated to store the bias. The schematic diagram is as follows Figure 18 As shown. Among them, the weight parameters of the FC1 module are arranged as follows Figure 19 shown.
[0094] The FC2 module is used to multiply and accumulate 100 nodes to obtain 6 nodes.
[0095] The FC2 module is the second fully connected layer, which inputs 100 nodes and outputs 6 nodes (magnified 8192 times). Its implementation is similar to that of the FC1 module and will not be repeated here.
[0096] The out_label module is used to output the identification label, that is, to output the maximum value among the six nodes of the FC2 module, and output the human action label corresponding to the maximum value to the display module for display.
[0097] The display module includes an OLED display, a voice module, and a wireless transmission module. The OLED display module and the voice module drive the OLED screen to display the action name of the corresponding tag, while the speaker can broadcast the action name of the corresponding tag.
[0098] In an optional embodiment, the present application is also provided with a wireless sending module, which can send the tag of the identification action to the ESP_01S module through the serial port. The module can communicate with the server and realize identification communication between multiple devices, which is convenient for developing abnormal action alarm functions or somatosensory game action interaction functions on the mobile phone.
[0099] Example 1: One-time simulation and functional verification The interface module receives and unpacks data packets from the DCA1000 data acquisition card and outputs specific sampling points (here the 32nd sampling value of each chirp). The sampling points are represented by 32 bits, with the upper 16 bits representing the real part and the lower 16 bits representing the imaginary part. The following simulation uses two data packets, as shown in the simulation diagram. Figure 20 As shown. Figure 20 Compare with the 32nd sampling value of the original radar echo data, such as Figure 21 As shown in the figure, the interface module functions normally.
[0100] Compare the output data of the radar signal preprocessing module with the data output by the signal processing algorithm. The binary data finally output by the radar signal preprocessing module is as follows: Figure 22 As shown in the figure, the first set of 256 data is used for comparison. The data output by the signal processing algorithm is as follows Figure 23 As shown in the figure, the functions of each module in the radar signal preprocessing module are normal. All modules in the neural network recognition module will be simulated and verified, and the final output simulation diagram is as follows: Figure 24 As shown in the figure, the label output by the neural network recognition model is 5, and the specific values of each node are as follows Figure 25 As shown. Figure 25 It can be seen that its maximum value corresponds to the 5th node, that is, the recognition result is label 5, so the functions of the neural network recognition model are normal.
[0101] Example 2: Measured Verification In an offline state, the entire system was placed on a 1-meter-high table, and a person began to perform a gesture 2 meters in front of it. The system controlled the data acquisition card in the millimeter-wave radar to continuously collect human motion data in real time. The radar signal preprocessing module converted the data into a feature map, which was then input into the neural network recognition module for recognition. The action label 1 was identified, which in this experiment corresponded to falling backward. Simultaneously, the LED light group began to flash, and the OLED display displayed the action name and voice announcement. After multiple user changes and multiple tests, the system achieved an overall recognition accuracy of 98.3% for the six actions in this experiment (falling backward, approaching the radar, moving horizontally, moving away from the radar, falling sideways, and standing still). This achieved an offline state by directly controlling the data acquisition card to obtain raw radar echo data using an FPGA to simulate a PC, resulting in enhanced real-time performance. Multiple tests achieved an accuracy rate of up to 98.3%, a higher accuracy rate. Furthermore, directly transmitting data to the FPGA via Ethernet reduces packet loss or corruption due to line interference, improving robustness.
[0102] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0103] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A real-time human behavior recognition system based on FPGA and millimeter-wave radar, including millimeter-wave radar, characterized in that: Also included are an electrically connected radar signal pre-processing module and a neural network recognition module; The radar signal preprocessing module is used to enable the FPGA to directly connect to and control the data acquisition card in the millimeter wave radar to collect radar raw echo data, extract specific sampling points, and compress each frame of chirp signal to obtain a one-dimensional vector. The specific sampling points and the one-dimensional vector are windowed to obtain windowed data, and FFT is performed to obtain FFT spectrum data. After normalization and conversion, binarized distance time graph and velocity time graph are obtained, and the feature graph is generated by fusion. The neural network recognition module is used to store the parameters obtained by training the feature map into the neural network model and deploy it to the FPGA to achieve offline recognition; The FPGA is used to implement data encapsulation and transmission in Ethernet communication.
2. The real-time human behavior recognition system based on FPGA and millimeter-wave radar according to claim 1 is characterized in that: The radar signal preprocessing module includes an interface module; The interface module is used to enable the FPGA to directly control the data acquisition card to collect the human motion data, splice and obtain the complete radar original echo data, and at the same time, trigger the next time after the data acquisition card completes one acquisition, so as to capture the human motion data without interruption; The interface module includes a conversion module, a sending module and a receiving module; The conversion module is used to convert the GMII and RGMII of the Ethernet interface into each other; The sending module is used to implement sending a preparation command to the data acquisition card on the FPGA; The receiving module is used to receive the preparation completion signal, the radar original echo data and the sending completion signal sent by the data acquisition card.
3. The real-time human behavior recognition system based on FPGA and millimeter-wave radar according to claim 2 is characterized in that: The radar signal preprocessing module also includes a cache module, a sampling module, a range compression module, a windowing module, an STFT module, a normalization module, a threshold processing module and a feature fusion module; The cache module is used to align data of the interface module across clock domains; The sampling module is used to extract the specific sampling point of each chirp in the radar raw echo data; The distance compression module is used to add each frame of chirp in the slow time dimension and then average it to obtain the one-dimensional vector.
4. The real-time human behavior recognition system based on FPGA and millimeter-wave radar according to claim 3 is characterized in that: The windowing module is used to store the window function data through a ROM core, and the reading of the ROM data and the input of the one-dimensional vector and the specific sampling point are controlled by the input enable signal of the one-dimensional vector and the specific sampling point. The one-dimensional vector and the specific sampling point are divided into a real part and an imaginary part, and the real part and the imaginary part are respectively multiplied with the window function data in the multiplier IP core. The output result is then spliced with the real part and the imaginary part to obtain the windowed data.
5. The real-time human behavior recognition system based on FPGA and millimeter-wave radar according to claim 4 is characterized in that: The STFT module is used to perform FFT on each frame of the windowed data to obtain the frequency spectrum data after FFT.
6. The real-time human behavior recognition system based on FPGA and millimeter-wave radar according to claim 5, characterized in that: The normalization module compensates the amplitude of the spectrum data after the FFT by dividing by the length and average value of the window function data; The threshold processing module is used to convert the amplitude into a decibel scale, and convert the speed time graph and the distance time graph drawn by the spectrum data into the binarized distance time graph and speed time graph according to the threshold.
7. The real-time human behavior recognition system based on FPGA and millimeter-wave radar according to claim 6, characterized in that: The feature fusion module is used to perform transposition transformation and corresponding element fusion on the binarized distance-time graph and speed-time graph to generate a feature graph.
8. The real-time human behavior recognition system based on FPGA and millimeter-wave radar according to claim 7, characterized in that: The neural network recognition module includes a neural network module and a software and hardware combination module; The neural network module is used to store the weights and biases obtained after the feature map training into the neural network model; The software and hardware combination module is used to deploy the neural network model into the FPGA.
9. The real-time human behavior recognition system based on FPGA and millimeter-wave radar according to claim 8, characterized in that: The training preprocessing module is used to obtain the feature map, obtain training sets and test sets for different actions according to the action and threshold, build a neural network model after packaging, and test to obtain the weights and biases.
10. The real-time human behavior recognition system based on FPGA and millimeter-wave radar according to claim 9, characterized in that: The software and hardware combination module includes a data_ram_1 module, a conv_1 module, a pool_layer_1 module, a data_ram_2 module, a conv_2 module, a pool_layer_2 module, an FC1 module, an FC2 module, and an out_label module, which is used to deploy the neural network model into the FPGA; The data_ram_1 module is used to cache the feature map; The conv_1 module is used to perform a convolution operation on the feature map to obtain a convolved feature map; The pool_layer_1 module is used to perform a pooling operation on the convolved feature map to obtain a pooled feature map; The data_ram_2 module is used to cache the pooled feature map, and start the next convolution when the pooled feature map reaches the threshold; The conv_2 module is used to perform a secondary convolution operation on the pooled feature map to obtain a secondary convolution feature map; The pool_layer_2 module performs a secondary pooling operation on the feature map after the secondary convolution to obtain a secondary pooled feature map; The FC1 module is used to flatten the feature map after the secondary pooling and perform multiplication and accumulation to obtain 100 nodes; The FC2 module is used to multiply and accumulate 100 nodes to obtain 6 nodes; The out_label module is used to output the maximum value among the 6 nodes, and output the human action label corresponding to the maximum value to the display module for display.
Citation Information
Patent Citations
Fall detection method and device based on millimeter wave radar and FPGA
CN115204240A
Indoor human body posture recognition method based on millimeter wave radar
CN116561700A
Portable millimeter wave detection system and method
CN118033630A
Cited By
Lightweight millimeter wave radar two-dimensional feature map classification method and system oriented to FPGA hardware deployment
CN121637200A