Target detection method, integrated circuit, sensor, device, and medium
The target detection method using FFT processing and machine learning improves accuracy and efficiency in short-distance spaces by addressing low detection rates and false alarms, enabling precise detection of special targets like infants.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2026-04-14
AI Technical Summary
In short-distance enclosed spaces like car cabins, target detection faces challenges such as low detection rates, many false alarms, and inaccurate angle estimation due to strong static clutter and multipath interference, especially for low RCS targets like children, and the complexity of device placement and rule-based judgment.
A target detection method involving FFT processing of echo signals, followed by machine learning models, and a system comprising a radio frequency module, analog signal processing, and digital signal processing module, with an integrated circuit and electromagnetic wave sensor for improved target detection and communication.
Enhances target detection accuracy and efficiency by reducing false alarms and improving angle estimation, enabling precise detection of special targets like infants and weak targets, supporting applications like Child Presence Detection and Safety Belt Reminder.
Smart Images

Figure 0007844770000042 
Figure 0007844770000043 
Figure 0007844770000044
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of target detection, and particularly to a target detection method, an integrated circuit, an electromagnetic wave sensor, a device, and a computer-readable storage medium.
Background Art
[0002] For example, when performing target detection in a short-distance enclosed or relatively enclosed space area, such as target detection inside a car cabin, due to the influence of strong static clutter and multipath, there are technical problems such as low detection rate, many false alarm targets, and inaccurate angle estimation. Especially in the scenario where a child is left unattended in the car, the child's RCS (Radar Cross Section) is low, making the detection difficulty even higher. In addition, the judgment of the seat area based on predetermined rules often depends on the measurement of spatial geometric relationships, appropriate rule design, and the adjustment and optimization of various parameters, resulting in high difficulty in device placement and test driving.
Summary of the Invention
[0003] The embodiments of the present application provide a target detection method, an integrated circuit, an electromagnetic wave sensor, a device, and a computer-readable storage medium that can improve the accuracy and efficiency of target detection.
[0004] According to some embodiments of the present application, in the first aspect of the embodiments of the present application, performing FFT processing on the echo signal so as to generate a feature vector of the echo signal, and processing the feature vector based on a machine learning model so as to realize target detection in a region of interest. A target detection method is provided that includes these steps.
[0005] According to some embodiments of the present application, in the second aspect of the embodiments of the present application, it includes a sequentially connected radio frequency module, an analog signal processing module, and a digital signal processing module. The radio frequency module is used to generate a radio frequency transmission signal and receive a radio frequency reception signal. The analog signal processing moduleRadio frequency received signal From intermediate frequency signal Get The digital signal processing module is used for analog-to-digital conversion of an intermediate frequency signal, and based on this, an integrated circuit is provided that realizes the target detection method described in the first embodiment of the present invention.
[0006] According to some embodiments of the present application, a third embodiment of the present application further provides a support, an integrated circuit as described in the second embodiment of the present application provided on the support, and an antenna provided on the support or integrated with the integrated circuit as an integrated device provided on the support, wherein the integrated circuit is further provided as an electromagnetic wave sensor connected to the antenna for transmitting radio frequency transmission signals and / or receiving radio frequency reception signals.
[0007] According to some embodiments of the present application, a fourth embodiment of the present application further provides a device comprising a main body and an electromagnetic wave sensor as described in the third embodiment of the present application, wherein the electromagnetic wave sensor is used for target detection and / or communication to provide reference information for the operation of the main body of the application. [Brief explanation of the drawing]
[0008] [Figure 1] This is a flowchart of the target detection method according to the embodiment of the present invention. [Figure 2] This is a schematic diagram of the flow of the target detection method according to an embodiment of the present invention. [Figure 3] This is a flowchart of the target detection method according to the embodiment of the present invention. [Figure 4] This is a flowchart of the target detection method according to the embodiment of the present invention. [Figure 5] This is a flowchart of the target detection method according to the embodiment of the present invention. [Figure 6] This is a schematic diagram of the division of the region of interest for the target detection method according to an embodiment of the present application. [Figure 7] This is a schematic diagram of the data rearrangement for the target detection method according to an embodiment of the present invention. [Figure 8] This is a schematic diagram of the structure of a deep learning model for a target detection method according to an embodiment of the present invention. [Figure 9] This is a flowchart of the target detection method according to the embodiment of the present invention. [Figure 10] This is a schematic diagram of the flow of the target detection method according to an embodiment of the present invention. [Figure 11] This is a schematic diagram of the flow of the target detection method according to an embodiment of the present invention. [Figure 12] This is a schematic diagram of the flow of the target detection method according to an embodiment of the present invention. [Figure 13] This is a schematic diagram of a target detection method according to another embodiment of the present invention. [Figure 14] This is a schematic diagram of a target detection method according to another embodiment of the present invention. [Figure 15] This is a schematic diagram of a target detection method according to another embodiment of the present invention. [Figure 16] This is a schematic diagram of a target detection method according to another embodiment of the present invention. [Figure 17] This is a schematic diagram of a target detection method according to another embodiment of the present invention. [Figure 18] This is a schematic diagram of a target detection method according to another embodiment of the present invention. [Figure 19] This is a schematic diagram of a target detection method according to another embodiment of the present invention. [Figure 20] This is a schematic diagram of a target detection method according to another embodiment of the present invention. [Figure 21] This is a schematic diagram of a target detection method according to another embodiment of the present invention. [Figure 22] This is a schematic diagram of a target detection method according to another embodiment of the present invention. [Figure 23] This is a schematic diagram of the data rearrangement for the target detection method according to an embodiment of the present invention. [Figure 24] This is a flowchart of the target detection method according to the embodiment of the present invention. [Figure 25] This is a flowchart of the target detection method according to the embodiment of the present invention. [Figure 26A] This is a schematic diagram illustrating the effects of the target detection method according to the embodiment of the present invention. [Figure 26B] It is a schematic diagram of the effect of the target detection method according to an embodiment of the present application. [Figure 26C] It is a schematic diagram of the effect of the target detection method according to an embodiment of the present application. [Figure 26D] It is a schematic diagram of the effect of the target detection method according to an embodiment of the present application.
Mode for Carrying Out the Invention
[0009] Although multiple embodiments are described in the present disclosure, the description is illustrative and not restrictive. It is obvious to those skilled in the art that within the scope included in the embodiments described in the present disclosure, more embodiments and implementation means may exist. Many possible combinations of features are shown in the drawings and discussed in specific embodiments, but many other combinations of the disclosed features are also possible. Unless explicitly restricted, any feature or element of any embodiment can be used in combination with, or replace, any other feature or element of any other embodiment.
[0010] The present disclosure includes and contemplates combinations of features and elements known to those skilled in the art. The embodiments, features, and elements already disclosed in the present disclosure can be combined with any conventional feature or element so as to form a unique inventive means defined by the claims. Any feature or element of any embodiment can be combined with features or elements from other inventive means so as to form another unique inventive means defined by the claims. Therefore, it should be understood that any feature shown and / or discussed in the present disclosure can be realized individually or in any suitable combination. Therefore, the embodiments are not subject to other restrictions except those based on the limitations of the appended claims and their equivalent replacements. Also, various modifications and changes can be made within the protection scope of the appended claims.
[0011] Furthermore, when describing representative embodiments, the specification may already present a method and / or process as a specific sequence of steps. However, to the extent that such method or process does not depend on a specific order of steps described herein, the method or process should not be limited to the steps in the specific order described. As those skilled in the art will understand, other orders of steps are also possible. Therefore, a specific order of steps described in the specification should not be construed as limiting the claims. Moreover, claims for such method and / or process are not limited to the steps performed in the order they are written. Those skilled in the art will readily understand that these orders are modifiable and still fall within the spirit and scope of the embodiments of this application.
[0012] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as those generally understood by those skilled in the art. Terms used in this disclosure are for illustrative purposes only and do not limit the disclosure.
[0013] In some arbitrary embodiments, when target detection is performed in a sealed or relatively sealed spatial region, target detection can be achieved by performing region determination based on high-speed and low-speed time processing of single frames, CFAR (constant false alarm) detection, or point clouds obtained from angle measurements. Alternatively, after high-speed time processing of single frames, operations such as DoA (direction of arrival estimation) can be performed using algorithms such as Capon to form an RA (range-azimuth) diagram, after which features such as outputs within a predetermined region can be extracted, recognized, and judged to achieve target detection. Alternatively, target detection can be achieved by filtering the received echo phase and then detecting human respiration and heartbeat. For example, when target detection is achieved by performing region determination based on high-speed and low-speed time processing of single frames, CFAR detection, or point clouds obtained from angle measurements, 1D-FFT (range-dimensional FFT, etc.) processing can be performed on the echo signal, followed by time correlation analysis based on the echo data to ultimately achieve human target discrimination. Subsequently, by combining point cloud information and other data to determine the position, the objective of target detection within the cabin can finally be achieved.
[0014] In some embodiments, the target detection flow is shown in Figure 1. The radar (taking frequency-modulated continuous wave (FMCW) radar target detection as an example) processes the echo signal received by the antenna by performing operations such as analog-to-digital conversion and sampling. Then, a main control module (which can be implemented based on, for example, a microcontroller unit (MCU)) and a baseband (BB) module (which can be implemented based on, for example, a baseband chip) provide further processing to realize target detection. The baseband module sequentially performs range-dimensional DC removal and range-dimensional Fourier transform (which can also be expressed as Range FFT, 1D-FFT, or fast time-dimensional FFT). Next, the baseband module transmits a signal to the main control module, which then performs multi-frame storage. After the data has been accumulated in a predetermined number of frames, such as 128 frames, the main control module moves the multi-frame accumulated data from the CPU Static Random-Access Memory (SRAM) to the baseband module. Based on the baseband module, it then performs Doppler DC removal, inter-frame Fourier transform (Frame FFT), and inter-channel storage (non-coherent storage, correlated storage, etc.). Finally, the baseband module performs constant false-alarm rate (CFAR), such as peak-value-based detection with peak selection (CFAR), based on noise obtained by noise estimation (using a variance-based noise estimator, for example). Finally, it performs angle detection and target classification to obtain target range, speed, and / or angle information.
[0015] Here, the relevant steps in Figure 1 can be implemented as follows.
[0016] Range-dimensional DC rejection targets the data corresponding to each chirp in the data of each frame. Specifically, range-dimensional DC rejection involves calculating the average along the fast time dimension for the data collected by each receiving (RX) channel for each chirp, and then subtracting its DC component from all sampling points of each RX channel.
[0017] Range-dimensional FFT can be achieved by windowing the data from which the range-dimensional DC has been removed and then performing a 1D-FFT.
[0018] In the target detection process, the data for each chirp in each frame obtained by the radar processing the echo signal using ADC or sampling is manipulated as described above and transmitted to the main control module. Here, the transmission of data to the main control module can be achieved by moving the data from Direct Memory Access (DMA) to the CPU SRAM.
[0019] Doppler DC rejection can be achieved by calculating a complex number average along the frame dimension for multi-frame data stored in the main control module (i.e., averaging data from different frames in the same range unit (bin) of the same channel) so that the DC component of the corresponding range unit of the corresponding channel is the average value of different ranges of different channels, and then subtracting the respective DC components from each range unit of each channel.
[0020] Interframe Fourier transform can be performed by windowing the data, from which Doppler DC has been removed, along the frame dimension, and then performing a 2D-FFT along the frame dimension. In this case, range Doppler information for each transmit and receive channel can be obtained. For example, in some embodiments, this can be achieved as follows: After the echo signal is processed by range-dimensional FFT to obtain a range Doppler spectrum, the results obtained from the range-dimensional FFT processing are slide-windowed, and a 2D FFT is performed on the windowed data. Here, the length of the sliding window and the period of the target's periodic motion have the same time length. By performing digital signal processing on multiframe data that is close to the period of the target's periodic motion in this way, target detection can be performed more appropriately in combination with the characteristics of the target's motion period, helping to avoid errors and omissions while simultaneously eliminating the effects of multipath and other factors, thereby improving the accuracy of target detection.
[0021] Non-coherent storage can be achieved by the following equation:
[0022]
number
number
[0023] However, P(r,v) and A(r,v) are the output and amplitude of the r-th range unit and the v-th Doppler unit, respectively. s(r,t,a,v) is the complex value obtained by the inter-frame Fourier transform of the signal echoes of the r-th range unit and the v-th Doppler unit in the t-th transmit channel and the a-th receive channel.
[0024] The detection of constant false alarms based on the noise floor obtained through noise estimation is described in conventional CFAR technology and will not be explained here.
[0025] Angle detection can be achieved by performing azimuthal-dimensional digital beamforming (DBF) and DOA, as well as elevation-dimensional DBF and DOA, on target points obtained by CFAR. Here, DBF and DOA are described in conventional DBF and DOA techniques and will not be explained here.
[0026] Target classification can be performed by counting target points detected in each frame (or a predetermined amount of data, i.e., each time) according to a predetermined (or divided) region unit design, and by determining the counted results according to predetermined rules. This allows for obtaining whether or not a physical target of interest (such as an adult, child, infant, or pet) is present in the process and identifying its location region.
[0027] Note that Figure 1 is merely an example for radars equipped with a baseband module and a main control module. Specifically, the main control module performs operations such as frame data storage and target classification (which can be achieved by Region HIST, classification, etc.), while the baseband module performs other operations such as range-dimensional DC rejection and Doppler-dimensional DC rejection. In other embodiments, frame data storage can be performed in the baseband module, and Doppler DC rejection and other operations can be performed in the main control module.
[0028] Note that while Figure 1 only illustrates caching 128 frames of data as an example, other embodiments can also cache 56 frames or 32 frames. The embodiments of this application are not limited to these and can be specified according to demand and hardware support capabilities. As a result, the cached multi-frame data is subjected to Doppler DC rejection and inter-frame FFT to obtain a Range-Doppler (RD) spectrum.
[0029] Note that in the flow shown in Figure 1, some steps such as range-dimensional DC rejection, windowing, etc., may be skipped, or some steps may be simplified. For example, non-coherent storage can be skipped and CFAR detection can be performed only on specific channels. Alternatively, for example, non-coherent storage can be performed on only some channels, and CFAR detection can be performed only on those channels.
[0030] Furthermore, based on the above means, the embodiments of this application further propose processing by combining it with a deep learning model. By deep mining target information using a deep learning model, it is possible to achieve more accurate target detection in spatial areas such as car cabins, indoors, and factory buildings, effectively improving the detection rate, reducing the number of false alarm targets, and simultaneously significantly improving the accuracy of angle estimation. As a result, special targets such as infants in the cabin or weak targets can be accurately detected, enabling applications such as CPD (Child Presence Detection) and SBR (Safety Belt Reminder). The deep learning model can be combined with the above embodiments in different forms, and these different combinations are described below.
[0031] In some embodiments, as shown in Figure 2, the target detection method includes the following steps.
[0032] In step 101, FFT processing and beamforming processing are performed on the echo signal so that a feature vector of the echo signal is generated. Here, the feature vector includes characterizing the energy features of the echo signal in different range units.
[0033] In step 102, feature vectors are processed based on a deep learning model so that target detection in the region of interest is achieved.
[0034] Thus, beamforming effectively improves the signal-to-noise ratio between personnel targets and static strong clutter, effectively constructing feature vectors for empty vehicles (with or without interference), adult presence scenes, and child presence scenes. By deep mining the target information contained in the feature vectors using a deep learning model, more accurate target detection is achieved, realizing a means of target detection in spatial domains such as car cabins, indoors, and factory buildings. This effectively improves the detection rate, reduces the number of false alarm targets, and significantly improves the accuracy of angle estimation. As a result, special targets such as infants in cabins and weak targets can be accurately detected, enabling the application of technologies such as CPD and SBR. Simultaneously, the introduction of a deep learning module further reduces the necessary digital signal processing, for example, eliminating the need for processing such as CFAR or DOA after feature extraction.
[0035] To better understand the embodiment shown in Figure 2, the steps are described below.
[0036] Step 101 does not limit the implementation of FFT processing and beamforming processing. Since beamforming processing essentially concentrates the directional gain (the ability to respond to energy) of a signal in a specific direction, beamforming is effectively equivalent to extracting energy features. In other words, different implementation methods of FFT processing and beamforming processing do not affect the extraction of energy features, and therefore do not affect the generation of feature vectors, including characterizing the energy features of echo signals in different range units.
[0037] As can be seen from the flow shown in Figure 1, the target detection process requires beamforming to enable angle estimation, and subsequent processing such as target classification to output intuitive and accurate results. In the embodiment shown in Figure 2, beamforming is used to generate feature vectors, and subsequent angle estimation and target classification are replaced by a deep learning model, resulting in higher efficiency. Furthermore, instead of directly using time-series data as input, the deep learning model extracts target information from the echo signal through FFT and beamforming processing, forming a feature vector with higher dimensions and richer information content. This allows the deep learning model to perform analysis and processing better, output more accurate results, improve the accuracy and efficiency of target detection, and enable more accurate target detection and recognition within the cabin. Here, noise and target reflection signals have different energy characteristics, and using feature vectors that characterize these energy characteristics helps the deep learning model perform better target detection and output more accurate and reliable results.
[0038] The application of deep learning models involves replacing processes such as CFAR and angle estimation, and directly outputting detection results. That is, based on the flow shown in Figure 1, beamforming is performed after interframe FFT, and processes such as correlation accumulation, CFAR, and angle estimation are replaced by the deep learning model's processing. In other words, in some embodiments, as shown in Figure 3, performing FFT processing and beamforming processing on the echo signal can be achieved by the following steps.
[0039] In step 1011, the echo signal is subjected to FFT processing so that a range Doppler spectrum is generated.
[0040] In step 1012, beamforming is performed based on the range Doppler spectrum.
[0041] As shown in Figure 1, the FFT processing includes range-dimensional FFT and interframe FFT. After the range-dimensional FFT, beamforming processing can be supported without waiting for the interframe FFT. Also, as mentioned above, the order of implementation of beamforming processing differs, energy features can still be obtained, and feature vectors can still be generated. Therefore, in some embodiments, as shown in Figure 4, the FFT processing and beamforming processing of the echo signal can be realized by the following steps.
[0042] In step 1013, the echo signal is subjected to range FFT processing.
[0043] In step 1014, beamforming is performed on the results obtained by range FFT processing.
[0044] In step 1015, the beamforming results are converted into a sliding window, and the windowed data is subjected to a 2D FFT. Optionally, a sliding window 2D-FFT may be used here.
[0045] Furthermore, beamforming is already supported after range-dimensional FFT processing, and the subsequent inter-frame FFT is performed to improve target detection through digital signal processing. Also, because deep learning models have the ability to learn hidden features, good target detection results can still be obtained even without inter-frame FFT, and good effects can still be achieved even without inter-frame FFT. Based on this, in some embodiments, as shown in Figure 5, FFT processing and beamforming processing on the echo signal can be achieved by the following steps.
[0046] In step 1016, the echo signal is subjected to range FFT processing.
[0047] In step 1017, beamforming is performed on the results obtained by range FFT processing.
[0048] Note that the related processes performed in the embodiments shown in Figures 3 to 5 (including range FFT, frame data storage, inter-frame FFT, etc.) are described in the embodiment shown in Figure 2, and will not be explained here.
[0049] Of course, the above is merely an illustrative explanation of the implementation methods for FFT processing and beamforming processing, and in some embodiments, processing can be carried out by other appropriate methods, but these will not be listed here.
[0050] In the embodiments shown in Figures 4 and 5, beamforming is performed based on the results of range-dimensional FFT processing, whereas in the embodiment shown in Figure 3, beamforming is performed based on the results of inter-frame FFT processing. Therefore, the implementation methods for beamforming are also different. Specifically, in the embodiments shown in Figures 4 and 5, beamforming is implemented in the range dimension, that is, forming and accumulation are performed in different range units, whereas in the embodiment shown in Figure 3, beamforming may be implemented in both the range dimension and the Doppler dimension, that is, forming and accumulation are performed in different Doppler units of different range units. To make it easier to understand, the implementation of beamforming in the embodiment shown in Figure 3 will be explained below as an example.
[0051] In some embodiments, beamforming (for example, performing Nc-point beamforming) based on the range Doppler spectrum obtained by interframe FFT can be realized by the following equation.
[0052]
number
[0053] However, P(r,v,b) is the energy characteristic of the r-th range unit and v-th Doppler unit based on beam b. sv(c,b) is the steering vector of beam b in virtual channel c. x(c,r,v) is the data corresponding to the r-th range unit and v-th Doppler unit in virtual channel c in the range-Doppler spectrum. Nc is the total number of beams.
[0054] In response to this, beamforming based on data obtained by range-dimensional FFT replaces x(c,r,v) in the above equation with x(c,r), i.e., it does not have a Doppler dimension. Of course, in other embodiments, this can be achieved by other methods, but we will omit the explanation here.
[0055] In the embodiments of this application, the steering vector of the beam used for beamforming is not limited and may be specified by the user or automatically generated according to the region of interest. For example, in some embodiments, beamforming can be achieved by identifying a target direction pointing to the region of interest, generating a steering vector based on the target direction, and performing beamforming based on the steering vector. Here, the identification of the target direction is also not limited to the embodiments of this application. For example, in some examples, identifying a target direction pointing to the region of interest includes identifying the target direction based on the azimuth and elevation angles of the region of interest relative to the radar. In some other examples, identifying a target direction pointing to the region of interest includes acquiring measured target data in the region of interest and performing clustering based on the measured target data, thereby setting the azimuth and elevation angles of the clustering center points in each region of interest as the target direction.
[0056] For example, considering the interior space of a vehicle shown in Figure 6, if we consider the need to determine whether or not there are people in the six areas in total—the three rear seats and the aisles in front of those seats—and if there are people, we can determine whether they are adults or children. By performing a 6-point DBF, we can identify six target directions pointing to the three seats and three aisles shown in Figure 6. In this case, we will obtain P(r,v,b) corresponding to the six beams. Of course, we can also add other directions, such as three auxiliary beams pointing to the other three areas. In this case, we perform a 9-point DBF to obtain nine beams.
[0057] The steering vector can be generated by the following equation.
[0058]
number
[0059] However, sv(c,b) is the steering vector of beam b in virtual channel c. λ is the center wavelength of the radar detection signal. xc and zc are the positions corresponding to the horizontal and vertical directions of the antenna in virtual channel c, respectively. θb is the azimuth angle corresponding to the target direction. φb is the elevation angle corresponding to the target direction.
[0060] Furthermore, since the radar coverage area is usually larger than the range required by the user for target detection, target detection typically occurs in the user's region of interest. Based on this, the embodiments of this application will mainly describe target detection in the region of interest as an example, but this does not mean that the target detection method according to the embodiments of this application can only target the region of interest; for example, it can target all detectable areas, and such details will be omitted here and thereafter.
[0061] Furthermore, the input to the deep learning model also affects the accuracy of the deep learning model. It is important to consider that the target information in the echo signal also has characteristics in a specific order. Therefore, to allow the deep learning model to process the feature vectors better, the energy features of the feature vectors can be further arranged. Based on this, in some embodiments, generating the feature vectors of the echo signal can be achieved by rearranging the energy features obtained by FFT and beamforming processing of the echo signal according to predetermined rules, so that the feature vectors are obtained. In this way, rearrangement strengthens the correlation of the feature vector data, which helps the model to detect better.
[0062] In the embodiments of this application, the specific rules for rearrangement are not limited. For example, in some embodiments, a predetermined rule includes continuously arranging data with the same first dimension according to the gradient direction of the second dimension, and arranging the data as a whole continuously according to the gradient direction of the first dimension. Here, the first and second dimensions are different dimensions in the beam dimension and range dimension, respectively. In this case, it is possible to accommodate situations in which only energy features in the range dimension occur. Also, in some embodiments, a predetermined rule includes continuously arranging data with the same third dimension and the same fourth dimension according to the gradient direction of the fifth dimension, continuously arranging data with the same third dimension according to the gradient direction of the fourth dimension, and arranging the data as a whole continuously according to the gradient direction of the third dimension. Here, the third, fourth, and fifth dimensions are different dimensions in the beam dimension, range dimension, and Doppler dimension, respectively. In this case, it is possible to accommodate situations in which energy features in the range dimension and Doppler dimension occur.
[0063] To facilitate understanding, the predetermined rules for rearrangement will be explained in conjunction with the rearranged feature vectors shown in Figure 7.
[0064] As shown in Figure 7, the energy features are arranged sequentially in physical memory in the order of v-dimensional, r-dimensional, and b-dimensional. In the figure, v(0) and v(Nv) represent the 0th and Nvth Doppler units (Doppler bins) (totaling Nv+1 Doppler bins), r(0) and r(Nr) represent the 0th and Nrth range units (range bins) (totaling Nr+1 range bins), and b0 represents the 0th beam, totaling N beams. This array of (Nv+1)×(Nr+1)×N energy features is the feature vector input to the deep learning model.
[0065] As mentioned above, since target detection is implemented based on hardware, it can be adapted to the hardware's capabilities by performing further data processing to match the hardware's capabilities.
[0066] For example, in some embodiments, generating feature vectors of the echo signal can be achieved by logarithmically normalizing the energy features obtained by FFT processing and beamforming processing of the echo signal. Therefore, logarithmic normalization helps to convert the data from floating-point numbers to fixed-point numbers, improving the efficiency of subsequent processing and, consequently, the efficiency of target detection.
[0067] For example, in some embodiments, logarithmically normalizing the energy features obtained by FFT and beamforming of the echo signal can be achieved by obtaining a logarithm of the energy features obtained by FFT and beamforming of the echo signal, and then linearly normalizing the result obtained by the logarithm under predetermined upper and lower bound constraints. That is, by further introducing upper and lower bounds and linear parameters based on logarithmic normalization, the data format is further optimized and more useful in forming characteristic data for subsequent use.
[0068] Here, obtaining a logarithm of the energy features obtained by FFT processing and beamforming processing on the echo signal, and then linearly normalizing the obtained logarithm by predetermined upper and lower bound constraints, is achieved by the following equation.
[0069]
number
[0070] However, Pnorm(r,v,b) is the result of log-normalization. P(r,v,b) is the energy characteristic of beam b in range unit r and Doppler unit v. a0, b0, Pmax, and Pmin are all predetermined parameters.
[0071] In the embodiments of this application, the specific values of the above-mentioned parameters are not limited. In some possible embodiments, b0 can be set to the mean or median of all possible log2(P(r,v,b)). Alternatively, in some embodiments, a0 is the difference between the 95th percentile and the 5th percentile of all possible log2(P(r,v,b)) values, divided by 256. This allows Pnorm(r,v,b) to be normalized to the range of -127 to 127, i.e., a fixed-point number.
[0072] The above is merely an illustrative explanation of related data processing, and other processing methods may be used in some embodiments. For example, when obtaining a logarithm, the logarithm of 10 is obtained instead of 2. Here, the data can be normalized to -127 to 127 by taking the logarithm of 2 and then using the above formula, and other methods of obtaining logarithms or other processing methods can achieve similar effects. Alternatively, instead of quantification, converting to double-precision floating-point or single-precision floating-point numbers such as int32, int16, int8 or uint32, uint16, uint8 can also improve the efficiency of subsequent processing, so this will not be explained here.
[0073] Here, the implementation of the normalization embodiments described above can take various forms. For example, the complex number data may be transmitted to the MCU Mem via DMA and then performed in the MCU, or it may be implemented by a hardware engine in the BB and then directly transmitted to the MCU Mem via DMA, or part of it may be implemented in the BB and then transmitted to the MCU Mem via DMA and the remaining calculations performed in the MCU (the operation of finding the norm and the operation of obtaining the logarithm of 2 can be implemented in the BB, and subtraction, division and saturation truncation can be implemented in the MCU), but the explanation is omitted here.
[0074] In step 102, the deep learning model used is not limited; any deep learning model such as Convolutional Neural Networks (CNNs) or Transformers may be used, and their explanation is omitted here.
[0075] Furthermore, the processing steps and output of deep learning models can be adjusted through the model's structure and training.
[0076] For example, in some embodiments, the result of target detection is a confidence matrix. Each element of the confidence matrix corresponds to at least one of the following pieces of information: the confidence that a target is detected, the confidence that a target exists in the corresponding region of interest, the confidence that an adult target exists in the corresponding region of interest, and the confidence that a child target exists in the corresponding region of interest.
[0077] Furthermore, in some embodiments, the structure of the deep learning model is a CNN with three convolutional kernels. The type, input, and output size of each layer are shown in Figure 8. In this case, the input feature dimension of the deep learning model is 6 × 64 × 20, corresponding to 6 beampoints, 64 Doppler units, and 20 range units. Each convolution operation includes a two-dimensional convolutional layer, a batch normalization layer, a ReLU nonlinear activation layer, and a max pooling layer. After three convolution operations, the resulting 128 × 2 feature vectors are flattened to one dimension to obtain a 256 × 1 one-dimensional feature vector, which is then randomly deactivated (dropped out) and the detection results are obtained through a 256 × 10 fully connected layer. For example, the detection could determine whether or not there are people inside a car, and if so, the confidence level that adults and children are present in each seat and aisle inside the car.
[0078] In this case, further processing may be applied based on the output of the deep learning model before outputting to the user. For example, using the car interior space shown in Figure 6 as an example, let's assume that the output of the deep learning model is a 1×10 vector yi formed by the confidence levels indicating whether there are people in the car, whether there are children in the three seats, whether there are children in the three aisles, and whether there are adults in the three seats. In this case, the following determinations can be made: If yi(1)>0, it is directly determined that there are people in the car without considering the other digits of yi, or if yi(k)=0, k=1,2,...,10, it is determined that there are no people in the car. If k=2,3,...,10 exists and yi(k)>0, it is determined that there are children or adults in the corresponding region. If the confidence levels for the presence of adults and children in seat A and aisle A are both greater than 0, it is determined that there is only an adult in seat A, and by analogy, seat B and aisle B and seat C and aisle C are inferred. If yi=[-0.1,7.2,-5,-7.2,-1.5,-1.9,-3.1,-6.4,-6.2,-5.5], it is determined that a child is present in seat A. If yi=[-0.1,-5,-4.2,7.2,-1.5,-1.9,-3.1,6.4,-6.2,-5.5], it is determined that a child is in aisle C and an adult is present in seat A. If yi=[-0.1,-5,-4.2,7.2,-1.5,-1.9,-3.1,-6.4,-6.2,5.5], it is determined that an adult is present only in seat C. If yi = [0.1, -5, -4.2, 7.2, -1.5, -1.9, -3.1, 6.4, -6.2, -5.5], it is determined that the vehicle is empty. In this case, it is possible to determine situations where multiple people are inside the vehicle, and false alarms can be significantly reduced, although the risk of detection failure increases slightly.
[0079] Regarding the training method for deep learning models, in some embodiments, as shown in Figure 9, the method further includes the following steps before processing feature vectors based on the trained deep learning model so that the target detection results output by the deep learning model are obtained.
[0080] In step 103, a training feature vector is obtained that has a one-hot encoding label and characterizes the energy features of the signal in different range units.
[0081] In step 104, the training feature vectors are processed by a predetermined data augmentation mode, which includes at least one of random increase of white noise, inversion of data along the Doppler dimension, and translation of data along the range dimension, so that a training set is obtained.
[0082] In step 105, the deep learning model is trained based on the training set.
[0083] To help those skilled in the art better understand the above embodiments, the training of a deep learning model will be explained below with an example.
[0084] First, training feature vectors are obtained according to the method for obtaining feature vectors described above, and the input feature data set S{P} is constructed by accumulating the processing data from multiple trials in multiple sets of experiments. Here, each element Pi is the (Nv+1)×(Nr+1)×N array described above.
[0085] Next, S{P} is labeled based on the experimental conditions, i.e., a label Li corresponding to Pi is generated, and the labeled input feature data set S{P,L} is obtained. In this case, one-hot encoding can be used to generate the label Li. That is, in the car interior scene shown in Figure 6, a total of 10 classification problems are performed, considering whether there are people in the three rear seats and three aisle areas of the car, as well as considering the classification of adults and children, while not considering the situation of adults crouching in the three aisle areas. Li is a 1×10 vector, each value being either 0 or 1. The first digit of Li represents whether there are people in the car; if there are no people, this digit is 1, and if there are people, this digit is 0. The second to fourth digits represent whether there are children in seats A to C; if there are, this digit is 1, and if there are no children, this digit is 0. The fifth to seventh digits represent whether there are children in aisles A to C; if there are, this digit is 1, and if there are no children, this digit is 0. The 8th to 10th digits indicate whether or not there are adults in seats A through C. If there are adults, these digits are 1; otherwise, they are 0. If there are no people in the vehicle, the corresponding Li is [1,0,0,0,0,0,0,0,0,0]. If there is an adult in seat A, the corresponding Li is [0,0,0,0,0,0,0,1,0,0]. If there is a child in aisle A and an adult in seat B, the corresponding Li is [0,0,0,0,1,0,0,0,1,0]. The first digit of Li, which indicates whether or not there are people in the vehicle, may be omitted. In other words, Li is a 1x9 vector, and if all digits are zero, it means there are no people in the vehicle.
[0086] Next, data augmentation is performed on S{P,L} using one or more methods such as random increase of white noise, random inversion along the Doppler dimension, and random small-number translation along the range dimension, thereby obtaining the augmented data set S'{P,L}. The number of samples in the augmented data set is significantly increased compared to the remote data set. For example, if S{P,L} has 5,000 samples each for empty cars, the presence of children, and the presence of adults, then after 10 rounds of data augmentation, an augmented data set S'{P,L} with 50,000 samples each can be obtained (it is possible to choose whether or not to include the original data set S{P,L}). Here, S'{P,L} can be randomly rearranged to obtain the validation data set Sv{P,L} with a ratio of α=0.3, and the remaining data can be used as the training data set St{P,L}.
[0087] Then, the deep learning model is trained as follows.
[0088] We build an uninitialized CNN model and then randomly initialize the parameters of each layer.
[0089] Repeat the following steps 20 times (each time is called an epoch).
[0090] St{P,L} is randomly divided into Nb batches, each batch containing Ns samples, and the following steps are performed for each batch.
[0091] For each batch of data, one or more data augmentation processes are performed again, such as random increase of white noise, random inversion along the Doppler dimension, and random small-scale translation along the range dimension.
[0092] The augmented data is sent to a CNN for forward computation, and three types of confidence levels are obtained: no people, adult presence, and child presence. The corresponding loss CEt is then obtained based on the loss function.
[0093] Backpropagation is performed on the CNN network model, and the parameters of each layer of the CNN are updated based on the gradient descent optimization algorithm.
[0094] After performing the above operation for all St{P,L}, the updated CNN is used to perform forward calculations for all samples of Sv{P,L} to obtain the loss CEv for Sv{P,L}, and if CEv does not decrease within the set Nes cycles, the cycle is terminated early.
[0095] The final model output is obtained as the CNN model with the minimum number of cycles for which CEv is minimized.
[0096] Here, during the training process, the Adam Optimizer can be used as the gradient descent optimizer, and cross-entropy is selected as the loss function. However, for the network output, normalization is first performed using sigmoid, and then the loss is calculated using cross-entropy. The learning rate is set to 0.001, the weight decay is set to 0.0001, Ns is set to 128, and Nes is set to 5.
[0097] Of course, the above is merely an illustrative explanation of a deep learning model. Other examples may employ models with different structures, CNNs with different parameters, or different training parameters and training methods, but these will not be explained here.
[0098] It should be noted that the embodiments described above are merely illustrative, and in some embodiments, the output and label assignment can take other forms without employing one-hot encoding. For example, if we consider only scenarios where there are no people in the car or only one person, Li=1,2,...,10 will be directly generated as labels. For the generated yi, we search for the digit in which the maximum value is located as the detection and recognition result, but we will omit the explanation here.
[0099] While radar coverage is large, users do not necessarily have detection needs for the entire coverage area. Therefore, in some embodiments, feature vectors are used to characterize the energy features of echo signals in different range units within a given range and / or Doppler range. For example, when beamforming, only certain Doppler units and range units are used, such as 0-32 and 96-127 Doppler units, and 10-42 range units. In this way, only the parts of interest to the user are retained in order to reduce the occupation of resources such as computation and storage, and to improve detection efficiency.
[0100] To help those skilled in the art better understand the application of the target detection method according to the above embodiment in a radar-based target detection process, the application of the embodiment shown in Figures 3 to 5 will be described below as an example.
[0101] In the embodiment shown in Figure 3, after being applied to the complete target detection flow, as shown in Figure 10, i.e., based on the flow shown in Figure 1, inter-frame FFT (Frame FFT) processing is performed, followed directly by two-dimensional digital beamforming (2D-DBF) processing and DL-based target classification (DL-based classification) operations. That is, CFAR processing is not performed after the inter-frame FFT, and the azimuth and elevation dimensions are calculated for each region of interest (or unit interval), and the 2D-DBF operation is performed so that a number of RD spectra corresponding to the region of interest are obtained. Subsequently, the above RD spectra are input to the constructed deep learning model (DL) to perform target determination. In some arbitrary embodiments, the relationship between the region of interest and the RD spectra may be one-to-many. That is, at least two RD spectra can be obtained based on one region of interest, and the specific number can be adjusted according to the actual demand.
[0102] Here, the DL input can be linear or in the dB range, or amplitude or power, and can be subjected to operations such as normalization. At the same time, the constructed deep learning model can classify based on the input, and the categories to be distinguished include the presence or absence of a target, the attributes of the target, and the specific unit interval location (i.e., region location) where each target is located.
[0103] For example, in an application scenario where the rear row area of a vehicle is called the region of interest or target region (i.e., the scene shown in Figure 6), and is divided into three seat areas (section units) and three corresponding aisle areas, after an interframe FFT, a 2D-DBF of azimuth and elevation is performed on the center position of the six regions of interest (the regions of interest are also set to three seat areas) so that six RD spectra (or three RD spectra) corresponding to the regions of interest are obtained. At least a portion of these six RD spectra (or three RD spectra) are input into the constructed deep learning model to perform operations such as target determination and distinction. For example, the constructed deep learning model can classify based on the input to determine whether or not there is a biological target in the region of interest. If there is a biological target, it can further determine whether the biological target is an adult, child, infant, or pet, and can further determine specific location information such as which seat or aisle area the biological target is located in.
[0104] In the flow shown in Figure 10, the target detection method is a processing flow based on model data dual drive, which has low requirements for signal processing. Therefore, it effectively avoids the selection of signal processing parameters, domain parameters, and the design of CFAR logic and domain judgment logic, and consequently, the difficulty of implementing and designing the method is greatly reduced.
[0105] The embodiment shown in Figure 4, after being applied to the complete target detection process, becomes as shown in Figure 11. As shown in Figure 11, based on the flow shown in Figure 1, after the range-dimensional Fourier transform (Range FFT, also called 1D-FFT), a 2D-DBF is first performed, and then operations such as frame data accumulation and inter-frame FFT processing are carried out. That is, before the inter-frame FFT, a 2D-DBF of the azimuth and elevation angles is first performed with respect to the center position of the region of interest (such as three seating areas). This embodiment can realize a means of parallel processing of data processing and radio wave transmission time, and consequently, the processing time can be effectively reduced.
[0106] The embodiment shown in Figure 5, after being applied to the complete target detection process, becomes as shown in Figure 12. As shown in Figure 12, based on the flow shown in Figure 1, after a range-dimensional Fourier transform (Range FFT, also called 1D-FFT), a 2D-DBF is performed, followed by a direct Complex DL-based Classification operation. That is, throughout the entire target detection flow, inter-frame FFT processing is not performed so that a corresponding number of range-frame spectra (e.g., three range-frames) are obtained. After the 1D-FFT, a 2D-DBF is directly performed on the center positions of the regions of interest (e.g., three seat regions), and then the above range-frame spectra are input to a constructed (e.g., complex numbers) deep learning model to perform operations such as target distinction and judgment.
[0107] In the embodiment shown in Figure 12, the FFT processing step between frames is avoided, thereby effectively reducing the computational complexity. At the same time, when employing a complex deep learning model, a higher processing gain can be obtained than with FFT, thus achieving the objective of improving target detection performance. Furthermore, this embodiment effectively shortens the processing hierarchy, making efficient adjustment operations and other operations more convenient.
[0108] As can be seen from the above, by combining 2D-FFT and DL-based classification in the flow shown in Figures 4 and 6, the accuracy of target recognition can be effectively improved. At the same time, the steps of 2D-FFT and DL-based classification can be flexibly set between each step of signal processing according to actual needs, for example, by setting the 2D-FFT step after Range FFT or after Frame FFT.
[0109] In the above embodiment, all FFT processing can be performed as windowed FFT, and can also be achieved by combining SVA and FFT. Alternatively, the multi-frame sliding window FFT can be replaced with other time-frequency transformation processing such as short-time Fourier transform or fractional-order Fourier transform, enabling effective detection of the target of interest. Alternatively, the multi-frame sliding window FFT can be replaced with high-pass FIR, Comb FIR, or filters optimized and synthesized based on an optimization function, thereby achieving performance equivalent to or better than multi-frame FFT processing with fewer hardware resources. Alternatively, DBF can be replaced with Capon, MUSIC, ESPRINT, or their variant algorithms, and beamforming at the center of the region of interest can be optimized and synthesized, and / or the antenna array can be optimized and synthesized, further improving the system's target detection performance.
[0110] It should be noted that the embodiments shown in Figures 10 to 12 are merely means of target detection combined with multi-frame storage technology, but this does not mean that multi-frame storage is necessarily incorporated. For example, in the embodiment shown in Figure 12, inter-frame storage is omitted, and of course, its explanation is omitted here.
[0111] As shown in Figures 13-21, the embodiments of this invention further provide another target detection method. By combining interframe storage and deep learning modes, coherent storage is performed on the frequency of human respiration, improving the signal-to-noise ratio between personnel targets and statically strong clutter, and effectively constructing feature data for empty vehicles (with or without interference), adult presence scenes, and child presence scenes. Simultaneously, by deep mining the constructed feature data in combination with a deep learning model, effective detection of whether or not there are people inside the vehicle and accurate determination of adults and children can be achieved, enabling more accurate target detection and recognition in spatial domains such as car cabins, indoors, and factory buildings. Compared to point cloud information based on signal processing, the method according to this embodiment utilizes higher data feature dimensions and richer information, thus enabling more accurate target detection and recognition inside cabins. Specifically, it is as follows:
[0112] As shown in Figure 13, the target detection method includes the following steps.
[0113] In step S11, for each chirp data in each frame, for example in BB, the downlink ADC data can be DC filtered. Optionally, first, the average can be calculated along the fast time dimension for the data collected by each RX channel of each chirp, and the DC component can be subtracted from all sampling points of each RX channel. Next, the DC filtered data is windowed, and then a 1D-FFT is performed to obtain range dimension information. The 1D-FFT data is then moved from BB to the CPU SRAM cache by DMA. Finally, the above operations are performed similarly for different chirps.
[0114] In step S12, if the data cached in the CPU SRAM reaches a certain number of frames, such as 128 frames, the data for these 128 frames is moved from the CPU SRAM to the BB using DMA, and the next step is executed.
[0115] First, the multi-frame 1D data is DC filtered, and a complex number average along the frame dimension is calculated for this 128 frames of data to obtain the average value of different ranges across different channels as its DC component. This DC component is then subtracted from each range unit of each channel.
[0116] Next, the DC-filtered data is windowed along the frame dimension so that the range Doppler complex value x(c,r,v) for each transmit and receive channel is obtained, and then a 2D-FFT calculation is performed along the frame dimension.
[0117] Subsequently, some Doppler units and range units, such as the 0-32 and 96-127 Doppler units and the 10-42 range unit, are extracted. The energy is then calculated for the complex value x(c,r,v) of each Doppler unit in each range unit of each channel, followed by a logarithmic operation to normalize the result. The resulting mathematical formula is as follows:
[0118]
number
[0119] In the formula, P(c,r,v) is the normalized "energy" of the virtual channel c, range unit r, and Doppler unit v. b0 and a0 are normalized predetermined parameters. Pmax and Pmin are the upper and lower limits of a given P(c,r,v). Typically, the same b0, a0, Pmax, and Pmin values are selected for different c, r, and v (in possible modifications, different b0, a0, Pmax, and Pmin values can be selected for different c and different v and r partitions). In one possible embodiment, b0 can be set to the mean or median of all possible log2(P(r,v,b)) values, and a0 can be set to the difference between the 95th percentile and the 5th percentile of all possible log2(P(r,v,b)) values divided by 256. This allows P(c,r,v) to be normalized between -127 and 127.
[0120] Next, P(c,r,v) is rearranged into the format shown in Figure 7, that is, in physical memory, it is contiguous along the v dimension, then the r dimension, and finally the c dimension. In the figure, v0 and vNv represent the 0th and Nvth Doppler bins (totaling Nv+1 Doppler bins), r0 and rNr represent the 0th and Nrth range bins (totaling Nr+1 range bins), and c0 represents the 0th virtual channel, totaling Nc+1 virtual channels. This (Nv+1)×(Nr+1)×(Nc+1) array is the input feature data for the deep learning model.
[0121] Finally, the above data is sent to the constructed deep learning model to determine whether there is a person inside the cabin, and whether they are an adult or a child, and the result is output.
[0122] The deep learning model may be a CNN or a Transformer. As shown in Figure 14, the main steps in one embodiment of the model construction process are as follows.
[0123] In step S21, input feature data is obtained using the same steps as above, and the input feature data set is constructed by accumulating the processing data from multiple sets of experiments. Here, each element Pi is an (Nv+1)×(Nr+1)×(Nc+1) array, and its definition is the same as in step 2.c) above.
[0124] In step S22, S{P} is labeled based on the experimental conditions, i.e., a label Li corresponding to Pi is generated, and the labeled input feature data set S{P,L} is obtained.
[0125] In step S23, data augmentation is performed on S{P,L} using one or more methods such as random increase of white noise, random inversion along the Doppler dimension, and random translation of a small number along the range dimension to obtain the augmented data set S'{P,L}. The number of samples in the augmented data set is significantly increased compared to the remote data set. For example, if S{P,L} has 5,000 samples each for empty cars, the presence of children, and the presence of adults, then after 10 rounds of data augmentation, an augmented data set S'{P,L} can be obtained with 50,000 samples each for each category (it is possible to choose whether or not to include the original data set S{P,L}).
[0126] In step S24, S'{P,L} is randomly rearranged to obtain the validation data set Sv{P,L} with a ratio of α=0.3, and the remaining data is used as the training data set St{P,L}.
[0127] In step S25, the deep learning model is trained based on St{P,L} and Sv{P,L}.
[0128] As an option, the main steps of one embodiment are as follows:
[0129] First, we build an uninitialized CNN model and then randomly initialize the parameters of each layer.
[0130] Next, the following steps are repeated 20 times (each time referred to as one epoch). First, St{P,L} can be randomly divided into Nb batches, each containing Ns samples. For each batch, the following steps are performed. First, one or more data augmentation processes are performed again on the data in each batch, such as random increase of white noise, random inversion along the Doppler dimension, and random small number of parallel shifts along the range dimension. Next, the augmented data is sent to the CNN for forward computation to obtain three confidence levels: no people, adult presence, and child presence, and the corresponding loss CEt is obtained based on the loss function. Finally, backpropagation is performed on the CNN network model, and the parameters of each layer of the CNN are updated based on the gradient descent optimization algorithm. Then, after performing the above operations for all St{P,L}, forward computation is performed on all samples of Sv{P,L} using the updated CNN to obtain the loss CEv for Sv{P,L}, and if CEv does not decrease after a set Nes cycles, the cycle is terminated early.
[0131] Next, we obtain the CNN model with the fewest cycles as the final model output, which minimizes CEv.
[0132] In the model training process, in possible implementations, the Adam Optimizer can be used as a gradient descent optimizer, and cross-entropy can be selected as the loss function. The learning rate is set to 0.001, the weight decay is set to 0.0001, Ns is set to 128, and Nes is set to 5.
[0133] As shown in Figure 15, possible CNN structures are described below. Here, a CNN with three convolutional kernels is shown, and the type of each layer and the input and output sizes are also shown in the figure. Here, the input feature dimension is 16 × 64 × 20, corresponding to 16 virtual channels, 64 Doppler units, and 20 range units. Each convolution operation includes a two-dimensional convolutional layer, a batch normalization layer, a ReLU nonlinear activation layer, and a max pooling layer. After three convolution operations, the resulting 128 × 2 feature vector is flattened to one dimension to obtain a 256 × 1 one-dimensional feature vector, randomly deactivated (drop-out), and then three confidence levels for empty cars, the presence of children, and the presence of adults are obtained through a 256 × 10 fully connected layer. These confidence levels can be calculated using softmax to obtain the maximum value as the classification of the input, that is, it can be determined that the input belongs to one of the categories of empty cars, the presence of children, or the presence of adults.
[0134] Furthermore, 1. Unlike conventional low-speed time processing between chirps in radar signal processing, the method in this embodiment accumulates multi-frame data and performs inter-frame FFT processing using a sliding window.
[0135] 2. The method in this embodiment proposes a process that detects and classifies data using interframe FFT processing or CNN, without requiring the acquisition of point clouds using conventional signal processing methods such as CFAR or DoA.
[0136] 3. The method in this embodiment details the construction of a CNN input feature for the first time, and the points of interest and possible changes include the following:
[0137] 1. When calculating P(c,r,v), the log2 operation can be skipped, or the log10 or ln operation can be used instead.
[0138] 2. When calculating P(c,r,v), the calculation may be performed in the MCU after the complex number data is sent to the MCU Mem via DMA, or it may be performed by the hardware engine in the BB and then directly sent to the MCU Mem via DMA, or part of the calculation may be performed in the BB and then sent to the MCU Mem via DMA, with the remaining calculation performed in the MCU (the operation to find the norm and the log2 operation can be performed in the BB, while subtraction, division and saturation truncation can be performed in the MCU).
[0139] 3. P(c,r,v) may be expressed as a double-precision floating-point or single-precision floating-point number without quantification, or it may be expressed in the quantification format in int32, int16, int8, or uint32, uint16, uint8.
[0140] 4. The above data array may be continuous along the r dimension, then the v dimension and c dimension (as shown in Figure 23), or it may be vcr or cvr dimension.
[0141] 5. In a deep learning model, virtual channels may be selected as the model's channels, or the virtual channel dimension may be tiled to adopt the input of a single-channel deep learning model. The input of the deep learning model is 1 × 64 × 320, containing 64 Doppler units, and the 20 range units of 16 channels are tiled as 320.
[0142] 6. Regarding channels, to avoid the impact of noisy channels or channels with poor radio frequency simulation characteristics on processing effectiveness, only some channels may be selected. For example, 13 out of 16 channels (4 transmit and 4 receive) can be used to construct 13 × 64 × 20 CNN input characteristic data to determine whether a vehicle is empty, a child, or an adult.
[0143] 7. Data after interframe FFT may be truncated to save memory and reduce computational load, or it may not be truncated to obtain optimal performance. If truncation is performed, it can be done after the 1D-FFT or 2D-FFT, or when calculating P, or a combination thereof.
[0144] As shown in Figure 16, the target detection method in the embodiment of the present invention may further employ the following processing flow shown in Figure 16: non-coherent storage of data after interframe FFT according to the channels so that an array is obtained, and after performing the same operation as in the above step of cutting out some Doppler units and range units, the data is sent to a deep learning model as input feature data for processing to obtain a classification result, and the number of input channels of the corresponding deep learning model is selected as 1.
[0145] As shown in Figure 17, the target detection method in the embodiment of the present invention may further employ the processing flow shown in Figure 17. After performing non-coherent storage of the data after interframe FFT according to the channel so that two-dimensional data can be obtained, the same operation as in the above step is employed to select only the energy of points that exceed a predetermined threshold using CFAR and cut out some Doppler units and range units, then the range units and Doppler units that do not exceed the threshold are filled as input feature data, and then the input feature data is sent to a deep learning model for processing to obtain a classification result, and the number of input channels of the corresponding deep learning model is selected as 1.
[0146] Furthermore, a combination of the two embodiments shown in Figures 16 and 17 can be adopted. That is, after performing CFAR on the data after interframe FFT according to the channel, selecting only the energy of points that exceed a predetermined threshold for each channel, and employing an operation similar to the above step to cut out some Doppler units and range units, the range units and Doppler units that do not exceed the threshold are filled as input feature data, then the data of all channels is arranged in the same manner as the rearrangement of P(c,r,v), the input feature data is sent to a deep learning model for processing to obtain the classification result, and the number of input channels of the corresponding deep learning model is selected as the number of virtual channels to be used.
[0147] The deep learning model in the embodiment of this application can be any other CNN or Transformer, as long as it is adapted to the input. As shown in Figures 18 and 19, it can be based on a CNN model modified by ResNet18 and an extension of its Sequential module. The output here can be of two types: empty cars and non-empty cars.
[0148] The two figures in Figure 20 show examples of CNN input feature diagrams corresponding to adults and children, respectively. The left figure shows the input features for adults, and the right figure shows the input features for children. Each sub-figure corresponds to the data for each channel. The horizontal axis represents Doppler units, and the vertical axis represents range units. In this example, 20 range units and 64 Doppler units have been extracted.
[0149] Figure 21 shows the changes in the training model loss and model classification accuracy with each training epoch. It can be seen that the model converges in approximately 7 epochs and can achieve an accuracy of 98%.
[0150] Figure 22 shows the test confidence for adult and child samples. Here, the horizontal axis represents the sample number, and the vertical axis represents the confidence score. In the experiment, the first 150 samples are samples of adults in the car, and the last 150 samples are samples of children in the car. For each sample, a possible deep learning model calculates the confidence score corresponding to adults and children, respectively. That is, two points, blue and orange, are generated for each sample. If the confidence score for adults is greater than the confidence score for children, it is determined that there are adults in the car, and if the confidence score for children is greater than the confidence score for adults, it is determined that there are children in the car. As can be seen from the figure, the trained CNN model can correctly determine whether the person in the car is an adult or a child.
[0151] In some embodiments, as shown in Figure 24, the target detection method includes the following:
[0152] In step 10, after performing a range-dimensional FFT based on the echo signal, 1D-FFT data is obtained.
[0153] In step 20, digital beamforming is performed on the 1D-FFT data to acquire a predetermined number of frame data.
[0154] In step 30, a target classification process is performed on a predetermined number of frame data based on a machine learning model so that target detection can be achieved.
[0155] The above-mentioned target detection in the region of interest includes at least one operation from among judgment, localization, and recognition.
[0156] By performing beamforming on a predetermined region after range-dimensional FFT, and then using multi-frame beamformed data to perform target classification processing directly based on a machine learning model, effective detection of whether or not there are people inside the vehicle and accurate determination of whether they are adults or children can be achieved. Compared to point cloud information based on conventional signal processing, the method of this embodiment has lower requirements for processing units (such as BB units), utilizes higher data feature dimensions, and has a richer amount of information, thus enabling more accurate target detection and recognition inside the cabin and facilitating the application of CPD and SBR.
[0157] In the exemplary embodiment, the machine learning model may be a complex machine learning model, such as a complex neural network model.
[0158] Note that the embodiment shown in Figure 24 differs from the embodiment described above primarily in that the model used in the embodiment shown in Figure 24 is a machine learning model, while the model used in the embodiment described above is a deep learning model. To adapt to the input of a complex machine learning model, some embodiments can be implemented by transforming the feature vector. The transformation method involves converting the power and phase into the real and imaginary parts.
[0159]
number
number
number
[0160] In the formula, zreal and zimg are the normalized real and imaginary parts, respectively. φ(r,v,b) is the phase of the range unit r chirp v in beam b. z(r,v,b) is the normalized complex signal. In some embodiments, double-precision floating-point or single-precision floating-point numbers may be used for z(r,v,b), or quantitative forms such as int32, int16, int8, uint32, uint16, uint8, or other quantification processes may be performed.
[0161] Furthermore, different loss functions can be used based on different models. For example, in a complex neural network model, the absolute value (abs) layer of the last layer of a complex-valued CNN can be removed. In this case, the output of the complex-valued CNN is a complex number. In this case, the following loss function can be adopted.
[0162]
number
number
number
[0163] Here, y is the forward computation output of a Complex-valued CNN for a given sample, and is a vector with a length of the classification count. yi is the i-th scalar of y. t is the one-hot encoded label of the sample, and is also a vector with a length of the classification count. ti is the i-th scalar of t. Other hyperparameters, such as the learning rate and weight decay, can have other values.
[0164] The embodiment shown in Figure 24 and the embodiments described above have essentially the same other features or methods for realizing those features, and therefore will not be explained here.
[0165] Of course, in the above embodiments, the use of the model replaces corresponding digital signal processing processes, such as angle estimation or CFAR processing. In some embodiments, the model may be used to further process targets obtained by digital signal processing.
[0166] Based on this, in some embodiments, as shown in Figure 25, the target detection method includes the following steps.
[0167] In step 30, after performing range-dimensional FFT processing based on the echo signal, 1D-FFT data is acquired, and the 1D-FFT data is processed in a multi-frame linked manner so that the RD spectrum is obtained as the feature vector.
[0168] In step 40, target point cloud data is acquired based on the RD spectrum, and machine learning-based target classification processing is performed on the target point cloud data, so that the target in the region of interest can be detected.
[0169] The above-mentioned target detection in the region of interest includes at least one operation from among judgment, localization, and recognition.
[0170] The target detection method and related device according to this embodiment utilize multi-frame collaborative processing technology, i.e., an inter-frame storage method, to realize a means for target detection in enclosed spatial areas such as cabins, indoors, and factory buildings. This effectively improves the detection rate, reduces the number of false alarm targets, and significantly improves the accuracy of angle estimation. As a result, it can accurately detect special targets such as infants in cabins or weak targets, enabling applications such as CPD (Child Presence Detection) and SBR (Safety Belt Reminder). Furthermore, by performing machine learning-based target classification processing on point cloud data, more precise area determination and more accurate recognition of targets within an area can be achieved. For example, it can improve the accuracy of judgments regarding the presence or absence of personnel and their locations in enclosed environments such as cabins.
[0171] In an exemplary embodiment, step 30 includes performing a range-dimensional FFT in a chirp on the echo signal to obtain 1D-FFT data, accumulating frame data on the 1D-FFT data until a predetermined amount of data is accumulated, and then performing an inter-frame FFT to obtain an RD spectrum. For example, a sliding window can be used to read a predetermined number of 1D-FFT data each time and perform an inter-frame FFT. That is, accurate detection of targets inside the cabin can be achieved based on multi-frame collaborative processing technology. The multi-frame collaborative processing means performs a sliding window FFT on the multi-frame data to obtain a range Doppler spectrum, then processes it using FIR (Finite Impulse Response) or other complex time-frequency transformations, and then performs processing operations such as CFAR and DOA on the region of interest. By combining this with a set region judgment logic and region parameters, detection and localization operations for biological targets such as adults, children, pets, or other non-biological targets are realized. The CFAR may be a Doppler-dimensional NR-CFAR, RD-CFAR, or DAE (Doppler-Azimuth-Elevation)-CFAR, etc. After the CFAR and DoA, post-processing such as clustering, false alarm suppression, and association of point clouds processed multiple times can be employed. Region parameters can be identified by at least one or a combination of at least two means, such as a clustering operation on point clouds detected after a sliding window multi-frame processing, or an outlier detection and removal operation on point clouds detected after a sliding window multi-frame processing.
[0172] In an exemplary embodiment, the step of acquiring target point cloud data in step 40 above includes performing non-coherent storage on the RD spectrum of each channel so that candidate target detection points are obtained, performing constant false alarm processing based on noise estimation based on the results of the non-coherent storage, estimating the azimuth midpoint angle and elevation angle for the candidate target detection points, and acquiring target point cloud data based on the azimuth midpoint angle and elevation angle. By performing non-coherent storage on the RD spectrum of each channel, the signal-to-noise ratio between personnel targets and statically strong clutter can be effectively improved.
[0173] In some embodiments, the constant false alarm processing includes performing non-coherent accumulation on the RD spectrum of each channel so that an estimated noise floor value for each range unit is obtained, estimating the noise floor of each range unit after non-coherent accumulation, and performing non-coherent constant false alarm processing based on the estimated noise floor value. Here, when estimating the noise floor of each range unit after non-coherent accumulation, the noise floor of each range unit is corrected by the global noise floor.
[0174] In some embodiments, the correction method includes correcting the noise floor of each range unit by ni' = min(ni, α * ng), where ni is the original noise floor estimate of the i-th range unit, ng is the average of the noise floor estimates of multiple range units, and ni' is the corrected noise floor estimate of the i-th range unit, where α > 1. By correcting the noise floor, it is possible to avoid situations where the target is not detected due to a high noise floor.
[0175] In an exemplary embodiment, step 40, which performs machine learning-based target classification processing on the target point cloud data so that target detection (e.g., judgment) in a region of interest is achieved, includes inputting the target point cloud data, including coordinate data, into a pre-trained first machine learning classifier so that a first detection result is obtained indicating whether or not a candidate target detection point is a valid candidate target detection point. The coordinate data refers to coordinates in a Cartesian coordinate system, including x, y, and z coordinates. In some embodiments, the data input to the first machine learning classifier includes one or more of range data, signal-to-noise ratio, amplitude value, azimuthal midpoint angle, and elevation angle.
[0176] In exemplary embodiments, step 40, which processes the target point cloud data for machine learning-based target classification so that target detection (e.g., localization or recognition) in a region of interest is achieved, includes processing the target point cloud data, extracting region information of candidate target detection points, and inputting the region information into a pre-trained second machine learning classifier so that a second detection result is obtained indicating whether or not there is a target in a given region. In some embodiments, the region information of candidate target detection points includes one or more of the following: the number of valid candidate target detection points in each given region, the ratio of valid candidate target detection points in each given region to the total number of valid candidate target detection points, the range of all valid candidate target detection points in each given region, the azimuth midpoint angle, the elevation angle, and the mean and variance of the constant false alarm signal to noise ratio.
[0177] By combining this with a machine learning classifier for classification and / or recognition, more precise region segmentation than a predetermined logic can be achieved, allowing for accurate classification of seats corresponding to radar-processed point clouds.
[0178] The first or second machine learning classifier may be one or more of the following: support vector machines, random forests, decision trees, Gaussian mixture models, KNNs, hidden Markovs, and multilayer perceptrons.
[0179] Here, when applied to a sealed or relatively sealed spatial area, different regions of interest (divided regions) can be defined and detected by pre-dividing the spatial area. For example, in the case of target detection inside a vehicle or cabin, the area monitored by radar can be divided into seating sections and aisle sections, etc. Then, corresponding parameter types and thresholds can be pre-set for different types of regions, and by combining corresponding processing steps, accurate detection of targets of interest in a specific region can be achieved.
[0180] In some embodiments, in the case of a private car, the interior space can generally be simply divided into headroom, rear-row space, and trunk space. If the rear-row space is the primary monitoring area, the rear-row space can be further divided into seating sections and aisle sections, etc. The seating sections may be divided into a corresponding number of seating section units based on the seats. Similarly, the aisle sections may be divided into a corresponding number of aisle section units corresponding to the seating sections. For example, in the case of a five-seater private car, the three seats in the rear-row space can be divided into three seating section units and three corresponding aisle section units. By pre-setting corresponding parameter types and thresholds for different types of areas (or sections or section units), and by employing appropriate signal data processing methods and steps, accurate detection of targets of interest (specific targets) in areas such as specific sections or section units can be achieved. Here, adjacent section units may have partially overlapping areas, and gaps of a predetermined width may be adjacent or separated.
[0181] The following is based on the technical concept of frame-level Fourier transform, and the technical means of this disclosure will be described in detail with reference to the drawings. Specifically, after performing a fast time-dimension Fourier transform on the chirp of the echo signal, the Fourier transform data of at least two frames in the fast time dimension is accumulated, and the data of at least two frames is performed a frame-dimension Fourier transform so that a Range-Doppler (RD) spectrum is obtained, and the target range and / or velocity is estimated based on the RD spectrum.
[0182] As shown in Figure 26A, after performing operations such as ADC (analog-to-digital) conversion and sampling on the echo signal, range DC removal and range Fourier transform (Range FFT, 1D-FFT, or fast time-dimension FFT) are performed sequentially. Then, multi-frame storage is performed (e.g., storing 128 frames of data), followed by Doppler DC removal, inter-frame Fourier transform (Frame FFT), and non-coherent integration. Next, based on the noise obtained by noise estimation (e.g., using a noise variance estimator), constant false alarm detection (CFAR (detection with peak selection) based on peak values) is performed. Finally, processing such as angle detection and target classification is performed so that target range, velocity, and / or angle information can be obtained. For example, based on the CFAR results, azimuth and elevation angles can be estimated first. Specifically, this can be achieved through techniques such as DBF (digital beam forming) and DoA (direction of arrival estimation).
[0183] After estimating the azimuth and elevation angles, the output target point cloud data can be processed in combination with a machine learning model to perform accurate target detection in the region of interest. For example, in the case of a single target point cloud data obtained after azimuth and elevation angle estimation, it can be combined with ML (Machine Learning, such as SVM or RF algorithms) false alarm suppression techniques to suppress non-ideal data such as noise, static clutter, and false alarms generated by target multipath. That is, one or more operations such as ML-based false-alarm suppression, clustering, and ML-based target classification are subsequently performed on the target point cloud data output by elevation DBF&DoA. The clustering process here may be an agglomerative clustering algorithm or an algorithm such as DBSCAN. In an exemplary embodiment, the point cloud data after clustering can be input into a pre-trained machine learning model (such as SVM or RF) to determine whether a person is present in the current processing. If present, its location can be further determined, enabling operations such as distinguishing and judging physical targets such as adults, children, and infants.
[0184] In the embodiment shown in Figure 26A, the influence of non-ideal factors such as noise, static clutter, and target multipath can be effectively reduced, and the difficulty of selecting domain parameters and designing judgment logic for complex domains can be effectively lowered. As a result, point cloud information can be used more effectively, and better target detection and discrimination performance can be obtained, especially for applications in complex scenarios.
[0185] In some arbitrary embodiments, for an electromagnetic wave sensor equipped with a BB (Baseband) unit and an MCU (Microcontroller Unit) unit, operations such as storing frame data (store 128 frames) and classifying targets (e.g., classification can be based on a Region Histogram) can be performed in the MCU module based on the target detection flowchart shown in Figure 26A. Range DC removal and Doppler DC removal can be performed in either the BB module or the MCU module. The flow shown in Figure 26A illustrates the example of DC filtering of downlink ADC data in the BB module.
[0186] As shown in Figure 26A, after performing 1D-FFT processing or acquiring range dimension information in the BB module, the 1D-FFT data is cached in the MCU module. After a predetermined amount of data (such as 128, 56, or 32 frames) has been cached, DC rejection and frame-level FFT can be performed on the cached data in a predetermined number of frames to obtain an RD (Range-Doppler) spectrum.
[0187] As shown in Figure 26A, if the system includes multiple channels, non-coherent storage can be performed on the multi-channel RD spectrum, and the noise floor of each range unit (range bin) of the data stored by the NVE (noise variance estimator) module can be estimated. For example, noise floor estimation can be performed based on a predetermined formula, and subsequent NR (Noise Reference)-CFAR processing can be performed based on the noise floor estimate. The above predetermined formula may be ni = min(ni, α * ng), where ni is the noise floor estimate for the i-th range bin, and ng is the noise floor estimate for the last range bin of interest. α is a coefficient greater than 1, which can be set and updated based on demand, engineering data, experience, etc.
[0188] In the case of angle estimation, as shown in Figure 26A, azimuthal-dimensional DBF and DoA, and elevation-dimensional DBF and DoA can be performed on the CFAR detection point data, respectively, so that azimuthal-elevation angle estimation data can be obtained for each CFAR detection point.
[0189] Figure 26B is a schematic diagram of another workflow for achieving target detection based on frame-level FFT in an embodiment of the present invention, and differs from Figure 2A in that the subsequent machine learning-based classification process is slightly different. One or more of the following operations are performed on the target point cloud data output by DBF&DoA: ML-based points classification, region feature extraction, and ML-based region classification. Here, ML-based points classification is used to classify the target point cloud data so that effective target detection points (i.e., the effective target candidate detection points mentioned above) can be obtained, enabling the determination of targets in the region of interest. Region feature extraction is useful for classification, for example, by enhancing classification performance through clustering, which helps improve the accuracy and performance of classification. ML-based region classification is used to locate and recognize targets in the region of interest, that is, to determine whether or not a target exists in a given region.
[0190] In the above embodiment, to achieve accurate detection of the target of interest, the conventional chirp-level Doppler FFT is replaced with a frame-level FFT to identify the target's velocity. Furthermore, by performing sliding window FFT processing using multiple frames, the frequency of result updates is further improved, thereby enhancing the real-time capabilities of the system. It is also possible to reduce the amount of data processed by processing only the range of interest and / or the Doppler region. When estimating background noise, the background noise obtained by combining the estimation based on the current range bin and the estimation of the last range bin of interest to obtain the minimum value is more suitable for application in relatively enclosed environments such as inside a car or cabin. When determining the region logic, the targets detected in each frame are counted according to a predetermined region design, and the counted results are judged according to predetermined rules, thereby enabling accurate determination of whether or not a target exists in each predetermined region.
[0191] Furthermore, as with the embodiments described above, when performing target detection, the implementation or substitution of the corresponding processing is not limited. For example, the FFT processing can be a windowed FFT, or it can be implemented by a combination of SVA and FFT, or a multi-frame sliding window FFT can be used to effectively detect targets of interest by replacing other time-frequency transformation processes such as short-time Fourier transform or fractional-order Fourier transform, and these will not be explained here.
[0192] In some embodiments, after inter-frame FFT (frame FFT) processing, two-dimensional digital beamforming (2D-DBF) processing and ML-based target classification (ML-based classification) operations can be performed directly. That is, after inter-frame FFT, azimuth and elevation angles are applied to each region of interest (or unit interval) without performing CFAR processing, and 2D-DBF operations are performed so that a number of RD spectra corresponding to the regions of interest are obtained. Subsequently, the above RD spectra can be input into a pre-built machine learning model (ML) to perform target determination. In some arbitrary embodiments, if processing is performed using FIR, STFT (Short-time Fourier Transform), etc., instead of inter-frame FFT, the above RD spectra are a one-dimensional spectral output corresponding to FIR and a three-dimensional spectrum which is a time-frequency transformation corresponding to STFT.
[0193] The embodiments of this application can be referenced and interchangeable with one another where they do not contradict each other, and the order and configuration of each functional module can be adjusted according to the requirements. In the case of a system comprising a BB module and an MCU module, where the system performs target detection by emitting electromagnetic waves and receiving a corresponding echo signal, each step of the target detection method according to the embodiments of this application can be arranged in accordance with the BB module and / or MCU module for operation, taking into consideration the actual requirements and aspects such as the data processing capability and timeliness of the system operation, and the relevant examples in the figures can be used as reference for some of these options.
[0194] By employing the method according to the embodiment of this application and using multi-frame collaborative processing technology to perform coherent storage on the frequency of human respiration in the Doppler domain and on multiple antenna channels at a predetermined orientation in the spatial domain, the signal-to-noise ratio between personnel targets and statically strong clutter can be effectively improved. Furthermore, by combining machine learning models to perform clustering and classification on radar point clouds, and achieving more precise area determination than predetermined logic, seats corresponding to radar processed point clouds can be accurately classified. By improving the signal-to-noise ratio, self-labeling of unsupervised clustering, and precise machine learning classification, it is possible to improve the accuracy of determining the presence or absence of personnel and their location in enclosed environments such as cabins. In addition, compared to rule-based area determination, the model construction and processing flow reduces the need for rule design, eases the difficulty of parameter tuning, and increases the convenience and applicability of the method.
[0195] Based on the above methods, after testing using the data actually collected, the application of the top-mounted radar system can effectively achieve a high detection rate and a very low detection failure rate and false alarm rate. For example, in the case of target detection inside the cabin, it is possible to achieve a detection rate of over 99%, a detection failure rate of less than 0.5%, and a false alarm rate of less than 1%.
[0196] Embodiments of the present application provide a target detection method applicable to target detection in a specific target region (e.g., a sealed or semi-sealed region). The method includes inputting target point cloud data into a pre-trained first machine learning classifier to obtain valid candidate target detection points, extracting region information of the valid candidate target detection points, and inputting this data into a pre-trained second machine learning classifier to obtain target detection results. Specific steps can be found in the descriptions of the embodiments in the preceding or following sections, and are omitted here.
[0197] Note that the steps in the flow shown in Figure 26B and the steps corresponding to the flow in the embodiment shown in Figure 1 and other figures mentioned above are almost identical in their implementation methods, and therefore, the explanation of the same parts will be omitted here. The following will mainly explain the differences.
[0198] In the flow shown in Figure 26B, when estimating the azimuthal midpoint angle of a candidate target detection point, the following operations are performed for each candidate target detection point Ti: Extract the 2D-FFT data corresponding to the detection point, select an azimuthal dimension antenna, perform azimuthal dimension DBF (digital beamforming) based on the azimuthal dimension steering vector, obtain an azimuthal dimension DBF spectrum showing the situation where the signal power or intensity in different directions changes with frequency, and estimate the azimuthal midpoint angles θi,j based on the azimuthal dimension DBF spectrum. The relationship between the azimuthal midpoint angle and the azimuthal angle is azimuthal midpoint angle = sin(azimuthal angle)cos(elevation angle). When estimating the elevation angle of a candidate target detection point, for each candidate target point Ti,j, an elevation direction steering vector sv(i,j) is generated based on its azimuthal midpoint angle θi,j and arrangement, obtain an elevation dimension DBF spectrum by performing elevation dimension DBF for each candidate target point Ti,j, and estimate the elevation angles φi,j,k based on the elevation dimension DBF spectrum. Here, i represents the number of the CFAR point (range unit, etc.). j represents the numbers of multiple candidate target points for a given CFAR point at different orientations. k represents the numbers of multiple candidate target points for a given CFAR point at different elevations.
[0199] Furthermore, in some exemplary embodiments, when estimating the azimuth mid-angle and elevation angle, the azimuth mid-angle and elevation angle can be estimated by directly performing azimuth-elevation two-dimensional DBF, or by other super-resolution algorithms such as the minimum variance unbiased estimation (Capon) algorithm or the multiple signal classification (MUSIC) algorithm. Therefore, when extracting target information, for all candidate target points Ti,j,k that satisfy the candidate target point detection conditions (e.g., exceeding the CFAR threshold), information can be extracted that includes, but is not limited to, the target point's range bin index, Doppler bin index, range ri,j,k, Doppler frequency fi,j,k, azimuth mid-angle θi,j,k, elevation angle φi,j,k, CFAR SNRsnri,j,k, and the amplitude of each channel ampi,j,k,c (where c is the channel index). Next, a coordinate system transformation is performed. For each candidate target point Ti,j,k, a coordinate system transformation is performed on the range ri,j,k, azimuthal intermediate angle θi,j,k, and elevation angle φi,j,k to obtain the x, y, and z coordinates in the Cartesian coordinate system. The data format to be transformed is as follows:
[0200]
number
number
number
[0201] Therefore, a pre-trained first machine learning classifier is used to classify the region to which the candidate target belongs. That is, by taking multiple attributes of the candidate target points Ti,j,k as input, the first machine learning classifier classifies the region to which each candidate target point belongs, thereby determining whether the candidate target points Ti,j,k are valid target points or invalid interference points in a particular region. Possible input examples for the first machine learning classifier are ri,j,k, xi,j,k, zi,j,k, yi,j,k, snri,j,k, ampi,j,k,c. Examples of machine learning classifiers include support vector machines (SVM) and random forests (RF).
[0202] The first machine learning classifier may have multiple forms, or even be a fusion of multiple forms, such as a fusion of a support vector machine, a decision tree, and a Gaussian mixture model. Weights are set for the output of each classifier so that a weighted average of all classifiers is obtained, and a final judgment is made again based on the weighted average result.
[0203] The input to the first machine learning classifier described above is merely an example, and in other embodiments, the attributes of the selected candidate target point Ti,j,k may be all or part of "ri,j,k,xi,j,k,zi,j,k,yi,j,k,fi,j,k,snri,j,k,ampi,j,k,c", for example, only coordinate data (the above x,y,z values), or in addition to coordinate data, one or more of range data (the above r), Doppler frequency (the above f), signal-to-noise ratio (the above SNR), and amplitude value (the above AMP). Alternatively, those skilled in the art may add input values to the first machine learning classifier based on the concept of this disclosure. For example, coordinate data may be replaced with azimuth midpoint angle and elevation angle, or azimuth midpoint angle and elevation angle data may be input together to the first machine learning classifier.
[0204] Furthermore, each region is classified by a pre-trained second machine learning classifier, i.e., region features are extracted. The number of effective target points in each region and the proportion of effective target points to the total number of effective target points (wl) are aggregated, and the mean and variance of the range, azimuth midpoint angle, elevation angle, and SNR for all effective target points in each region are calculated, i.e., the mean of each range.
number
number
number
number
number
number
number
number
[0205] For example, when aggregating the number of valid target points in each domain and the percentage wl of valid target points to the total number of valid target points, valid points determined to belong to multiple domains may be counted multiple times (such as in situations where people are present in both domain A and domain B, which are suitable for multi-label classification) or not (such as when they are suitable for single-label classification), but this disclosure is not limited to this.
[0206] The second machine learning classifier may have multiple forms, or even be a fusion of multiple forms, such as a fusion of a support vector machine, a decision tree, and a Gaussian mixture model. Weights are set for the output of each classifier so that a weighted average of all classifiers is obtained, and a decision is made again based on the weighted average result so that a final judgment result is obtained.
[0207] The input to the second machine learning classifier described above is merely an example; in other embodiments, the attributes of the selected candidate target points Ti, j, k may be "wl,
number
number
number
number
[0208] The first machine learning classifier used to determine valid points is one of the following: SVM, RF, decision tree, Gaussian mixture model, K-nearest neighbors (KNN), hidden Markov, or multilayer perceptron. The training steps for this classifier model include the following:
[0209] 1. The process for generating valid point labels is as follows:
[0210] As illustrated in Figure 6, the area is divided in advance. The horizontal coordinates in the figure are the X direction (width of the car) and the Z direction (rear end to front end of the car). The area is divided into a total of six regions: the three rear seats and the three aisles in front of the three seats. These regions may or may not overlap. Some areas inside the car may belong to multiple predetermined regions simultaneously, or may not belong to any region. In other embodiments, when dividing the region, the region can be subdivided in combination with the target number of people. For example, the rear seats A, B, and C can be divided into adult seat A, adult seat B, adult seat C, child seat A, child seat B, and child seat C. In this way, when the application requires determining whether there are people in rear seats A, B, and C, and whether they are adults or children, the results can be obtained directly. Only training is required to add classifications corresponding to effective point and region determinations.
[0211] For data on various scenarios, including scenarios where there is a person in seat A, a person in seat B, a person in seat C, or interference scenarios where there is no person, the attributes of all candidate target points Ti,j,k in each scenario are obtained using steps 1 to 10 above, and all or part of, for example, ri,j,k, xi,j,k, zi,j,k, yi,j,k, fi,j,k, snri,j,k, ampi,j,k,c can be selected. Perform the following operations across various area scenarios such as seats A, B, and C in Figure 5, and aisle areas A, B, and C.
[0212] 1.1 Clustering analysis is performed on all candidate target point data for the scene in question, and excluded points are removed. In some embodiments, clustering analysis of these candidate target point data can be performed using a density clustering algorithm (DBSCAN). In other embodiments, valid point labels can be generated using other clustering algorithms, or they can be manually labeled.
[0213] 1.2 After clustering analysis, the category with the most types is selected as the true target for the region, and its label is set to 1. All candidate points in other categories are false targets, and their labels are set to 0. In other embodiments, all non-noisy candidate target points from DBSCAN clustering analysis can be selected as true target points, and their labels can be set to 1, while noisy candidate target points from DBSCAN clustering analysis can be selected as invalid target points, and their labels can be set to 0. Alternatively, some categories with a large number of samples can be artificially selected as the true targets for the region, and their labels can be set to 1, while all candidate points in other categories are false targets, and their labels can be set to 0.
[0214] In empty vehicle and interference scenarios, all candidate target points are set as invalid target points and their labels are set to 0.
[0215] 2. Train the model based on the generated labels. The main step is to perform the following operations across various domain scenarios, such as seats A, B, and C, and aisle areas A, B, and C.
[0216] 2.1 Read the attributes and labels of all candidate target points in the region and denote them as S{Xi} and S{Li}, respectively, and read the attributes and labels of all candidate target points in the empty and interfering scenes and denote them as S{X0} and S{L0}, respectively. Here, each element of S{Xi} is an N×1 vector. A possible embodiment is to perform valid point discrimination by any six scalars (e.g., ri,j,k, xi,j,k, zi,j,k, yi,j,k, fi,j,k, and snri,j,k) of the attributes of the candidate target points ri,j,k, xi,j,k, zi,j,k, yi,j,k, fi,j,k and snri,j,k). When considering a total of 16 MIMO virtual channels, each element of S{Xl} is a 22×1 feature vector.
[0217] 2.2. Combine the attributes and labels of these target points to obtain training input data S{X}=S{Xi}∪S{X0} and training labels S{L}=S{Li}∪S{L0}.
[0218] 2.3. Train a machine learning classifier on the input data S{X} and labels S{L} so that a classifier model can be obtained.
[0219] In an application scenario where six areas are to be distinguished—seats A, B, and C in the back row, and aisles A, B, and C—six machine learning classifier models Mv,l can be obtained after performing the above steps. Here, l=1,2,...,6 corresponds to the six areas of seats A, B, and C in the back row, and aisles A, B, and C, respectively. Each machine learning classifier model Mv,l determines whether a candidate target point is a valid target point in each area based on the attributes of a specific candidate target point, such as a 22×1 feature vector consisting of ri,j,k,xi,j,k,zi,j,k,yi,j,k,fi,j,k,snri,j,k, and ampi,j,k,c.
[0220] For example, the first machine learning classifier can perform multiple classifications using a single classifier, or it can perform hierarchical multiple classifications using fewer classifiers than the number of regions. For instance, when classifying six seating regions, classifier 1 can be designed to classify seats and aisles, and then classifiers 2 and 3 can be designed to classify aisles A, B, and C and seats A, B, and C, respectively.
[0221] For example, oversampling and / or undersampling operations can be performed when training a first machine learning classifier. Oversampling refers to balancing the quantitative differences between different categories in a data set by increasing the number of copies of minority samples or by synthesizing new samples. By adding data to minority samples, the sample distribution between each category becomes more balanced, thereby improving the model's ability to learn minority samples. Undersampling refers to balancing an unbalanced data set by reducing the number of majority samples, thereby bringing the sample amounts between each category closer together. By removing majority samples, the number of samples in each category becomes equal, thereby lowering the model's priority for the majority.
[0222] When the first machine learning classifier determines a valid point, a particular candidate target point may be determined to belong to multiple regions. For example, if the outputs of classifier models Mv,1 and Mv,2 are both 1, and the outputs of the other classifier models are all 0, the candidate target point will be determined to be a valid target point for both seat A and seat B.
[0223] The second machine learning classifier, which determines whether or not there are people in each area, may be any one of the following: SVM, RF, decision tree, Gaussian mixture model, KNN, hidden Markov, or multilayer perceptron. The training steps for the classifier model include the following:
[0224] Perform the following operations across various areas, such as seats A, B, and C, and aisle areas A, B, and C.
[0225] 1, wl,
number
number
number
number
[0226] 2. Read the region features of each frame sliding window generated by steps 1-11 above as input negative sample data, and set all sample labels to 0.
[0227] 3. Positive and negative sample data and labels are combined, and if the number of positive samples is less than the number of negative samples at the time of combination, the positive samples can be resampled. The iteration coefficient is γ = Nn / Np, where Nn and Np are the number of negative and positive samples, respectively.
[0228] 4. Train a second machine learning classifier to obtain the classifier model Mc,l.
[0229] In an application scenario where six areas are to be distinguished—seats A, B, and C in the back row, and aisles A, B, and C—six machine learning classifier models Mc,l can be obtained after performing the above steps. Here, l=1,2,...,6 corresponds to the six areas of seats A, B, and C in the back row, and aisles A, B, and C, respectively. Each machine learning classifier model Mc,l can determine whether or not a person is present in area l. If the area features of a sliding window process are input and the output of the second machine learning classifier model Mc,l is 0, it is determined that no person is present in area l. If the output of the machine learning classifier model Mc,l is 1, it is determined that a person is present in area l. Furthermore, based on this processing flow, it is possible to determine whether or not multiple people are present in a given scenario.
[0230] In an exemplary embodiment, the empty parking lot scenario can be output as a separate category.
[0231] When making a valid point determination by the second machine learning classifier, for a specific candidate target point, it may be determined to belong to multiple regions. For example, when the outputs of classifier models Mc,1 and Mc,2 are both 1 and the outputs of other classifier models are all 0, it is determined that there are people in both seat A and seat B in the scene. When the outputs of all classifier models Mc,l, l = 1, 2,..., 6 are all 0, it is determined that there are no people in the scene.
[0232] Exemplarily, the second machine learning classifier can perform multiple classifications using one classifier, or can perform hierarchical multiple classifications using fewer classifiers than the number of regions.
[0233] Exemplarily, an oversampling operation and / or an undersampling operation can be performed when training the second machine learning classifier. The oversampling refers to balancing the quantity difference between different categories in the data set by increasing the replication of minority samples or synthesizing new samples. By adding data to the minority samples, the sample distribution between categories becomes more balanced, thereby improving the learning ability of the minority sample model. The undersampling refers to balancing an unbalanced data set by reducing the number of majority samples, so that the sample amounts between categories become closer. By deleting majority samples, the number of samples in each category becomes equal, thereby reducing the priority of the model for the majority.
[0234] The method according to the embodiment of the present invention differs from the slow time processing between chirps in radar signal processing in related technologies. In the embodiment of the present invention, by accumulating multi-frame data and performing inter-frame FFT processing in the form of a sliding window, the signal-to-noise ratio between personnel targets and static strong clutter can be improved. In the embodiment of the present invention, by performing region segmentation using a machine learning model, the use of information is richer than spatial segmentation and region determination strategies based on predetermined rules, and it can be fully automated without human intervention, thus greatly reducing the pressure of parameter tuning.
[0235] The embodiments of this application are applicable not only to human target detection and location identification scenarios inside vehicles, but also to other similar application scenarios such as personnel detection indoors and personnel detection in factory buildings.
[0236] The above operations may be performed in the CPU or MCU, or in the baseband unit BB. Alternatively, some operations may be performed by the CPU or MCU, and some by the BB. For example, 1D-FFT, 2D-FFT, CFAR, DBF, and DOA may be performed in the BB, while target detection and recognition may be performed in the BB.
[0237] Figures 26A and 26B show the results of DBSCAN clustering performed on all points in multiple experiments in the seat C scene. The black dots in the figures are candidate target points identified as noise by DBSCAN, and the other colored dots are different classification clusters. According to the above example, the classification cluster with the largest number of samples (shown by the black border in the figure) is selected and represented as a valid point in the seat C area, and the other points are considered invalid points in seat C.
[0238] Figures 26C and 26D show the results of DBSCAN clustering and labeling, taking into account all points from multiple experiments in seat A, B, and C. The classification cluster with the most valid points in seat A, B, and C is indicated by the black border in the figure.
[0239] Compared to conventional technologies, the method according to the embodiment of this application employs a new inter-frame storage method and performs coherent storage for the frequency of human respiration, thereby improving the signal-to-noise ratio between personnel targets and statically strong clutter. Furthermore, by employing a clustering method to remove abnormal outliers, self-labeling of radar point cloud samples is achieved. By combining machine learning classifiers such as SVM, random forest, and mixture Gaussian models, more precise region partitioning than predetermined logic can be achieved, enabling accurate classification of seats corresponding to radar processed point clouds. By improving the signal-to-noise ratio, unsupervised clustering self-labeling, and precise machine learning classification, it is possible to improve the accuracy of determining the presence or absence and location of personnel in enclosed environments such as cabins. In addition, compared to rule-based region determination, the model construction and processing flow reduces the need for rule design, mitigates the difficulty of parameter tuning, and improves the convenience and applicability of the method.
[0240] The division of the steps in the above method is for the sole purpose of clarifying the explanation, and when implemented, they can be combined into a single step or some steps can be divided into multiple steps, as long as the same logical relationship is included, and both are within the scope of this patent. Adding non-essential modifications to the algorithm or flow, or introducing non-essential designs, without changing the core design of the algorithm or process, is also within the scope of this patent.
[0241] In embodiments of the present application, an integrated circuit is provided which processes a digital signal based on the target detection method according to the embodiments of the present application, including sequentially connected radio frequency modules, analog signal processing modules, and digital signal processing modules, wherein the radio frequency modules are used to generate radio frequency transmission signals and receive radio frequency reception signals, the analog signal processing modules are used to reduce the frequency of the radio frequency reception signals so that an intermediate frequency signal is obtained, and the digital signal processing modules are used to convert the intermediate frequency signals from analog to digital, so that the objective of target detection is achieved. For example, the integrated circuit may be a millimeter-wave radar chip (chip or die). Here, the digital processing modules may include a baseband module and a main control module, etc. Each module may be configured to perform the corresponding steps of the target detection method according to the embodiments.
[0242] In some optional embodiments, the integrated circuit may be an AiP (Antenna-In-Package) chip structure, an AoP (Antenna-On-Package) chip structure, or an AoC (Antenna-On-Chip) chip structure.
[0243] According to several other embodiments of the present invention, electromagnetic wave sensors have been further proposed. These electromagnetic wave sensors may include an antenna and the integrated circuit described above. Here, the integrated circuit is electrically connected to the antenna and is used for transmitting and receiving electromagnetic wave signals. For example, the electromagnetic wave sensor may include a support, the integrated circuit described in any of the embodiments above, and an antenna. The integrated circuit may be provided on the support. The antenna may be provided on the support, or integrated with the integrated circuit as an integrated device and provided on the support (i.e., in this case, the antenna may be an antenna provided in an AiP, AoP, or AoC structure). Here, the integrated circuit is connected to the antenna (i.e., in this case, the sensing chip or integrated circuit does not have an integrated antenna such as a conventional SoC) and is used for transmitting and receiving electromagnetic wave signals. Here, the support may be a printed circuit board (PCB), and the corresponding transmission line may be PCB wiring.
[0244] In this application, electromagnetic waves may include radio waves and light waves. Radio waves include short waves, medium waves, long waves, and microwaves. Microwaves include centimeter waves (i.e., electromagnetic waves from 3 GHz to 30 GHz, such as electromagnetic waves in the 3.1 GHz to 10.6 GHz frequency band and electromagnetic waves in the 24 GHz frequency band) and millimeter waves (i.e., electromagnetic waves from 30 GHz to 300 GHz, such as electromagnetic waves in the 60 GHz frequency band and electromagnetic waves in the 77 GHz frequency band (77 GHz to 81 GHz, etc.)). Light waves may include ultraviolet light, visible light, infrared light, and lasers. Here, the electromagnetic wave frequency band of lasers is (3.846 to 7.895) * 10^5 GHz, meaning that lasers are included in the frequency bands of ultraviolet light and some of the visible light.
[0245] In embodiments of the present application, a device is provided that includes a main body of the device and the electromagnetic wave sensor provided on the main body of the device. Here, the electromagnetic wave sensor is used for target detection and / or communication in order to provide reference information for the operation of the main body of the device.
[0246] Embodiments of the present application further provide an electronic device that can be represented in the form of a general-purpose computer. The assembly of the electronic device may include, but is not limited to, at least one processing unit, at least one storage unit, a bus connecting different system assemblies (including the storage unit and the processing unit), a display unit, and the like. The storage unit stores program code. The program code may be executed by the processing unit so that the processing unit performs the methods of various exemplary embodiments of the present application described herein. The storage unit may include, for example, a readable medium in the form of a volatile storage unit such as a random access storage unit (RAM) and / or a cache memory storage unit, and may further include a read-only storage unit (ROM).
[0247] A storage unit may include a program / utility having a set (at least one) of programming modules. Such programming modules may include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. One or more of these examples may include the implementation of a network environment.
[0248] A bus can represent one or more of several bus structures, including a storage unit bus or storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus that uses any of the bus structures in a plurality of bus structures.
[0249] The electronic device may communicate with one or more external devices (e.g., keyboards, pointing devices, Bluetooth devices, etc.), one or more devices that enable a user to interact with the electronic device, and / or any device that enables the electronic device to communicate with one or more other computing devices (e.g., routers, modems, etc.). This communication can be carried out via an input / output (I / O) interface. The electronic device may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs) such as the Internet, and / or public networks) via a network adapter. The network adapter can communicate with other modules of the electronic device via a bus. Although not shown in the diagram, other hardware and / or software modules, including but not limited to microcode, device drives, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, may be used in combination with the electronic device.
[0250] For example, the electronic device according to the embodiment of the present application may further include a main body of the device and an electromagnetic wave sensor provided on the main body of the device as described in any of the above embodiments. Here, the electromagnetic wave sensor is used to realize functions such as target detection and / or wireless communication.
[0251] Specifically, based on the above embodiments, in any embodiment of the present application, the electromagnetic wave sensor may be provided either outside or inside the main body of the device. Furthermore, in any other embodiment of the present application, the electromagnetic wave sensor may be partially provided inside the main body of the device and partially provided outside the main body of the device. The embodiments of the present application are not limited thereto and may be specifically determined on a case-by-case basis.
[0252] In any embodiment, the device body may be a member or product applicable to fields such as smart cities, smart homes, transportation, smart homes, household electrical appliances, security monitoring, industrial automation, in-cabin detection (such as smart cockpits), medical devices, and healthcare. For example, the device body may be a smart transportation device (such as cars, bicycles, motorcycles, ships, subways, trains, etc.), a security device (such as a camera), a liquid level / flow detection device, a smart wearable device (such as a bracelet, glasses, etc.), a smart home device (such as a cleaning robot, a door lock, a TV, an air conditioner, a smart light, etc.), various communication devices (such as mobile phones, tablet computers, etc.), and a barrier gate, a smart traffic signal, a smart sign, a traffic camera, and various industrial robot arms (or robots), etc. It may also be various instruments used to detect biometric parameters such as biometric detection inside the cabin of a car, indoor personnel monitoring, smart medical devices, household electrical appliances, etc., and various devices equipped with such instruments.
[0253] In an embodiment of the present application, a non-temporary computer-readable storage medium storing computer-readable instructions is further provided. When the instructions are executed by a processor, the processor executes the above-described feeder variable-length compensation method.
[0254] From the description of the above embodiments, for the ease of understanding by those skilled in the art, the exemplary embodiments described herein can be realized by software and can be realized by combining the software with the necessary hardware. The technical means according to the embodiments of the present application can be embodied in the form of a software product. The software product can be stored in a non-volatile storage medium (which may be a CD-ROM, a USB memory, a mobile hard disk, etc.) or on a network, and by including a plurality of instructions, a single computing device (which may be a personal computer, a server, or a network device, etc.) can execute the above method based on the embodiments of the present application.
[0255] Software products may employ any combination of one or more readable media. The readable media may be readable signal media or readable storage media. The readable storage media may be, but are not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any combination thereof. More specific examples (not exhaustive) of readable storage media include electrical connections with one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above.
[0256] A computer-readable storage medium may include data signals that propagate within the baseband or propagate as part of a carrier wave carrying readable program code. These propagated data signals may take multiple forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The readable storage medium may be any other readable medium than that which transmits, propagates, or transmits a program intended for use by or in combination with an instruction execution system, apparatus, or device. The program code contained in the readable storage medium may be transmitted by any suitable medium, including but not limited to wireless, wired, cable, RF, or any suitable combination thereof.
[0257] Program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as the C language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on a user device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device or to an external computing device via any type of network, including a local area network (LAN) or a wide area network (WAN) (for example, by connecting via the Internet using an Internet service provider).
[0258] The computer-readable medium described above contains one or more programs. When one or more of these programs are executed by the device, the computer-readable medium performs the functions described above.
[0259] Those skilled in the art will understand that each of the above modules may be distributed within the apparatus according to the description of the embodiment, and may be modified accordingly in one or more apparatuses uniquely different from this embodiment. The modules of the above embodiment may be combined into a single module, or further divided into multiple submodules.
[0260] According to embodiments of the present application, a computer program is proposed that includes a computer program or instructions capable of performing the above method when executed by a processor. In any embodiment, the integrated circuit may be a millimeter-wave radar chip. The type of digital function module within the integrated circuit can be specified according to the actual requirements. For example, in the case of a millimeter-wave radar chip, the data processing module is used for range-dimensional Doppler conversion, velocity-dimensional Doppler conversion, constant false alarm detection, wave arrival direction detection, point cloud processing, etc., and is used to acquire information such as the range, angle, velocity, height, micro-Doppler motion characteristics, shape, size, surface roughness, and dielectric properties of a target.
[0261] Furthermore, since wireless devices can perform functions such as target detection and / or communication by transmitting and receiving wireless signals, they can provide detected target information and / or communication information to the main unit of the device, and in turn, support or control the operation of the main unit of the device.
[0262] For example, when the above-mentioned device is applied to an advanced driver-assistance system (ADAS), the wireless device (such as millimeter-wave radar) as an in-vehicle sensor can support the ADAS system and enable applications such as adaptive cruise control, automatic emergency braking (i.e., AEB), blind spot detection warning (i.e., BSD), lane change assist warning (i.e., LCA), reverse assist warning (i.e., RCTA), parking assist, rear vehicle warning, collision avoidance, pedestrian detection, and in-cabin biometric detection (i.e., CPD).
[0263] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described, but as long as these combinations of technical features are inconsistent, they shall be considered to fall within the scope described herein.
[0264] The above embodiments merely represent preferred embodiments and the technical principles used in the present invention, and while their descriptions are specific and detailed, they should not be understood as limiting the scope of the patent. Those skilled in the art will be able to make various obvious modifications, readjustments, and substitutions without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail by the above embodiments, the present invention is not limited to these embodiments and may include many other equivalent embodiments without departing from the spirit of the invention, and the scope of protection of the patent for the present invention is determined by the appended claims.
[0265] This application is filed based on the Chinese patent application with application number "202310913653.8" and filing date of July 24, 2023, and claims priority from the said Chinese patent application, the entire contents of the said Chinese patent application are incorporated into this application by reference. [Claim 1] The echo signal is subjected to FFT processing so that a feature vector of the echo signal is generated, This includes processing the feature vectors based on a machine learning model so that target detection in the region of interest (ROI) can be achieved, A target detection method characterized by the following: [Claim 2] The feature vector includes characterizing the energy features of the echo signal in different range units, or The feature vector includes characterizing the energy features of the echo signal in different Doppler dimensions of different range units. The target detection method according to feature 1. [Claim 3] The feature vectors include characterizing the energy features of the echo signal in different Doppler dimensions of different range units, Performing FFT processing on an echo signal is The echo signal is subjected to FFT processing so that a range Doppler spectrum is generated, This includes performing beamforming based on the range Doppler spectrum so that a feature vector of the echo signal can be obtained, The target detection method according to feature 2. [Claim 4] Performing FFT processing on the echo signal so that a range Doppler spectrum is generated is The echo signal is subjected to range-dimensional FFT processing, This includes sliding windowing the results obtained by range-dimensional FFT processing to obtain the aforementioned range Doppler spectrum, and then processing the windowed data with 2D FFT (Doppler FFT). The duration of the sliding window is at the same level as the period of the target periodic motion. The target detection method according to feature 3. [Claim 5] Performing beamforming based on the range Doppler spectrum is achieved by the following equation:
number
number
number
number
Claims
1. The echo signal is subjected to FFT processing so that a feature vector of the echo signal is generated. This includes obtaining target detection point information by processing the feature vector based on a machine learning model, and realizing target detection in a region of interest (ROI) based on the target detection point information. The target detection point information includes distance information and angle information, The feature vector includes characterizing the energy features of the echo signal in different range units, or The feature vector includes characterizing the energy features of the echo signal in different Doppler dimensions of different range units. A target detection method characterized by the following:
2. The feature vectors include characterizing the energy features of the echo signal in different Doppler dimensions of different range units, Processing the echo signal with FFT is The echo signal is subjected to FFT processing so that a range Doppler spectrum is generated, This includes performing beamforming based on the range Doppler spectrum so that a feature vector of the echo signal can be obtained, The target detection method according to claim 1.
3. Performing FFT processing on the echo signal so that a range Doppler spectrum is generated is The echo signal is subjected to range-dimensional FFT processing, The process includes: sliding windowing the results obtained by range-dimensional FFT processing to obtain the range Doppler spectrum, and then processing the windowed data with 2D FFT (Doppler FFT), The duration of the sliding window is at the same level as the period of the target periodic motion. The target detection method according to feature 2.
4. Performing beamforming based on the range Doppler spectrum is achieved by the following equation: [Math 1] However, P(r,v,b) is the energy feature of the v-th Doppler unit of the r-th range unit based on beam b, sv(c,b) is the steering vector of beam b in virtual channel c, x(c,r,v) is the data corresponding to the v-th Doppler unit of the r-th range unit in the virtual channel c in the range Doppler spectrum, Nc is the total number of beams. The target detection method according to feature 2.
5. The feature vectors include characterizing the energy features of the echo signal in different Doppler dimensions of different range units, Performing FFT processing and beamforming processing on the echo signal is The echo signal is subjected to range FFT processing, Performing beamforming on the results obtained by range FFT processing, This includes converting the beamforming results into a sliding window and then performing 2D FFT processing on the windowed data. The target detection method according to claim 1.
6. The feature vector includes characterizing the energy features of the echo signal in different range units, Performing FFT processing and beamforming processing on the echo signal is The echo signal is subjected to range FFT processing, This includes performing beamforming on the results obtained by range FFT processing. The target detection method according to claim 1.
7. Applying beamforming to an echo signal is, Identifying a target direction that points towards the aforementioned region of interest, and generating a steering vector based on the aforementioned target direction, This includes performing beamforming on the echo signal based on the steering vector, A target detection method according to any one of claims 2 to 6.
8. Identifying the target direction that aligns with the aforementioned area of interest is, This includes determining the target direction based on the azimuth and elevation angles of the region of interest relative to the radar, The target detection method according to feature 7.
9. Identifying the target direction that aligns with the aforementioned area of interest is, To acquire target measurement data in the aforementioned area of interest, This includes clustering based on the aforementioned target measurement data, thereby setting the azimuth and elevation angles of the clustering center points in each of the aforementioned regions of interest to the target direction, The target detection method according to feature 7.
10. Generating a steering vector based on the aforementioned target direction is achieved by the following equation: [Math 2] However, sv(c,b) is the steering vector of beam b in virtual channel c, λ is the central wavelength of the radar detection signal. xc and zc are the positions corresponding to the horizontal and vertical directions of the antenna for virtual channel c, respectively. θb is the azimuth angle corresponding to the target direction, φb is the elevation angle corresponding to the target direction. The target detection method according to feature 7.
11. Generating the feature vector of the aforementioned echo signal is, This includes rearranging the energy features obtained by FFT processing and beamforming processing on the echo signal so that the feature vector is obtained according to predetermined rules, A target detection method according to any one of claims 1 to 6.
12. The aforementioned rules are, Data with the same first dimension are arranged continuously according to the gradient direction of the second dimension, and the entire dataset is arranged continuously according to the gradient direction of the first dimension. Or, This includes continuously arranging data that have the same third dimension and the same fourth dimension according to the gradient direction of the fifth dimension, continuously arranging data that have the same third dimension according to the gradient direction of the fourth dimension, and arranging the data as a whole continuously according to the gradient direction of the third dimension. The first and second dimensions are different dimensions in the beam dimension and range dimension, respectively. The third, fourth, and fifth dimensions are different dimensions in the beam dimension, range dimension, and Doppler dimension, respectively. The target detection method according to claim 11.
13. Generating the feature vector of the aforementioned echo signal is, This includes logarithmically normalizing the energy characteristics obtained by FFT processing and beamforming processing on the echo signal. A target detection method according to any one of claims 1 to 6.
14. Logarithmically normalizing the energy features obtained by FFT processing and beamforming processing on the echo signal is This includes obtaining a logarithm of the energy features obtained by FFT processing and beamforming processing on the echo signal, and then linearly normalizing the obtained logarithm according to predetermined upper and lower limit constraints. The target detection method according to claim 13, characterized by the features described above.
15. Obtaining a logarithm of the energy features obtained by FFT processing and beamforming processing on the echo signal, and then linearly normalizing the obtained logarithm by predetermined upper and lower limit constraints, is achieved by the following equation: [Math 3] However, Pnorm(r, v, b) is the result of log-normalization. P(r, v, b) is the energy characteristic of beam b in range unit r and Doppler unit v, a0, b0, Pmax, and Pmin are all predetermined parameters. The target detection method according to feature 14.
16. The aforementioned machine learning model is a deep learning model. Before processing the feature vectors based on the machine learning model, This involves obtaining a training feature vector to characterize the energy features of a signal in different range units, while also having a one-hot encoding label. Processing the training feature vectors so that a training set is obtained by a predetermined data augmentation mode, which includes at least one of random increase of white noise, inversion of data along the Doppler dimension, and translation of data along the range dimension. This further includes training a deep learning model based on the aforementioned training set. A target detection method according to any one of claims 1 to 6.
17. The result of the target detection is a confidence matrix, The elements of the confidence matrix correspond to and indicate at least one of the following pieces of information: the confidence that a target will be detected, the confidence that a target exists in the corresponding region of interest, the confidence that an adult target exists in the corresponding region of interest, and the confidence that a child target exists in the corresponding region of interest. A target detection method according to any one of claims 1 to 6.
18. The feature vector includes energy features that characterize the echo signal in different range units within a predetermined range range and / or predetermined Doppler range. A target detection method according to any one of claims 1 to 6.
19. The aforementioned machine learning model includes a neural network model, Processing the feature vector using a machine model is A predetermined number of frame data are normalized, and the normalized results are input into a pre-trained neural network model, or This includes normalizing a predetermined number of frame data, converting the normalized result into a complex number, and inputting the complex number, or the real part or imaginary part of the complex number, into the neural network model. The target detection method according to claim 1.
20. Normalizing a predetermined number of frame data includes normalizing them using the following formula: [Math 4] However, Pnorm(r,v,b) is the result obtained after beam b and range unit r chirp v have been normalized. P(r,v,b) is the power of the DBF data for beam b and range unit r chirp v. b0 and a0 are predetermined normalization parameters, Pmax and Pmin are the upper and lower limits of a predetermined P(r, v, b), respectively. The method according to feature 19.
21. The b0 is the mean or median of all possible log2(P(r,v,b)) values, a0 is the difference between the 95th percentile and the 5th percentile of all possible log2(P(r,v,b)) values, divided by 256. The method according to the present invention, characterized by the present invention.
22. Converting the normalized result to a complex number involves performing the conversion using the following formula: [Math 5] However, zreal and zimg are the converted real and imaginary parts, Pnorm(r, v, b) is the result after normalization. φ(r, v, b) is the phase of beam b and range unit r chirp v. The method according to feature 19.
23. The neural network model is a composite neural network model and / or includes a composite convolutional layer, a composite nonlinear activation layer, a pooling layer, a composite fully connected layer, and an absolute value layer. The method according to the feature of 22.
24. Performing FFT processing on the echo signal so that a feature vector of the echo signal is generated is: This includes processing the echo signal with a range-dimensional FFT, acquiring 1D-FFT data, and then processing the 1D-FFT (Range FFT) data in a multi-frame linked manner so that an RD spectrum is obtained as the feature vector. The target detection method according to claim 1.
25. Processing the feature vectors based on a machine learning model so that target detection in the region of interest is achieved is This includes inputting target point cloud data, including coordinate data, into a pre-trained first machine learning classifier so that a first detection result is obtained indicating whether or not a candidate target detection point is a valid candidate target detection point. The first machine learning classifier is one or more of the following: support vector machines, random forests, decision trees, Gaussian mixture models, KNNs, hidden Markovs, and multilayer perceptrons. The target detection method according to feature 24.
26. The data input to the first machine learning classifier further includes one or more of the following: range data, constant false alarm signal-to-noise ratio, amplitude value, Doppler frequency, azimuth intermediate angle, and elevation angle. The method according to the present invention of the present invention.
27. Processing the feature vectors based on a machine learning model so that target detection in the region of interest is achieved is This includes processing target point cloud data, extracting region information of candidate target detection points, and inputting the region information into a pre-trained second machine learning classifier so that a second detection result indicating whether or not a target exists in a predetermined region is obtained. The second machine learning classifier is one or more of the following: support vector machines, random forests, decision trees, Gaussian mixture models, KNNs, hidden Markovs, and multilayer perceptrons. The target detection method according to feature 24.
28. It includes sequentially connected radio frequency modules, analog signal processing modules, and digital signal processing modules, The aforementioned radio frequency module is used to generate a radio frequency transmission signal and to receive a radio frequency reception signal. The aforementioned analog signal processing module is used to perform frequency reduction processing on a radio frequency received signal so that an intermediate frequency signal can be obtained. The aforementioned digital signal processing module is used to process the digital data obtained by analog-to-digital conversion using the target detection method described in claim 1, so as to convert the intermediate frequency signal to analog-to-digital and enable target detection. An integrated circuit characterized by the following features.
29. Support and The integrated circuit according to claim 28 provided on the support, The antenna is provided on the support, or is integrated with the integrated circuit as an integrated device and provided on the support, The integrated circuit is connected to the antenna for transmitting radio frequency transmission signals and / or receiving radio frequency reception signals. An electromagnetic wave sensor characterized by the following features.
30. The main body of the device, The device body includes the electromagnetic wave sensor described in claim 29, The electromagnetic wave sensor is used for target detection and / or communication in order to provide reference information for the operation of the main body of the device. A terminal device characterized by the following features.
Citation Information
Patent Citations
Gesture recognition using sensors
JP2018516365A
Blink detection system, blink detection method
JP2019030582A
Radar device and radar signal processing method thereof
JP2019086464A
Radar system, and radar signal processing method
JP2023001662A
Smart home device using a single radar transmission mode for activity recognition of active users and vital sign monitoring of inactive users
WO2022060369A1