Target detection method, integrated circuit, sensor, device and medium

The target detection method enhances accuracy and efficiency in enclosed spaces by using FFT processing and machine learning for echo signal analysis, addressing low detection rates and false alarms, and enabling precise target identification.

JP2025530620AActive Publication Date: 2025-09-17CALTERAH SEMICON TECH (SHANGHAI) CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2024575826
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-07-24
Filing Date
2024-07-24
Publication Date
2025-09-17
Estimated Expiration
2044-07-24

AI Technical Summary

Technical Problem

Target detection in closed or relatively enclosed spaces, such as car cabins, faces challenges due to low detection rates, many false alarms, and inaccurate angle estimation, particularly for low radar cross section targets like children, exacerbated by strong static clutter and multipath effects.

Method used

A target detection method involving FFT processing of echo signals, followed by a machine learning model for feature vector processing, and an integrated circuit with RF, analog, and digital signal processing modules to enhance detection accuracy and efficiency.

Benefits of technology

Improves detection accuracy and reduces false alarms by leveraging deep learning models to analyze echo signals, enabling precise identification of targets like infants and adults in vehicle cabins, supporting applications like Child Presence Detection and Safety Belt Reminders.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025530620000001_ABST
    Figure 2025530620000001_ABST
Patent Text Reader

Abstract

The present application relates to a technical field of target detection, and discloses a target detection method, an integrated circuit, an electromagnetic wave sensor, an apparatus, and a computer-readable storage medium. The target detection method includes: performing FFT processing on an echo signal to generate a feature vector of the echo signal; and processing the feature vector based on a machine learning model to realize target detection in a region of interest.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] TECHNICAL FIELD Embodiments of the present application relate to the technical field of target detection, and in particular to a target detection method, an integrated circuit, an electromagnetic wave sensor, an apparatus, and a computer-readable storage medium. [Background technology]

[0002] For example, target detection in a closed or relatively enclosed space at short distances, such as inside a car cabin, can present technical challenges, such as low detection rates, many false alarms, and inaccurate angle estimation, due to the effects of strong static clutter and multipath. In particular, in situations where children are left unattended in a car, their radar cross section (RCS) is low, making them more difficult to detect. Furthermore, determining seating areas based on predetermined rules often relies on measuring spatial geometric relationships, designing appropriate rules, and adjusting and optimizing various parameters, which increases the difficulty of device deployment and commissioning. Summary of the Invention

[0003] In the embodiments of the present application, a target detection method, an integrated circuit, an electromagnetic wave sensor, an apparatus, and a computer-readable storage medium are provided that can improve the accuracy and efficiency of target detection.

[0004] According to some embodiments of the present application, a first aspect of the embodiments of the present application provides a target detection method including: performing FFT processing on an echo signal to generate a feature vector of the echo signal; and processing the feature vector based on a machine learning model to achieve target detection in a region of interest.

[0005] According to some embodiments of the present application, a second aspect of the embodiments of the present application provides an integrated circuit that includes a radio frequency module, an analog signal processing module, and a digital signal processing module connected in series, wherein the radio frequency module is used to generate a radio frequency transmission signal and receive a radio frequency reception signal, the analog signal processing module is used to frequency down-process the radio frequency reception signal so as to obtain an intermediate frequency signal, and the digital signal processing module is used to perform analog-to-digital conversion of the intermediate frequency signal, and based on this, there is further provided an integrated circuit that realizes the target detection method described in the first aspect of the embodiments of the present application.

[0006] According to some embodiments of the present application, a third aspect of the embodiment of the present application includes a support, an integrated circuit according to the second aspect of the embodiment of the present application provided on the support, and an antenna provided on the support or integrated with the integrated circuit as an integral device and provided on the support, wherein the integrated circuit further includes an electromagnetic wave sensor connected to the antenna for transmitting radio frequency transmission signals and / or receiving radio frequency reception signals.

[0007] According to some embodiments of the present application, a fourth aspect of the embodiment of the present application further provides a device including a device body and an electromagnetic wave sensor according to the third aspect of the embodiment of the present application provided on the device body, the electromagnetic wave sensor being used for target detection and / or communication to provide reference information for operation of the device body. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is a flowchart of a target detection method according to an embodiment of the present application. [Figure 2] 1 is a schematic diagram of a target detection method according to an embodiment of the present application; [Figure 3] 1 is a flowchart of a target detection method according to an embodiment of the present application. [Figure 4] 1 is a flowchart of a target detection method according to an embodiment of the present application. [Figure 5]1 is a flowchart of a target detection method according to an embodiment of the present application. [Figure 6] 1 is a schematic diagram of a region of interest segmentation for a target detection method according to an embodiment of the present application; [Figure 7] FIG. 2 is a schematic diagram of data rearrangement for a target detection method according to an embodiment of the present application; [Figure 8] 1 is a schematic diagram of the structure of a deep learning model for a target detection method according to an embodiment of the present application; FIG. [Figure 9] 1 is a flowchart of a target detection method according to an embodiment of the present application. [Figure 10] 1 is a schematic diagram of a target detection method according to an embodiment of the present application; [Figure 11] 1 is a schematic diagram of a target detection method according to an embodiment of the present application; [Figure 12] 1 is a schematic diagram of a target detection method according to an embodiment of the present application; [Figure 13] 4 is a schematic diagram of a target detection method according to another embodiment of the present application; [Figure 14] 4 is a schematic diagram of a target detection method according to another embodiment of the present application; [Figure 15] 4 is a schematic diagram of a target detection method according to another embodiment of the present application; [Figure 16] 4 is a schematic diagram of a target detection method according to another embodiment of the present application; [Figure 17] 4 is a schematic diagram of a target detection method according to another embodiment of the present application; [Figure 18] 4 is a schematic diagram of a target detection method according to another embodiment of the present application; [Figure 19] 4 is a schematic diagram of a target detection method according to another embodiment of the present application; [Figure 20] 4 is a schematic diagram of a target detection method according to another embodiment of the present application; [Figure 21] 4 is a schematic diagram of a target detection method according to another embodiment of the present application; [Figure 22] 4 is a schematic diagram of a target detection method according to another embodiment of the present application; [Figure 23] FIG. 2 is a schematic diagram of data rearrangement for a target detection method according to an embodiment of the present application; [Figure 24] 1 is a flowchart of a target detection method according to an embodiment of the present application. [Figure 25] 1 is a flowchart of a target detection method according to an embodiment of the present application. [Figure 26A] 1 is a schematic diagram of the effect of a target detection method according to an embodiment of the present application; [Figure 26B] 1 is a schematic diagram of the effect of a target detection method according to an embodiment of the present application; [Figure 26C] 1 is a schematic diagram of the effect of a target detection method according to an embodiment of the present application; [Figure 26D] 1 is a schematic diagram of the effect of a target detection method according to an embodiment of the present application; DETAILED DESCRIPTION OF THE INVENTION

[0009] While this disclosure describes multiple embodiments, the descriptions are illustrative and not limiting, and it will be apparent to those skilled in the art that many more embodiments and implementations may exist within the scope of the embodiments described in this disclosure. While many possible combinations of features are shown in the drawings and discussed in specific embodiments, many other combinations of the disclosed features are possible. Except as expressly limited, any feature or element of any embodiment can be used in combination with or substituted for any other feature or element of any other embodiment.

[0010] The present disclosure includes and contemplates combinations of features and elements known to those skilled in the art. The embodiments, features, and elements previously disclosed in the present disclosure may be combined with any conventional features or elements to form a unique inventive solution as defined by the claims. Any feature or element of any embodiment may be combined with features or elements from other inventive solutions to form another unique inventive solution as defined by the claims. Therefore, it should be understood that any feature shown and / or discussed in the present disclosure may be implemented individually or in any suitable combination. Therefore, the embodiments are not subject to any other limitations, except those based on the appended claims and their equivalent substitutions. Furthermore, various modifications and variations may be made within the scope of protection of the appended claims.

[0011] Additionally, when describing representative embodiments, the specification may have previously presented a method and / or process as a particular sequence of steps. However, to the extent that the method or process does not depend on the particular order of steps described herein, the method or process should not be limited to the particular order of steps described. As one of ordinary skill in the art will understand, other orders of steps are possible. Thus, the particular order of steps described in the specification should not be construed as limiting the claims. Additionally, claims to the method and / or process are not limited to performing them in the order in which they are written. One of ordinary skill in the art will readily understand that these orders can be changed and still be within the spirit and scope of the embodiments of the present application.

[0012] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the specification of the present disclosure is for the purpose of describing examples only and is not intended to be limiting of the disclosure.

[0013] In some embodiments, target detection in a sealed or relatively sealed spatial region can be achieved by performing region determination based on single-frame high-speed and low-speed time processing, CFAR (constant false alarm) detection, and point clouds obtained from angle measurement. Alternatively, target detection can be achieved by performing single-frame high-speed time processing followed by operations such as DoA (direction of arrival) estimation using an algorithm such as Capon to form a range-azimuth (RA) map and then extracting, recognizing, and determining features such as output within a predetermined region. Alternatively, target detection can be achieved by filtering the received echo phase and then detecting human breathing and heartbeat. For example, when target detection is achieved by performing region determination based on single-frame high-speed and low-speed time processing, CFAR detection, and point clouds obtained from angle measurement, 1D-FFT (e.g., range-dimensional FFT) processing on the echo signal can be performed, followed by time correlation analysis based on the echo data, thereby achieving human target discrimination. Position determination can then be performed by combining point cloud information, ultimately achieving the goal of target detection within the cabin.

[0014] In some embodiments, the target detection flow is shown in FIG. 1. A radar (taking target detection by a frequency modulated continuous wave (FMCW) radar as an example) processes echo signals received by an antenna through operations such as analog-to-digital conversion and sampling. Then, a main control module (e.g., a microcontroller unit (MCU)) and a baseband (BB) module (e.g., a baseband chip) provide further processing for target detection. The baseband module sequentially performs range-dimensional DC removal and range-dimensional Fourier transform (Range FFT, which can also be expressed as 1D-FFT or fast time-dimensional FFT). The baseband module then sends a signal to the main control module, causing the main control module to perform multi-frame accumulation. After the data has been accumulated for a predetermined number of frames, such as 128 frames, the main control module transfers the multi-frame accumulated data from the CPU static random-access memory (SRAM) to the baseband module, where it subsequently performs Doppler DC removal, frame-to-frame Fourier transform (Frame FFT), and inter-channel accumulation (non-coherent accumulation, correlation accumulation, etc.) processing in the baseband module. Then, the baseband module performs constant false-alarm rate (CFAR) detection, such as peak-based CFAR detection, based on the noise obtained by noise estimation (such as using a noise variance estimator), and finally performs angle detection and target classification processing to obtain target range, velocity, and / or angle information.

[0015] Here, the relevant steps of FIG. 1 can be implemented as follows:

[0016] Range-dimension DC removal applies to data corresponding to each chirp in each frame of data. In this case, range-dimension DC removal specifically involves calculating an average along the fast time dimension for the data collected by each receive (RX) channel of each chirp, and then subtracting the DC component from all sampling points of each RX channel.

[0017] The range-dimensional FFT can be realized by windowing the data from which the range-dimensional DC has been removed and then performing a 1D-FFT.

[0018] In the target detection process, the data for each chirp in each frame of data obtained by the radar performing ADC or sampling based on the echo signal is processed as described above and sent to the main control module. The data can be sent to the main control module by transferring it to the CPU SRAM via Direct Memory Access (DMA).

[0019] Doppler DC removal can be achieved by calculating a complex average along the frame dimension for the multi-frame data stored in the main control module (i.e., averaging data from different frames in the same range unit (bin) of the same channel) so that the average value of different ranges of different channels is obtained as the DC component of the corresponding range unit of the corresponding channel, and then subtracting the respective DC components for each range unit of each channel.

[0020] Inter-frame Fourier transform can be implemented by windowing the data from which Doppler DC has been removed along the frame dimension and performing a 2D-FFT along the frame dimension. In this case, range-Doppler information for each transmit and receive channel is obtained. For example, in some embodiments, this can be implemented as follows: After the echo signal is processed by a range-dimensional FFT to obtain the range-Doppler spectrum, the results obtained by the range-dimensional FFT processing are subjected to a sliding window, and a 2D-FFT is performed on the windowed data. Here, the length of the sliding window and the corresponding period of the target's periodic motion have the same time length. In this way, performing digital signal processing on multi-frame data that is close to the period of the target's periodic motion can be combined with the characteristics of the target's motion period to better detect targets. This helps to avoid errors and omissions while also eliminating the effects of multipath and other factors, thereby improving the accuracy of target detection.

[0021] Non-coherent accumulation can be achieved by the following equation:

[0022]

number

number

[0023] where P(r,v) and A(r,v) are the power and amplitude of the rth range unit and the vth Doppler unit, respectively, and s(r,t,a,v) is the complex value obtained by inter-frame Fourier transform of the signal echo of the rth range unit and the vth Doppler unit in the tth transmit channel and the ath receive channel.

[0024] Constant false alarm detection based on the noise floor obtained by noise estimation has been described in conventional CFAR techniques and will not be described here.

[0025] The angle detection can be realized by performing azimuth-dimensional digital beam forming (DBF) and DOA and elevation-dimensional DBF and DOA on the target point obtained by CFAR. Here, DBF and DOA are described in the conventional DBF and DOA technology, and the description thereof will be omitted here.

[0026] Target classification can be realized by counting target points detected in each frame (or a predetermined amount of data, i.e., each time) according to a predetermined (or divided) area unit design, and judging the counted results according to predetermined rules, thereby obtaining whether or not a physical target of interest (such as an adult, child, infant, or pet) exists in the process and identifying its location area.

[0027] 1 is merely an example of a radar having a baseband module and a main control module. That is, operations such as frame data accumulation and target classification (which can be achieved by Region HIST classification, etc.) are performed in the main control module, while other operations such as range dimension DC removal and Doppler dimension DC removal are performed in the baseband module. In other embodiments, frame data accumulation is performed in the baseband module, and Doppler dimension DC removal is performed in the main control module.

[0028] 1 illustrates caching 128 frames of data as an example, but in other embodiments, 56 frames of data, 32 frames of data, etc. may be cached. The embodiments of the present application are not limited to this, and may be specified depending on the needs and hardware support capabilities. As a result, the cached multi-frame data is subsequently subjected to Doppler DC removal and inter-frame FFT to obtain a range-Doppler (RD) spectrum.

[0029] In the flow shown in FIG. 1, some steps such as range-dimension DC removal, range-dimension DC removal, and windowing may be skipped or simplified. For example, non-coherent accumulation may be skipped and CFAR detection may be performed only on specific channels. Also, for example, only some channels may be selected, non-coherent accumulation may be performed, and CFAR detection may be performed only on those channels.

[0030] Furthermore, based on the above means, the embodiments of the present application further propose combining and processing a deep learning model. By deep mining target information using a deep learning model, more accurate target detection can be realized in spatial domains such as the interior of a vehicle cabin, indoors, and factory buildings. This effectively improves the detection rate and reduces the number of false alarm targets, while significantly improving the accuracy of angle estimation. This allows for accurate detection of special or weak targets such as infants in the cabin, and realizes applications such as CPD (Child Presence Detection) and SBR (Safety Belt Reminder). The deep learning model can be combined with the above embodiments in different forms, and different combinations are described below.

[0031] In some embodiments, as shown in FIG. 2, the target detection method includes the following steps:

[0032] In step 101, an FFT process and a beamforming process are performed on the echo signal, so that a feature vector of the echo signal is generated, where the feature vector includes characterizing the energy features of the echo signal in different range units.

[0033] In step 102, the feature vector is processed based on a deep learning model so as to achieve target detection in the region of interest.

[0034] In this way, beamforming processing effectively improves the signal-to-noise ratio between personnel targets and strong static clutter, effectively constructing feature vectors for empty vehicles (with or without interference), scenes with adults, and scenes with children. Deep learning models then deep-mining the target information contained in the feature vectors provides a more accurate means of target detection in spatial domains such as car cabins, indoors, and factory buildings. This effectively improves the detection rate and reduces the number of false alarms. At the same time, the accuracy of angle estimation is significantly improved, allowing for the accurate detection of special or weak targets, such as infants in cabins, enabling applications such as CPD and SBR. At the same time, the introduction of deep learning modules further reduces the amount of digital signal processing required; for example, eliminating the need for CFAR and DOA processing after feature extraction.

[0035] To better understand the embodiment shown in FIG. 2, its steps are described below.

[0036] Step 101 does not limit the implementation of FFT processing and beamforming processing. Note that beamforming processing essentially concentrates the directional gain (which can react energy) of a signal in a specific direction, so beamforming actually corresponds to extracting energy features. That is, different implementation methods of FFT processing and beamforming processing do not affect the extraction of energy features, and therefore do not affect the generation of feature vectors, which include characterizing the energy features of echo signals in different range units.

[0037] As can be seen from the flow diagram shown in FIG. 1 , the target detection process requires beamforming to achieve angle estimation, and subsequent processing such as target classification to output intuitive and accurate results. In the embodiment shown in FIG. 2 , beamforming is used to generate a feature vector, and subsequent angle estimation, target classification, and other processing are performed using a deep learning model, resulting in higher efficiency. Furthermore, rather than directly using time-series data as input, the deep learning model further extracts target information from the echo signal through FFT processing and beamforming processing to form a feature vector with higher dimensions and richer information. This allows the deep learning model to better analyze and process the data, output more accurate results, improve the accuracy and efficiency of target detection, and more accurately detect and recognize targets within the cabin. Here, noise and target reflection signals have different energy characteristics, and using a feature vector characterizing the energy characteristics helps the deep learning model better perform target detection and output more accurate and reliable results.

[0038] The application of the deep learning model involves replacing processes such as CFAR and angle estimation and directly outputting the detection results. That is, based on the flow shown in Figure 1, beamforming is performed after inter-frame FFT, and processes such as correlation accumulation, CFAR, and angle estimation are replaced by deep learning model processing. That is, in some embodiments, as shown in Figure 3, FFT processing and beamforming processing on echo signals can be achieved by the following steps:

[0039] In step 1011, the echo signals are FFT processed to generate a range-Doppler spectrum.

[0040] In step 1012, beamforming processing is performed based on the range-Doppler spectrum.

[0041] As shown in FIG. 1, the FFT process includes a range-dimensional FFT and an inter-frame FFT. After the range-dimensional FFT, beamforming can be performed without waiting for the inter-frame FFT. As mentioned above, the order of the beamforming process is different, and energy features can still be obtained and feature vectors can still be generated. Therefore, in some embodiments, as shown in FIG. 4, the FFT process and beamforming process of the echo signal can be further performed by the following steps:

[0042] In step 1013, the echo signal is subjected to range FFT processing.

[0043] In step 1014, beamforming processing is performed on the results obtained by the range FFT processing.

[0044] In step 1015, the beamforming processing result is made into a sliding window, and the windowed data is subjected to 2D FFT processing. Optionally, the 2D-FFT here may be a sliding window 2D-FFT.

[0045] Note that beamforming processing is already supported after range-dimensional FFT processing, and the subsequent inter-frame FFT is used to better perform target detection through digital signal processing. Furthermore, because the deep learning model has the ability to learn hidden features, good target detection results can still be obtained even without inter-frame FFT, and good results can still be achieved even without inter-frame FFT. Based on this, in some embodiments, as shown in FIG. 5, performing FFT processing and beamforming processing on echo signals can be achieved by the following steps:

[0046] In step 1016, the echo signal is subjected to range FFT processing.

[0047] In step 1017, beamforming processing is performed on the results obtained by the range FFT processing.

[0048] Note that the related processes (including range FFT, frame data accumulation, inter-frame FFT, etc.) performed in the embodiment shown in FIGS. 3 to 5 are described in the embodiment shown in FIG. 2, and will not be described here.

[0049] Of course, the above is merely an exemplary explanation of how to implement FFT processing and beamforming processing, and in some embodiments, processing can be performed using other appropriate methods, which will not be listed here.

[0050] In the embodiments shown in Figures 4 and 5, beamforming is performed based on the result of range-dimension FFT processing, while in the embodiment shown in Figure 3, beamforming is performed based on the result of inter-frame FFT processing. Therefore, the implementation methods of beamforming are also different. Specifically, in the embodiments shown in Figures 4 and 5, beamforming is realized in the range dimension, i.e., beamforming accumulation is performed in different range units, while in the embodiment shown in Figure 3, beamforming may be realized in both the range dimension and the Doppler dimension, i.e., beamforming accumulation is performed in different Doppler units for different range units. For ease of understanding, the implementation of beamforming in the embodiment shown in Figure 3 will be described below as an example.

[0051] In some embodiments, beamforming processing (taking Nc-point beamforming as an example) based on the range-Doppler spectrum obtained by inter-frame FFT can be realized by the following equation:

[0052]

number

[0053] where P(r,v,b) is the energy feature of the rth range unit and the vth Doppler unit based on beam b, sv(c,b) is the steering vector of beam b in virtual channel c, x(c,r,v) is the data corresponding to the rth range unit and the vth Doppler unit in virtual channel c in the range-Doppler spectrum, and Nc is the total number of beams.

[0054] Correspondingly, in beamforming based on data obtained by range-dimensional FFT, x(c,r,v) in the above equation is replaced with x(c,r), i.e., the Doppler dimension is not included. Of course, in other embodiments, other methods can be used to achieve this, but the description will be omitted here.

[0055] In some embodiments of the present application, the steering vector of the beam used in beamforming is not limited to a specific vector and may be specified by a user or automatically generated based on a region of interest. For example, in some embodiments, beamforming can be achieved by identifying a target direction toward a region of interest, generating a steering vector based on the target direction, and performing beamforming based on the steering vector. Here, the identification of the target direction is also not limited to a specific vector. For example, in some examples, identifying the target direction toward a region of interest includes identifying the target direction based on an azimuth angle and an elevation angle of the region of interest relative to the radar. In other examples, identifying the target direction toward a region of interest includes acquiring target measurement data for the region of interest and performing clustering based on the target measurement data to determine the azimuth angle and elevation angle of a clustering center point in each region of interest as the target direction.

[0056] For example, for the interior space of a vehicle shown in Figure 6, if we consider the scenario demand of determining whether people are present in a total of six areas, including the three seats in the back row and the aisle in front of the three seats, and if so, determining whether they are adults or children, we can perform 6-point DBF, i.e., identify six target directions pointing toward each of the three seats and three aisles shown in Figure 6. In this case, P(r,v,b) corresponding to six beams can be obtained. Of course, other directions can be added, such as adding three auxiliary beams pointing toward the other three areas. In this case, 9-point DBF is performed to obtain nine beams.

[0057] The generation of the steering vectors can be achieved by the following equation:

[0058]

number

[0059] where sv(c,b) is the steering vector of beam b in virtual channel c. λ is the center wavelength of the radar detection signal. xc and zc are positions corresponding to the horizontal and vertical directions of the antenna of virtual channel c, respectively. θb is the azimuth angle corresponding to the target direction. φb is the elevation angle corresponding to the target direction.

[0060] In addition, since the coverage area of ​​a radar is usually larger than the range required for target detection by a user, target detection usually occurs in the user's region of interest. Based on this, in the embodiments of the present application, target detection for the region of interest will be mainly described as an example, but this does not mean that the target detection method according to the embodiments of the present application can target only the region of interest, but can target, for example, all detectable regions, and the description thereof will be omitted here and hereinafter.

[0061] It should be noted that the input of the deep learning model also affects the accuracy of the deep learning model. Considering that the target information of the echo signal also has a certain arrangement characteristic. Therefore, in order for the deep learning model to better process the feature vector, the energy features of the feature vector can be further arranged. Based on this, in some embodiments, generating the feature vector of the echo signal can be achieved by rearranging the energy features obtained by FFT processing and beamforming processing of the echo signal according to a predetermined rule, so as to obtain the feature vector. In this way, rearrangement strengthens the data correlation of the feature vector, which helps the model to perform better detection.

[0062] Note that the embodiments of the present application do not limit the specific rules for rearrangement. For example, in some embodiments, the predetermined rule includes arranging data having the same first dimension consecutively according to the gradient direction of the second dimension, and arranging the data consecutively according to the gradient direction of the first dimension as a whole. Here, the first dimension and the second dimension are different dimensions in the beam dimension and the range dimension, respectively. In this case, it is possible to deal with a situation in which only energy features in the range dimension occur. In addition, in some embodiments, the predetermined rule includes arranging data having the same third dimension and the same fourth dimension consecutively according to the gradient direction of the fifth dimension, and arranging data having the same third dimension consecutively according to the gradient direction of the fourth dimension, and arranging the data consecutively according to the gradient direction of the third dimension as a whole. Here, the third dimension, the fourth dimension, and the fifth dimension are different dimensions in the beam dimension, the range dimension, and the Doppler dimension, respectively. In this case, it is possible to deal with a situation in which energy features in the range dimension and the Doppler dimension occur.

[0063] For ease of understanding, the predetermined rules for rearrangement will be explained in conjunction with the rearranged feature vectors shown in FIG.

[0064] As shown in Figure 7, the energy features are sequentially arranged in the v, r, and b dimensions in physical memory. In the figure, v(0) and v(Nv) represent the 0th and Nvth Doppler units (Doppler bins) (a total of Nv+1 Doppler bins), r(0) and r(Nr) represent the 0th and Nrth range units (range bins) (a total of Nr+1 range bins), and b0 represents the 0th beam, for a total of N beams. This array of (Nv+1) x (Nr+1) x N energy features is the feature vector input to the deep learning model.

[0065] As mentioned above, target detection is realized based on hardware, so that it can be adapted to the capabilities of the hardware by further performing data processing to adapt to the capabilities of the hardware.

[0066] For example, in some embodiments, generating a feature vector of an echo signal can be achieved by log-normalizing the energy features obtained by FFT processing and beamforming processing of the echo signal, which helps convert data from floating-point numbers to fixed-point numbers and improves the efficiency of subsequent processing, thereby improving the efficiency of target detection.

[0067] For example, in some embodiments, the logarithmic normalization of the energy features obtained by FFT processing and beamforming processing of the echo signals can be achieved by taking the logarithm of the energy features obtained by FFT processing and beamforming processing of the echo signals, and linearly normalizing the logarithmic result according to predetermined upper and lower limit constraints. That is, based on the logarithmic normalization, the upper and lower limits and linear parameters are further introduced, which further optimizes the data format and is more conducive to forming characteristic data to be subsequently used.

[0068] Here, the logarithm is taken of the energy features obtained by FFT processing and beamforming processing on the echo signal, and the result of the logarithm is linearly normalized according to predetermined upper and lower limit constraints, as realized by the following equation.

[0069]

number

[0070] where Pnorm(r,v,b) is the log-normalization result, P(r,v,b) is the energy feature of beam b in range unit r and Doppler unit v, and a0, b0, Pmax, and Pmin are all predetermined parameters.

[0071] Note that the embodiments of the present application do not limit the specific values ​​of the above-mentioned predetermined parameters. In some possible embodiments, b0 can be set to the mean or median of all possible log2(P(r,v,b)). Alternatively, in some embodiments, a0 is the difference between the 95th percentile and the 5th percentile of all possible log2(P(r,v,b)) values, divided by 256. This allows Pnorm(r,v,b) to be normalized to the range of -127 to 127, i.e., a fixed-point number.

[0072] Note that the above is merely an exemplary description of the related data processing, and other processing methods may be adopted in some embodiments. For example, when obtaining a logarithm, the logarithm of 10 is obtained instead of 2. Here, the data can be normalized to -127 to 127 using the above formula after taking the logarithm of 2, but other methods of obtaining logarithms or other processing methods can achieve the same effect. Alternatively, instead of quantization, conversion to double-precision floating-point or single-precision floating-point numbers such as int32, int16, int8, or uint32, uint16, or uint8 can also improve the efficiency of subsequent processing, so a description thereof will be omitted here.

[0073] Here, the normalization embodiments according to the above embodiments can be realized in various forms. For example, complex number data may be transmitted to the MCU Mem by DMA, and then performed in the MCU, and then realized by a hardware engine in the BB, and then directly transmitted to the MCU Mem by DMA, or a part of the normalization may be realized in the BB, and then transmitted to the MCU Mem by DMA, and then the remaining calculations may be performed in the MCU (the operation for obtaining the norm and the operation for obtaining the logarithm of 2 can be realized in the BB, and the subtraction, division, and saturation truncation can be realized in the MCU), etc., but a description thereof will be omitted here.

[0074] In step 102, the deep learning model used is not limited, and may be any deep learning model such as a convolutional neural network (CNN) or a transformer, and the description thereof will be omitted here.

[0075] The processing process and output results of the deep learning model can be adjusted by configuring the structure of the deep learning model and training the model.

[0076] For example, in some embodiments, the result of target detection is a confidence matrix, with elements of the confidence matrix correspondingly indicating at least one of the following: a confidence that a target will be detected, a confidence that a target is present in the corresponding region of interest, a confidence that an adult target is present in the corresponding region of interest, and a confidence that a child target is present in the corresponding region of interest.

[0077] In some embodiments, the deep learning model is structured as a three-layer convolutional kernel CNN. The type, input, and output size of each layer are shown in FIG. 8. In this case, the input feature dimensions of the deep learning model are 6x64x20, corresponding to 6 beam steering, 64 Doppler units, and 20 range units. Each convolution operation includes a two-dimensional convolutional layer, a batch normalization layer, a ReLu nonlinear activation layer, and a max pooling layer. After three convolution operations, the resulting 128x2 feature vector is flattened to one dimension to obtain a 256x1 one-dimensional feature vector. After random deactivation, the detection result is obtained through a 256x10 fully connected layer. For example, the detection result may be whether there is a person in the car, and if there is a person, the confidence that there are adults and children in each seat and aisle in the car.

[0078] In this case, further processing may be performed based on the output of the deep learning model and output to the user. For example, taking the vehicle interior space shown in FIG. 6 as an example, assume that the output of the deep learning model is a 1×10 vector yi formed by the respective confidence levels indicating whether there is a person in the vehicle, whether there are children in three seats, whether there are children in three aisles, and whether there are adults in three seats. In this case, the following determination can be made: If yi(1)>0, it is directly determined that there is a person in the vehicle without considering the other digits of yi; or if yi(k)=0, k=1, 2, ..., 10, it is also determined that there is no person in the vehicle. If yi(k)>0 due to the existence of k=2, 3, ..., 10, it is determined that there is a child or adult in the corresponding area. If the confidence levels of an adult and a child in seat A and aisle A are all greater than 0, it is determined that there is only an adult in seat A, and thereby seat B and aisle B, and seat C and aisle C are inferred. If yi=[-0.1,7.2,-5,-7.2,-1.5,-1.9,-3.1,-6.4,-6.2,-5.5], it is determined that a child is in seat A, and if yi=[-0.1,-5,-4.2,7.2,-1.5,-1.9,-3.1,6.4,-6.2,-5.5], it is determined that a child is in aisle C and an adult is in seat A. If yi=[-0.1,-5,-4.2,7.2,-1.5,-1.9,-3.1,-6.4,-6.2,5.5], it is determined that an adult is only in seat C. When yi=[0.1,-5,-4.2,7.2,-1.5,-1.9,-3.1,6.4,-6.2,-5.5], the vehicle is determined to be empty. In this case, it is possible to determine the presence of multiple people in the vehicle and significantly reduce false alarms, but the risk of missed detections increases slightly.

[0079] Regarding the deep learning model training method, in some embodiments, before processing the feature vector based on the trained deep learning model, the method further includes the following steps, so as to obtain the target detection result output by the deep learning model, as shown in FIG. 9 :

[0080] In step 103, a training feature vector having one-hot encoding labels and for characterizing the energy features of the signal in different range units is obtained.

[0081] In step 104, the training feature vectors are processed to obtain a training set by a predetermined data augmentation mode, including at least one of randomly increasing white noise, inverting the data along the Doppler dimension, and translating the data along the range dimension.

[0082] In step 105, a deep learning model is trained based on the training set.

[0083] To help those skilled in the art better understand the above embodiments, the training of a deep learning model is described below by way of an example.

[0084] First, the training feature vectors are obtained according to the above-mentioned feature vector obtaining method, and the input feature data set S{P} is constructed by accumulating the processed data from multiple sets of experiments, where each element Pi is the above-mentioned (Nv+1)×(Nr+1)×N array.

[0085] Next, S{P} is labeled based on the experimental conditions, i.e., a label Li corresponding to Pi is generated, and a labeled input feature data set S{P,L} is obtained. In this case, one-hot encoding can be used to generate the label Li. For example, in the car scene shown in Figure 6, we consider determining whether or not there are people in the three rear seats and three aisle areas in the car, while also considering classifying adults and children. If adults are squatting in the three aisle areas, a total of 10 classification problems are addressed. Li is a 1 × 10 vector, each with a value of 0 or 1. The first digit of Li indicates whether or not there are people in the car. If there are no people, the corresponding digit is 1, and if there are people, the corresponding digit is 0. The second to fourth digits indicate whether or not there are children in seats A to C. If there are children, the corresponding digit is 1, and if there are no children, the corresponding digit is 0. The fifth to seventh digits indicate whether or not there are children in aisles A to C. If there are children, the corresponding digit is 1, and if there are no children, the corresponding digit is 0. The 8th to 10th digits indicate whether an adult is present in seats A to C; if so, the digit is 1; if not, the digit is 0. If there is no one in the car, the corresponding Li is [1,0,0,0,0,0,0,0,0,0,0]. If there is an adult in seat A in the car, the corresponding Li is [0,0,0,0,0,0,0,0,1,0,0]. If there is a child in aisle A in the car and an adult in seat B, the corresponding Li is [0,0,0,0,1,0,0,0,1,0]. The first digit of Li, i.e., the indicator of whether there is a person in the car or not, may be omitted; Li is a 1x9 vector, and if all zeros, it indicates that there is no one in the car.

[0086] Next, data augmentation is performed on S{P,L} using one or more methods, such as randomly increasing white noise, randomly inverting along the Doppler dimension, and randomly translating a small number of samples along the range dimension, to obtain an augmented data set S'{P,L}. The number of samples in the augmented data set is significantly increased compared to the remote data set. For example, if S{P,L} contains 5,000 samples each of empty vehicles, the presence of children, and the presence of adults, 10 rounds of data augmentation can be performed to obtain an augmented data set S'{P,L} with 50,000 samples each (it is possible to choose whether to include the original data set S{P,L}). Then, S'{P,L} can be randomly permuted to obtain a validation data set Sv{P,L} with a ratio of α=0.3, and the remaining data can be used as a training data set St{P,L}.

[0087] Then, we train the deep learning model as follows:

[0088] We build an uninitialized CNN model and randomly initialize the parameters of each layer.

[0089] The following steps are repeated 20 times (each time is called an epoch).

[0090] Randomly divide St{P,L} into Nb batches, each with Ns samples, and perform the following steps for each batch:

[0091] Each batch of data is again subjected to one or more data augmentation operations such as a random increase in white noise, a random inversion along the Doppler dimension, or a random small number of translations along the range dimension.

[0092] The augmented data is sent to the CNN for forward calculation to obtain three types of confidence: no person present, adult presence, and child presence, and the corresponding loss CEt is obtained based on the loss function.

[0093] Backpropagation is performed on the CNN network model, and the parameters of each layer of the CNN are updated based on the gradient descent optimization algorithm.

[0094] After performing the above operation for all St{P,L}, the updated CNN is used to perform forward calculations for all samples of Sv{P,L} to obtain the loss CEv for Sv{P,L}. If CEv does not decrease within the set Nes number of cycles, the cycle is terminated early.

[0095] The CNN model with the cycle that minimizes CEv is obtained as the final model output.

[0096] Here, in the training process, the Adam Optimizer can be used as the gradient descent optimizer, and the cross-entropy is selected as the loss function. For the network output, the sigmoid is first used for normalization, and then the cross-entropy is used for loss calculation. The learning rate is set to 0.001, the weight decay is set to 0.0001, Ns is set to 128, and Nes is set to 5.

[0097] Of course, the above is merely an exemplary description of a deep learning model, and in other embodiments, a model with a different structure, a CNN with different parameters, or different training parameters or training methods may be adopted, but the description will be omitted here.

[0098] It should be noted that the above-described embodiments are merely illustrative, and in some embodiments, the assignment of outputs and labels may take other forms than one-hot encoding. For example, when considering only scenes in which there is no person in the vehicle or only one person present, Li = 1, 2, ..., 10 is directly generated as a label. For the generated yi, the digit where its maximum value is located is searched for as a detection and recognition result, but a description thereof will be omitted here.

[0099] Although the coverage area of ​​a radar is large, a user may not necessarily need to detect the entire coverage area. Therefore, in some embodiments, a feature vector is used to characterize the energy characteristics of echo signals at different range units within a predetermined range range and / or a predetermined Doppler range. For example, when performing beamforming, only a portion of the Doppler units and range units, such as 0-32 and 96-127 Doppler units and 10-42 range units, are used. In this way, only the portion of interest to the user is retained to reduce resource occupancy, such as computational complexity and storage, and improve detection efficiency.

[0100] To help those skilled in the art better understand the application of the target detection method according to the above embodiment in the target detection process by radar, the application of the embodiment shown in FIGS. 3 to 5 will be described below as an example.

[0101] For the embodiment shown in FIG. 3, after application to the complete target detection flow, as shown in FIG. 10, i.e., based on the flow shown in FIG. 1, two-dimensional digital beamforming (2D-DBF) processing and DL-based target classification processing are directly performed after frame FFT processing. That is, CFAR processing is not performed after frame FFT, but azimuth and elevation dimensions are performed for each region of interest (or unit section), and 2D-DBF processing is performed so that a number of RD spectra corresponding to the region of interest are obtained. The RD spectra are then input into the constructed deep learning model (DL) for target determination. In some optional embodiments, the region of interest and the RD spectra may have a one-to-many relationship. That is, at least two RD spectra can be obtained based on one region of interest, and the specific number can be adjusted according to actual needs.

[0102] Here, the DL input can be amplitude or power in the linear or dB domain, and can be subjected to operations such as normalization. At the same time, the constructed deep learning model can classify based on the input, and the distinguished categories include the presence or absence of targets, target attributes, and the specific unit section location (i.e., region location) where each target is located.

[0103] For example, in the case of an application scenario in which the rear row area of ​​a vehicle interior is referred to as the region of interest or target area (i.e., the scene shown in FIG. 6 ) and is divided into three seating areas (section units) and three corresponding aisle areas, an azimuth-elevation 2D-DBF is performed on the center positions of the six regions of interest (the regions of interest are also set to the three seating areas) after frame-to-frame FFT to obtain six RD spectra (or three RD spectra) corresponding to the regions of interest, and at least a portion of the six RD spectra (or three RD spectra) is input into the constructed deep learning model to perform operations such as target determination and differentiation. For example, the constructed deep learning model performs classification based on the input to determine whether a biological target is present in the region of interest, and if a biological target is present, it can further determine whether the biological target is an adult, child, infant, pet, etc., and further determine specific location information such as which seat or aisle area the biological target is located in.

[0104] In the flow shown in Figure 10, the target detection method is a processing flow based on model data dual drive, which has low requirements for signal processing, so it can effectively avoid the selection of signal processing parameters, region parameters, CFAR logic and region judgment logic design, etc., which in turn greatly reduces the difficulty of realizing and designing the means.

[0105] The embodiment shown in FIG. 4, after being applied to the complete target detection process, becomes as shown in FIG. 11. As shown in FIG. 11, based on the flow shown in FIG. 1, after the range-dimensional Fourier transform (Range FFT, also called 1D-FFT), 2D-DBF is first performed, and then operations such as frame data accumulation and inter-frame FFT processing are performed. That is, before the inter-frame FFT, 2D-DBF of azimuth and elevation angles is first performed for the center position of the region of interest (e.g., three seating areas). This embodiment can realize a means for parallel processing of data processing and radio wave transmission time, thereby effectively shortening processing time.

[0106] The embodiment shown in Figure 5, after being applied to the complete target detection process, becomes as shown in Figure 12. As shown in Figure 12, based on the flow shown in Figure 1, after a range-dimensional Fourier transform (Range FFT, also known as 1D-FFT), 2D-DBF is subsequently performed, and then a DL-based target classification operation (Complex DL-based Classification) is directly performed. That is, in the entire target detection process, an inter-frame FFT processing operation is not performed, and 2D-DBF is directly performed on the center positions of the region of interest (e.g., three seat areas) after 1D-FFT so that a corresponding number of range frame spectra (e.g., three range frames) are obtained. Then, the range frame spectra are input into a constructed deep learning model (e.g., complex numbers) to perform operations such as target discrimination and judgment.

[0107] In the embodiment shown in Figure 12, the FFT processing step between frames is avoided, thereby effectively reducing the amount of calculation. At the same time, when using a complex deep learning model, a higher processing gain than FFT can be obtained, thereby achieving the goal of improving target detection performance. In addition, this embodiment effectively shortens the processing hierarchy and facilitates efficient adjustment operations.

[0108] As can be seen from the above, by combining 2D-FFT and DL-based classification (DL-based classification) in the flow shown in Figures 4 and 6, the accuracy of target recognition can be effectively improved. At the same time, the 2D-FFT step and the DL-based classification step can be flexibly set between each signal processing step according to actual needs, for example, by setting the 2D-FFT step after the Range FFT or after the Frame FFT.

[0109] In the above embodiments, all FFT processes can be windowed FFTs, and can be implemented by combining SVA and FFT. Alternatively, the multi-frame sliding window FFT can be replaced with other time-frequency transform processes, such as short-time Fourier transform and fractional Fourier transform, to achieve effective target detection. Alternatively, the multi-frame sliding window FFT can be replaced with high-pass FIR, comb FIR, or a filter optimized and synthesized based on an optimization function, to achieve the goal of achieving equivalent or better performance than multi-frame FFT processing with fewer hardware resources. Alternatively, the DBF can be replaced with algorithms such as Capon, MUSIC, ESPRINT, and their variants, to optimize and synthesize beamforming at the center of the region of interest and / or to optimize and synthesize antenna arrays, thereby further improving the system's target detection performance.

[0110] It should be noted that the embodiments shown in Figures 10 to 12 are merely means of target detection combined with multi-frame accumulation techniques, but this does not mean that multi-frame accumulation is incorporated. For example, in the embodiment shown in Figure 12, inter-frame accumulation is omitted, and of course, this will not be described here.

[0111] As shown in Figures 13 to 21, another target detection method is further provided in the present embodiment. By combining inter-frame accumulation with deep learning mode and performing coherent accumulation on the frequency of human breathing, the signal-to-noise ratio between human targets and strong static clutter is improved, effectively constructing feature data for empty vehicles (with or without interference), scenes with adults, and scenes with children. At the same time, by combining with a deep learning model and deep mining the constructed feature data, effective detection of whether a person is present in a vehicle and accurate discrimination between adults and children can be achieved, thereby achieving more accurate target detection and recognition in spatial domains such as inside a vehicle cabin, indoors, and factory buildings. Compared to point cloud information based on signal processing, the method of this embodiment utilizes a higher data feature dimension and richer information, thereby achieving more accurate target detection and recognition within a cabin. Specifically, the method is as follows.

[0112] As shown in FIG. 13, the target detection method includes the following steps:

[0113] In step S11, for each chirp data of each frame, the downlink ADC data can be DC filtered, for example, at the baseband buffer. Optionally, the data collected by each RX channel of each chirp can first be averaged along the fast time dimension, and the DC component can be subtracted from all sampling points of each RX channel. Next, the DC filtered data can be windowed, and then a 1D-FFT can be performed to obtain range dimension information. The 1D-FFT data can then be moved from the baseband buffer to the CPU SRAM cache by DMA. Finally, the above operations can be repeated for different chirps.

[0114] In step S12, when the data cached in the CPU SRAM reaches a certain number of frames, such as 128 frames, the 128 frames of data are moved from the CPU SRAM to the BB by DMA, and the next step is executed.

[0115] First, the multi-frame 1D data is DC filtered, and then the complex average along the frame dimension is calculated for these 128 frames of data to obtain the average values ​​of different ranges of different channels as their DC components, and then the DC components are subtracted from each range unit of each channel.

[0116] The DC filtered data is then windowed along the frame dimension to obtain the range-Doppler complex values ​​x(c,r,v) for each transmit and receive channel, followed by a 2D-FFT calculation along the frame dimension.

[0117] Then, cut out some Doppler units and range units, such as 0~32 and 96~127 Doppler units and 10~42 range units, and find the energy for the complex value x(c,r,v) of each Doppler unit of each range unit of each channel, and then obtain the logarithm to normalize, and the mathematical formula is as follows:

[0118]

number

[0119] In the formula, P(c,r,v) is the normalized "energy" of virtual channel c, range unit r, and Doppler unit v. b0 and a0 are predetermined normalized parameters. Pmax and Pmin are the upper and lower bounds of the predetermined P(c,r,v). Typically, the same b0, a0, Pmax, and Pmin values ​​are selected for different c, r, and v (in a possible variant, different b0, a0, Pmax, and Pmin values ​​can be selected for different c and different v and r partitions). In one possible embodiment, b0 is the mean or median of all possible log2(P(r,v,b)) values, and a0 can be set to the difference between the 95th percentile minus the 5th percentile of all possible log2(P(r,v,b)) values ​​divided by 256. This allows P(c,r,v) to be normalized between -127 and 127.

[0120] Next, P(c,r,v) is rearranged into the format shown in Figure 7, i.e., so that it is consecutive along the v dimension in physical memory, then along the r dimension, and finally along the c dimension. In the figure, v0 and vNv represent the 0th and Nvth Doppler bins (a total of Nv+1 Doppler bins), r0 and rNr represent the 0th and Nrth range bins (a total of Nr+1 range bins), and c0 represents the 0th virtual channel, for a total of Nc+1 virtual channels. This (Nv+1) x (Nr+1) x (Nc+1) array is the input feature data for the deep learning model.

[0121] Finally, the data is sent to the deep learning model, which determines whether there is someone in the cabin, whether they are an adult or a child, and outputs the results.

[0122] The deep learning model may be a CNN or a Transformer. As shown in Figure 14, in one embodiment of the model construction flow, the main steps are as follows:

[0123] In step S21, input feature data is acquired in the same manner as above, and the processed data from multiple experiments is accumulated to construct an input feature data set, where each element Pi is an (Nv+1) × (Nr+1) × (Nc+1) array, and its definition is the same as in step 2.c) above.

[0124] In step S22, S{P} is labeled based on the experimental conditions, that is, a label Li corresponding to Pi is generated, and a labeled input feature data set S{P, L} is obtained.

[0125] In step S23, data augmentation is performed on S{P,L} using one or more of the following methods: randomly increasing white noise, randomly flipping along the Doppler dimension, and randomly translating a small number of samples along the range dimension, to obtain an augmented data set S'{P,L}. The number of samples in the augmented data set is significantly increased compared to the remote data set. For example, if S{P,L} contains 5,000 samples each of empty vehicles, the presence of children, and the presence of adults, 10 rounds of data augmentation can be performed to obtain an augmented data set S'{P,L} with 50,000 samples each of each type (it is possible to choose whether to include the original data set S{P,L}).

[0126] In step S24, S'{P,L} is randomly rearranged to obtain S'{P,L} with a ratio of α=0.3 as the validation data set Sv{P,L}, and the remaining data is the training data set St{P,L}.

[0127] In step S25, a deep learning model is trained based on St{P, L} and Sv{P, L}.

[0128] Optionally, the main steps of one embodiment are as follows:

[0129] First, we build an uninitialized CNN model and randomly initialize the parameters of each layer.

[0130] Next, the following steps are repeated 20 times (each time is called an epoch). First, St{P,L} is randomly divided into Nb batches, each containing Ns samples. The following steps are performed for each batch: First, one or more data augmentation processes are performed on each batch of data, such as randomly increasing white noise, randomly inverting along the Doppler dimension, and randomly translating a small number of samples along the range dimension. Next, the augmented data is sent to the CNN for forward calculation to obtain three confidence levels: absence of a person, presence of an adult, and presence of a child. The corresponding loss CEt is obtained based on the loss function. Finally, backpropagation is performed on the CNN network model, and the parameters of each layer of the CNN are updated based on the gradient descent optimization algorithm. After performing the above operations on all St{P,L}, the updated CNN is used to perform forward calculation on all samples of Sv{P,L} to obtain the loss CEv for Sv{P,L}. If CEv does not decrease after the set number of cycles, the cycle is terminated early.

[0131] Next, we obtain the CNN model with the cycle that minimizes CEv as the final model output.

[0132] In addition, in the model training process, in a possible embodiment, the Adam Optimizer can be used as the gradient descent optimizer, and the cross-entropy can be selected as the loss function, the learning rate is set to 0.001, the weight decay is set to 0.0001, Ns is set to 128, and Nes is set to 5.

[0133] A possible CNN structure is shown below, as shown in Figure 15. This shows a CNN with three layers of convolution kernels, with the type of each layer and the input and output sizes also shown. Here, the input feature dimensions are 16x64x20, corresponding to 16 virtual channels, 64 Doppler units, and 20 range units. Each convolution operation includes a two-dimensional convolution layer, a batch normalization layer, a ReLu nonlinear activation layer, and a max pooling layer. After three convolution operations, the resulting 128x2 feature vector is flattened to one dimension to obtain a one-dimensional 256x1 feature vector. After random deactivation, the vector is passed through a 256x10 fully connected layer to obtain three confidence levels: empty vehicle, presence of a child, and presence of an adult. These confidence levels can be calculated using softmax to obtain the maximum value as the classification of the input, i.e., it can be determined to be one of the categories: empty vehicle, presence of a child, or presence of an adult.

[0134] Note that the method in this embodiment differs from the slow time processing between chirps in conventional radar signal processing in that it accumulates multi-frame data and performs inter-frame FFT processing using a sliding window.

[0135] 2. The method in this embodiment proposes a detection and classification flow using frame-to-frame FFT processing and CNN, without the need to acquire point clouds using traditional signal processing such as CFAR or DoA for processing.

[0136] 3. The method in this embodiment details the construction of CNN input features for the first time, and notable points and possible changes include:

[0137] 1. When calculating P(c,r,v), we can skip the log2 operation or adopt the log10 or ln operation.

[0138] 2. When calculating P(c,r,v), the complex data can be sent to the MCU memory by DMA and then performed in the MCU, or it can be realized by a hardware engine in the BB and then directly sent to the MCU memory by DMA, or part of it can be realized in the BB, sent to the MCU memory by DMA, and the remaining calculations can be performed in the MCU (the norm calculation operation and log2 operation can be realized in the BB, and the subtraction, division and saturation truncation can be realized in the MCU).

[0139] 3. P(c, r, v) may be double-precision floating-point or single-precision floating-point without quantification, or may be quantified in the form of int32, int16, int8 or uint32, uint16, uint8.

[0140] 4. The data array may be continuous along the r dimension, then along the v and c dimensions (as shown in FIG. 23), or along the vcr or cvr dimensions.

[0141] 5. In the deep learning model, a virtual channel can be selected as the model channel, and the virtual channel dimension can be tiled to adopt the input of a single-channel deep learning model. The input of the deep learning model is 1 x 64 x 320, which includes 64 Doppler units and 20 range units of 16 channels tiled as 320.

[0142] Regarding channels, to avoid the impact of noisy channels or channels with poor radio frequency simulation characteristics on the processing results, only some channels can be selected. For example, 13 out of 16 channels (4 transmit and 4 receive) are used to construct 13 x 64 x 20 CNN input characteristic data to determine whether a vehicle is empty, whether it is a child, or whether it is an adult.

[0143] 7. The data after the inter-frame FFT may be truncated to save memory and reduce the amount of calculation, or may not be truncated for optimal performance. If truncation is required, it can be performed after the 1D-FFT, 2D-FFT, or when calculating P, or a combination of these.

[0144] As shown in FIG. 16 , the target detection method in the embodiment of the present application may further adopt the processing flow shown in FIG. 16 , that is, perform non-coherent accumulation of the data after inter-frame FFT according to the channel to obtain an array, perform the same operation as the above step of cutting out some Doppler units and range units, and then send it to the deep learning model as input feature data for processing to obtain the classification result, and select the number of input channels of the corresponding deep learning model as 1.

[0145] As shown in Fig. 17, the target detection method in the embodiment of the present application may further adopt the processing flow shown in Fig. 17. After performing non-coherent accumulation on the data after inter-frame FFT according to the channel to obtain two-dimensional data, and using CFAR to select only the energy of points exceeding a predetermined threshold, and adopting the same operation as the above step of cutting out some Doppler units and range units, fill in the range units and Doppler units that do not exceed the threshold as input feature data, and then send the input feature data to a deep learning model for processing to obtain the classification result, and select the number of input channels of the corresponding deep learning model as 1.

[0146] It is also possible to combine the above two embodiments of Figures 16 and 17. That is, after performing CFAR on the data after inter-frame FFT according to the channel, only the energy of points exceeding a predetermined threshold is selected for each channel, and the same operation as the above step of cutting out some Doppler units and range units is adopted, and then the range units and Doppler units that do not exceed the threshold are filled in as input feature data, and then the data of all channels are arranged in the same manner as the rearrangement of P(c, r, v), and the input feature data is sent to a deep learning model for processing to obtain the classification result, and the number of input channels of the corresponding deep learning model is selected as the number of virtual channels to be used.

[0147] The deep learning model in this embodiment can be selected from other CNNs or Transformers as long as it is adapted to the input. As shown in Figures 18 and 19, it can be based on a CNN model modified by ResNet18 and its extension of the Sequential module. The output here can be one of two types: vacant and non-vacant.

[0148] The two diagrams in Figure 20 show example CNN input feature maps corresponding to an adult and a child, respectively. The left diagram shows the input features for an adult, and the right diagram shows the input features for a child. Each sub-plot corresponds to the data for each channel. The horizontal axis is Doppler units, and the vertical axis is range units. In this example, 20 range units and 64 Doppler units are cropped.

[0149] Figure 21 shows the change in model loss and classification accuracy during training over the course of epochs. It can be seen that the model converges in approximately 7 epochs, achieving an accuracy of 98%.

[0150] Figure 22 shows the test confidence for adult and child samples. Here, the horizontal axis represents the sample number, and the vertical axis represents the confidence (i.e., score). In the experiment, the first 150 samples are samples of adults in the car, and the last 150 samples are samples of children in the car. For each sample, the possible deep learning model calculates the confidence corresponding to adults and children, respectively. That is, two dots, blue and orange, are generated for each sample. If the confidence of the adult is greater than the confidence of the child, it is determined that the person in the car is an adult; if the confidence of the child is greater than the confidence of the adult, it is determined that the person in the car is a child. As can be seen from the figure, the trained CNN model can correctly determine whether the person in the car is an adult or a child.

[0151] In some embodiments, as shown in FIG. 24, a target detection method includes:

[0152] In step 10, after range-dimensional FFT processing based on the echo signal, 1D-FFT data is obtained.

[0153] In step 20, digital beamforming is performed on the 1D-FFT data to obtain a predetermined number of frame data.

[0154] In step 30, a target classification process is performed on a predetermined number of the frame data based on a machine learning model so as to realize target detection.

[0155] The target detection in the region of interest includes at least one of the following operations: determining, locating, and recognizing.

[0156] By performing beamforming on a predetermined region after range-dimensional FFT, and then using the multi-frame beamformed data to perform target classification processing directly based on a machine learning model, it is possible to effectively detect whether a person is in the vehicle and accurately distinguish between adults and children. Compared to point cloud information based on traditional signal processing, the method of this embodiment has lower requirements on the processing unit (such as a BB unit), uses higher data feature dimensions, and has a richer amount of information, thereby more accurately achieving target detection and recognition within the cabin and enabling applications such as CPD and SBR.

[0157] In an exemplary embodiment, the machine learning model may be a complex-valued machine learning model, for example a complex-valued neural network model.

[0158] The embodiment shown in FIG. 24 differs from the previously described embodiments mainly in that the model used in the embodiment shown in FIG. 24 is a machine learning model, while the model used in the previously described embodiments is a deep learning model. To adapt to the input of the complex machine learning model, some embodiments can be realized by transforming the feature vector. The transformation method is to convert the power and phase into real and imaginary parts.

[0159]

number

number

number

[0160] In the formula, zreal and zimg are the normalized real and imaginary parts, respectively. φ(r,v,b) is the phase of range unit r chirp v in beam b. z(r,v,b) is the normalized complex signal. Also, in some embodiments, double-precision floating-point or single-precision floating-point numbers are employed for z(r,v,b), or quantification formats such as int32, int16, int8, uint32, uint16, uint8, or other quantification processes can be used.

[0161] Furthermore, different loss functions can be used based on different models. For example, in a complex neural network model, the last layer of the complex-valued CNN, the absolute value (abs) layer, can be removed. In this case, the output of the complex-valued CNN is a complex output. In this case, the following loss function can be adopted:

[0162]

number

number

number

[0163] Here, y is the forward computation output of the Complex-valued CNN for a sample and is a vector with a length equal to the number of classifications. yi is the i-th scalar of y. t is the one-hot encoding label for the sample and is also a vector with a length equal to the number of classifications. ti is the i-th scalar of t. Other hyperparameters such as learning rate and weight decay can have other values.

[0164] Other features or methods of realizing the features of the embodiment shown in FIG. 24 are basically the same as those of the above-described embodiments, and therefore a description thereof will be omitted here.

[0165] Of course, in the above embodiments, the use of the model replaces the corresponding digital signal processing process, e.g., replacing the processes of angle estimation, CFAR, etc. In some embodiments, the model may be used to further process the target obtained by digital signal processing.

[0166] Based on this, in some embodiments, as shown in FIG. 25, the target detection method includes the following steps:

[0167] In step 30, after performing range-dimensional FFT processing based on the echo signal, 1D-FFT data is obtained, and the 1D-FFT data is subjected to multi-frame joint processing so as to obtain an RD spectrum as the feature vector.

[0168] In step 40, target point cloud data is obtained based on the RD spectrum, and a machine learning-based target classification process is performed on the target point cloud data, so as to realize target detection in the region of interest.

[0169] The target detection in the region of interest includes at least one of the following operations: determining, locating, and recognizing.

[0170] The target detection method and related device according to the present embodiment utilize multi-frame collaborative processing technology, i.e., an inter-frame accumulation method, to realize a means for target detection in enclosed spatial areas such as cabins, indoors, and factory buildings. This effectively improves the detection rate, reduces the number of false alarms, and significantly improves the accuracy of angle estimation, thereby enabling accurate detection of special targets such as infants and young children in cabins or weak targets, enabling applications such as child presence detection (CPD) and safety belt reminders (SBR). Furthermore, machine learning-based target classification processing of point cloud data enables more precise area determination and more accurate recognition of targets in the area, thereby improving the accuracy of determining the presence or absence and location of personnel in enclosed environments such as cabins.

[0171] In an exemplary embodiment, step 30 includes performing range-dimensional FFT processing on the echo signal within a chirp to obtain 1D-FFT data, accumulating the 1D-FFT data by frame until a predetermined amount of data is accumulated, and then performing inter-frame FFT processing to obtain an RD spectrum. For example, a predetermined number of frames of 1D-FFT data can be read each time using a sliding window and subjected to inter-frame FFT processing. That is, accurate detection of targets within the cabin can be achieved based on multi-frame joint processing technology. The multi-frame joint processing means performs a sliding window FFT on the multi-frame data to obtain a range-Doppler spectrum, then processes it using FIR (Finite Impulse Response) or other complex time-frequency transform, and then performs processing operations such as CFAR and DOA on the region of interest. By combining this with the set region judgment logic and region parameters, detection and location operations for living targets such as adults, children, and pets, or other non-living targets, can be achieved. The CFAR may be Doppler-dimensional NR-CFAR, RD-CFAR, or DAE (Doppler-Azimuth-Elevation)-CFAR, etc. After CFAR and DoA, post-processing such as clustering, false alarm suppression, and association of multiple-processed point clouds may be employed. The region parameters may be determined by at least one or a combination of at least two of a clustering operation on point clouds detected after a certain sliding window multi-frame processing, an outlier detection and removal operation on point clouds detected after a certain sliding window multi-frame processing, etc.

[0172] In an exemplary embodiment, the target point cloud data acquisition step in step 40 includes performing non-coherent accumulation on the RD spectrum of each channel to obtain candidate target detection points, performing constant false alarm processing based on noise estimation based on the result of the non-coherent accumulation, estimating the azimuth mid-angle and elevation angle for the candidate target detection points, and acquiring target point cloud data based on the azimuth mid-angle and elevation angle. By performing non-coherent accumulation on the RD spectrum of each channel, the signal-to-noise ratio between personnel targets and strong static clutter can be effectively improved.

[0173] In some embodiments, the constant false alarm processing includes performing non-coherent accumulation on the RD spectrum of each channel to obtain a noise floor estimate for each range unit, estimating the noise floor of each range unit after the non-coherent accumulation, and performing non-coherent constant false alarm processing according to the noise floor estimate, wherein when estimating the noise floor of each range unit after the non-coherent accumulation, the noise floor of each range unit is corrected by the global noise floor.

[0174] In some embodiments, the correction method includes correcting the noise floor of each range unit by n i '=min(n i ,α*ng), where n i is the original noise floor estimate of the i-th range unit, n g is the average value of the noise floor estimates of multiple range units, and n i ' is the corrected noise floor estimate of the i-th range unit, where α>1. Correcting the noise floor can prevent targets from going undetected due to a high noise floor.

[0175] In an exemplary embodiment, the step 40 of performing machine learning-based target classification processing on the target point cloud data to realize target detection (e.g., determination) in the region of interest includes inputting the target point cloud data including coordinate data to a pre-trained first machine learning classifier to obtain a first detection result indicating whether a candidate target detection point is a valid candidate target detection point. The coordinate data refers to coordinates in a Cartesian coordinate system including x-coordinates, y-coordinates, and z-coordinates. In some embodiments, the data input to the first machine learning classifier includes one or more of range data, signal-to-noise ratio, amplitude value, azimuth mid-angle, and elevation angle.

[0176] In exemplary embodiments, step 40 of subjecting the target point cloud data to machine learning-based target classification processing to achieve target detection (e.g., location, recognition, etc.) in the region of interest includes processing the target point cloud data to extract area information of candidate target detection points and inputting the area information to a pre-trained second machine learning classifier to obtain a second detection result indicating whether a target is present in the predetermined region. In some embodiments, the area information of the candidate target detection points includes one or more of the following information: the number of valid candidate target detection points in each predetermined region; the proportion of valid candidate target detection points in each predetermined region to the total number of valid candidate target detection points; the range, azimuth mid-angle, elevation angle, mean value and variance of a constant false alarm signal-to-noise ratio of all valid candidate target detection points in each predetermined region.

[0177] By combining classification and / or recognition with a machine learning classifier, it is possible to achieve more precise area division than with a predetermined logic, thereby accurately classifying seats corresponding to radar processing point clouds.

[0178] The first machine learning classifier or the second machine learning classifier may be any one or more of a support vector machine, a random forest, a decision tree, a Gaussian mixture model, a KNN, a hidden Markov model, and a multi-layer perceptron.

[0179] When applied to a closed or relatively closed space, the space can be divided in advance to define and detect different regions of interest (divided regions). For example, in the case of target detection inside a vehicle or cabin, the area monitored by the radar can be divided into a seat section, an aisle section, etc., and then corresponding parameter types and thresholds can be set in advance for different types of regions, and the corresponding processing steps can be combined to accurately detect targets of interest in specific regions.

[0180] In some embodiments, for a private vehicle, the interior space can generally be simply divided into a headroom, a rear-row space, and a trunk space. If the rear-row space is the primary monitoring area, the rear-row space can be subsequently divided into a seat section, an aisle section, etc. The seat section can be divided into a corresponding number of seat section units based on the number of seats. Similarly, the aisle section can be divided into a corresponding number of aisle section units corresponding to the seat space. For example, in the case of a five-seater private vehicle, the three seats in the rear-row space can be divided into three seat section units and three corresponding aisle section units. By pre-setting corresponding parameter types and thresholds for different types of areas (or sections or section units) and employing appropriate signal data processing methods and steps, accurate detection of a target of interest (specific target) in a specific section or section unit can be achieved. Here, adjacent section units may have partially overlapping areas, or may be adjacent or separated by a gap of a predetermined width.

[0181] The following is a detailed description of the technical means of the present disclosure based on the technical concept of frame-level Fourier transform, with reference to the drawings: After performing a fast time-dimensional Fourier transform on the chirp of an echo signal, at least two frames of fast time-dimensional Fourier transform data are accumulated, and the data of at least two frames are subjected to a frame-dimensional Fourier transform to obtain a range-Doppler (RD) spectrum, and then a target range and / or velocity estimation operation is performed based on the RD spectrum.

[0182] As shown in Figure 26A, after performing operations such as ADC (analog-to-digital) conversion and sampling on the echo signal, range-dimensional DC removal (Range DC removal) and range-dimensional Fourier transform (Range FFT, also known as 1D-FFT or fast time-dimensional FFT) are sequentially performed, followed by multi-frame accumulation (e.g., storing 128 frames of data), Doppler DC removal, frame Fourier transform (Frame FFT), and non-coherent integration. Next, based on the noise obtained by noise estimation (e.g., using a noise variance estimator), constant false alarm detection (peak detection with peak selection (CFAR)) is performed. Finally, processing such as angle detection and target classification is performed to obtain target range, velocity, and / or angle information. For example, based on the CFAR results, azimuth and elevation can be estimated first. Specifically, this can be achieved by technologies such as DBF (digital beam forming) and DoA (direction of arrival estimation).

[0183] After estimating the azimuth / elevation angles, the output target point cloud data can be processed in combination with a machine learning model to perform accurate target detection for the region of interest. For example, in the case of a single target point cloud data obtained after estimating the azimuth / elevation angles, non-ideal data such as noise, static clutter, or false alarms generated by target multipath can be suppressed by combining it with a false alarm suppression technique based on machine learning (ML, such as an SVM or RF algorithm). That is, the target point cloud data output by the DBF & DoA can subsequently undergo one or more of ML-based false alarm suppression, clustering, and ML-based target classification. The clustering process may be an algorithm such as agglomerative clustering or DBSCAN. In an example embodiment, the point cloud data after clustering can be input to a pre-trained machine learning model (such as an SVM or RF) to determine whether a person is present in the current process. If present, its location can be further determined, and operations such as distinguishing and determining physical targets such as adults, children, and infants can be realized.

[0184] In the embodiment shown in FIG. 26A, the effects of non-ideal factors such as noise, static clutter, and target multipath can be effectively reduced, and the difficulty of selecting region parameters and the difficulty of designing complex region judgment logic can be effectively reduced, so that point cloud information can be more effectively utilized and better target detection and discrimination performance can be obtained, especially for applications in complex scenes.

[0185] In some optional embodiments, in the case of an electromagnetic wave sensor equipped with a BB (Baseband) unit and an MCU (Microcontroller Unit) unit, operations such as storing 128 frames of frame data and target classification (e.g., classification based on a Region HIST) can be performed in the MCU module based on the target detection flowchart shown in FIG. 26A. Range DC removal and Doppler DC removal can be performed in the BB module or the MCU module. In the flow shown in FIG. 26A, DC filtering of downlink ADC data in the BB module will be described as an example.

[0186] As shown in FIG. 26A, after the BB module performs 1D-FFT processing or acquires range dimension information, the 1D-FFT data is cached in the MCU module. After a predetermined amount of data (such as 128, 56, or 32 frames) is cached, the cached data for a predetermined number of frames can be subjected to DC removal and frame-level FFT to obtain a Range-Doppler (RD) spectrum.

[0187] As shown in FIG. 26A, when a system includes multiple channels, non-coherent accumulation of the multi-channel RD spectrum can be performed, and a noise variance estimator (NVE) module can estimate the noise floor of each range bin of the accumulated data. For example, the noise floor estimation can be performed based on a predetermined formula, and subsequent Noise Reference (NR)-CFAR processing can be performed based on the noise floor estimate. The predetermined formula can be n i = min(n i,α*ng), where n i is the noise floor estimate for the i-th range bin, and ng is the noise floor estimate for the last range bin of interest. α is a coefficient greater than 1 that can be set and updated based on demand, engineering data, experience, etc.

[0188] For angle estimation, as shown in FIG. 26A, azimuth dimension DBF and DoA and elevation dimension DBF and DoA can be performed on the CFAR detection point data, respectively, so that azimuth and elevation angle estimation data for each CFAR detection point is obtained.

[0189] FIG. 26B is a schematic diagram of another flow for realizing target detection based on frame-level FFT in an embodiment of the present application, which differs from FIG. 2A in that the subsequent machine learning-based classification process is slightly different. The target point cloud data output by the DBF & DoA is subjected to one or more of ML-based target point classification, region feature extraction, and ML-based target region classification. Here, the ML-based target point classification process is used to determine targets in the region of interest, i.e., to classify the target point cloud data so as to obtain valid target detection points (i.e., the aforementioned valid target candidate detection points). The region feature extraction is useful for classification, and the classification performance can be enhanced, for example, by clustering, which helps to improve the accuracy and performance of the classification. The ML-based target region classification process is used to locate and recognize targets in the region of interest, i.e., to determine whether a target exists in a specified region.

[0190] In the above embodiment, to accurately detect targets of interest, a frame-level FFT replaces the conventional chirp-level Doppler FFT to determine target velocity, and multi-frame sliding window FFT processing further improves the update frequency of results, thereby improving the real-time performance of the system. Processing only the range and / or Doppler regions of interest can also reduce the amount of data to be processed. Furthermore, background noise estimation can be achieved by combining an estimate based on the current range bin with an estimate based on the last range bin of interest to obtain the minimum value, which is more suitable for applications in relatively enclosed environments such as inside a vehicle or cabin. When determining the region logic, targets detected for each frame are counted according to a predetermined region design, and the counted results are judged according to predetermined rules, thereby enabling accurate determination of whether a target is present in each predetermined region.

[0191] Similarly to the above-described embodiment, when target detection is performed, the implementation or replacement of the corresponding processing is not limited. For example, the FFT processing can be a windowed FFT, or can be realized by combining SVA and FFT. Alternatively, the multi-frame sliding window FFT can be replaced with other time-frequency transform processing such as short-time Fourier transform or fractional Fourier transform to effectively detect targets of interest, and the description thereof is omitted here.

[0192] In some embodiments, two-dimensional digital beamforming (2D-DBF) processing and ML-based target classification processing (ML-based classification) operations can be performed directly after inter-frame FFT (frame FFT) processing. That is, after inter-frame FFT, azimuth and elevation angles are calculated for each region of interest (or unit section) without CFAR processing, and 2D-DBF operations are performed so that the number of RD spectra corresponding to the region of interest is obtained. The RD spectra can then be input into a pre-built machine learning model (ML) for target determination. In some optional embodiments, when processing is performed using FIR, STFT (Short-time Fourier Transform), etc. instead of inter-frame FFT, the RD spectrum is a one-dimensional spectrum output corresponding to FIR and a three-dimensional spectrum that is a time-frequency transform corresponding to STFT.

[0193] The embodiments of the present application may refer to each other and be compatible when there is no contradiction, and the order and configuration of each functional module may be adjusted according to needs. For a system including a BB module and an MCU module, when the system performs target detection by emitting electromagnetic waves and receiving corresponding echo signals, each step of each target detection method according to the embodiments of the present application may be configured to operate in the BB module and / or the MCU module according to actual needs and consideration of aspects such as data processing capacity and timeliness of system operation, and the relevant examples in the figures may serve as reference for some of these options.

[0194] The method according to the present invention employs multi-frame joint processing technology to perform coherent accumulation of human breathing frequencies in the Doppler domain and multi-antenna channels in a predetermined direction in the spatial domain, thereby effectively improving the signal-to-noise ratio between personnel targets and strong static clutter. Furthermore, a machine learning model is combined with clustering and classification of radar point clouds to achieve more precise area determination than predefined logic, thereby accurately classifying seats corresponding to radar processed point clouds. By improving the signal-to-noise ratio, self-labeling of unsupervised clustering, and sophisticated machine learning classification, it is possible to more accurately determine the presence and location of personnel in enclosed environments such as cabins. Furthermore, the model construction and processing flow requires fewer rules and less difficulty in parameter adjustment compared to rule-based area determination, thereby increasing the convenience and applicability of the method.

[0195] Based on the above method, and after testing using actual collected data, it has been found that the application scenario of the top-mounted radar can effectively achieve a high detection rate and very low false alarm and false miss rates. For example, when detecting targets inside the cabin, it can achieve a detection rate of over 99%, a false miss rate of less than 0.5%, and a false alarm rate of less than 1%.

[0196] In an embodiment of the present application, a target detection method is provided that is applicable to target detection in a specific target area (e.g., an enclosed or semi-enclosed area). The method includes inputting target point cloud data into a first pre-trained machine learning classifier to obtain valid candidate target detection points, extracting area information of the valid candidate target detection points, and inputting the extracted area information into a second pre-trained machine learning classifier to obtain a target detection result. For specific step contents, please refer to the description in the preceding or following embodiments, and further description will be omitted here.

[0197] 26B and the steps corresponding to the flow of the embodiment shown in FIG. 1 are realized in almost the same way, and the same parts will not be described here. The following mainly describes the differences.

[0198] In the flow shown in FIG. 26B, when estimating the azimuth intermediate angle of a candidate target detection point, the following operations are performed for each candidate target detection point Ti: 2D-FFT data corresponding to the detection point is extracted, an azimuth-dimension antenna is selected, azimuth-dimension DBF (digital beamforming) is performed based on the azimuth-dimension steering vector, an azimuth-dimension DBF spectrum indicating how signal power or intensity in different directions changes with frequency is obtained, and the azimuth intermediate angle θi,j is estimated based on the azimuth-dimension DBF spectrum. The relationship between the azimuth intermediate angle and the azimuth angle is azimuth intermediate angle = sin (azimuth angle) cos (elevation angle). When estimating the elevation angle of a candidate target detection point, an elevation direction steering vector sv(i,j) is generated for each candidate target point Ti,j based on its azimuth intermediate angle θi,j and arrangement, an elevation-dimension DBF is performed for each candidate target point Ti,j to obtain an elevation-dimension DBF spectrum, and an elevation angle φi,j,k is estimated based on the elevation-dimension DBF spectrum. Here, i represents the number of the CFAR point (e.g., range unit). j represents the number of candidate target points in different orientations for a CFAR point, and k represents the number of candidate target points in different elevations for a CFAR point.

[0199] Furthermore, in some exemplary embodiments, when estimating the azimuth mid-angle and elevation angle, the azimuth mid-angle and elevation angle can be estimated by directly performing azimuth-elevation two-dimensional DBF, or by other super-resolution algorithms such as the minimum variance unbiased estimation (Capon) algorithm or the multiple signal classification (MUSIC) algorithm. Therefore, when extracting target information, for all candidate target points Ti,j,k that meet the candidate target point detection condition (e.g., exceeding the CFAR threshold), information can be extracted, including but not limited to the target point's range bin index, Doppler bin index, range ri,j,k, Doppler frequency f,j,k, azimuth mid-angle θ, elevation angle φ, CFAR SNR sn, and the amplitude of each channel a (c is the channel index). Then, a coordinate system transformation is performed. The coordinate system is converted for the range ri,j,k, azimuth intermediate angle θi,j,k, and elevation angle φi,j,k of each candidate target point Ti,j,k to obtain the position x, y, and z in the Cartesian coordinate system. The data format to be converted is as follows:

[0200]

number

number

number

[0201] Therefore, a pre-trained first machine learning classifier is used to classify the region to which the candidate target belongs. That is, the first machine learning classifier uses multiple attributes of the candidate target points Ti,j,k as input to classify the region to which each candidate target point belongs, thereby determining whether the candidate target point Ti,j,k is a valid target point or an invalid interference point for a specific region. Possible input examples for the first machine learning classifier are ri,j,k, xi,j,k, zi,j,k, yi,j,k, snri,j,k, and ampi,j,k,c. The machine learning classifier may be, for example, a support vector machine (SVM) or a random forest (RF).

[0202] In addition, the first machine learning classifier may have multiple forms, or even a fusion of multiple forms, such as a fusion of a support vector machine, a decision tree, and a Gaussian mixture model, and a weight is set for the output result of each classifier so that a weighted average of all classifiers is obtained, and a judgment is made again based on the weighted average result so that a final judgment result is obtained.

[0203] The inputs to the first machine learning classifier are merely examples. In other embodiments, the attributes of the selected candidate target points Ti,j,k may be all or part of "ri,j,k, xi,j,k, zi,j,k, yi,j,k, fi,j,k, snri,j,k, ampi,j,k,c." For example, they may include only coordinate data (the x, y, and z values) or may include, in addition to the coordinate data, one or more of range data (the r), Doppler frequency (f), signal-to-noise ratio (SNR), and amplitude value (AMP). Alternatively, a person skilled in the art may add input values ​​to the first machine learning classifier based on the concepts of the present disclosure. For example, the coordinate data may be replaced with azimuth mid-angle and elevation angle, or the azimuth mid-angle and elevation angle data may be input to the first machine learning classifier together.

[0204] Furthermore, each region is classified by a pre-trained second machine learning classifier, i.e., region features are extracted. The number of valid target points in each region and the ratio wl of valid target points to the total number of valid target points are counted, and the mean and variance of the range, azimuth mid-angle, elevation angle, and SNR of all valid target points in each region, i.e., the mean value of the range,

number

number

number

number

number

number

number

number

[0205] For example, when aggregating the number of valid target points in each region and the proportion wl of valid target points to the total number of valid target points, valid points determined to belong to multiple regions when calculating the total number of valid points may be counted redundantly (in cases suitable for multi-label classification, such as a scene where a person is present in both region A and region B), or may not be counted redundantly (in cases suitable for single-label classification), but the present disclosure is not limited to this.

[0206] The second machine learning classifier may have multiple forms, or even a fusion of multiple forms, such as a fusion of a support vector machine, a decision tree, and a Gaussian mixture model, and a weight is set for the output result of each classifier so as to obtain a weighted average of all the classifiers, and a judgment is made again based on the weighted average result so as to obtain a final judgment result.

[0207] The input of the second machine learning classifier is merely an example. In other embodiments, the attributes of the selected candidate target points Ti,j,k are

number

number

number

number

[0208] The first machine learning classifier for determining the valid points may be one of SVM, RF, decision tree, Gaussian mixture model, K-nearest neighbor (KNN), hidden Markov, multi-layer perceptron, etc. The training step of the classifier model includes:

[0209] 1. Generate valid point labels. A possible process is as follows:

[0210] As shown in the example in Figure 6, the region is first divided. The horizontal axis in the figure is the X direction (the width direction of the vehicle) and the Z direction (from the rear end of the vehicle to the front end of the vehicle). The region is divided into six regions: the three rear seats and the three aisles in front of the three seats. These regions may or may not overlap. Some regions in the vehicle may belong to multiple specific regions simultaneously, or may not belong to any region. In other embodiments, the region division can be combined with the target occupant to further subdivide the region. For example, rear seats A, B, and C can be divided into adult seat A, adult seat B, adult seat C, child seat A, child seat B, and child seat C. In this way, if an application requires determining whether there is a person in rear seats A, B, or C, and whether they are an adult or a child, the results can be obtained directly. All that is required is to add and train classifications corresponding to the valid point and region determinations.

[0211] For data of various scenes, including a scene with a person in seat A, a scene with a person in seat B, a scene with a person in seat C, or an interference scene with no person, the above steps 1 to 10 can be used to obtain the attributes of all candidate target points Ti,j,k in each scene, and then select all or part of, for example, ri,j,k, xi,j,k, zi,j,k, yi,j,k, fi,j,k, snri,j,k, ampi,j,k,c, across various area scenes such as seats A, B, and C and aisle areas A, B, and C in Figure 5, and perform the following operations:

[0212] 1.1 Perform clustering analysis on all candidate target point data of the scene to remove outliers. In some embodiments, a density clustering algorithm (DBSCAN) can be used to perform clustering analysis on these candidate target point data. In other embodiments, other clustering algorithms can be used to generate valid point labels, or the data can be manually labeled.

[0213] 1.2 After clustering analysis, the category with the most types is selected as the true target of the region, and their labels are set to 1. Candidate points in other categories are all false targets, and their labels are set to 0. In another embodiment, all non-noise candidate target points in DBSCAN clustering analysis can be selected as true target points, and their labels are set to 1, and noisy candidate target points in DBSCAN clustering analysis can be selected as invalid target points, and their labels are set to 0. Alternatively, some categories with a large number of samples can be artificially selected as the true target of the region, and their labels are set to 1, and candidate points in other categories are all false targets, and their labels are set to 0.

[0214] In the empty and interference scenes, all candidate target points are set as invalid target points and their labels are set to 0.

[0215] 2. Based on the generated labels, the model is trained. The main step is to traverse various scene areas such as seats A, B, and C and aisles A, B, and C, and perform the following operations:

[0216] 2.1. The attributes and labels of all candidate target points in the region are read, denoted as S{Xi} and S{Li}, respectively. The attributes and labels of all candidate target points in the empty vehicle and interference scenes are read, denoted as S{X0} and S{L0}, respectively. Here, each element of S{Xi} is an N×1 vector. In a possible embodiment, valid point discrimination is performed by any six scalars (e.g., ri,j,k, xi,j,k, zi,j,k, yi,j,k, fi,j,k, snri,j,k, ampi,j,k,c) of the attributes of the candidate target points ri,j,k, xi,j,k, zi,j,k, yi,j,k, fi,j,k, and snri,j,k. When a total of 16 MIMO virtual channels are considered, each element of S{Xl} is a 22×1 feature vector.

[0217] 2.2, Combine the attributes and labels of these target points to obtain the training input data S{X} = S{Xi} ∪ S{X0} and training labels S{L} = S{Li} ∪ S{L0}.

[0218] 2.3. Train a machine learning classifier on the input data S{X} and labels S{L} to obtain a classifier model.

[0219] For example, in the case of an application scenario in which six regions, i.e., seats A, B, and C in the back row and aisles A, B, and C, are distinguished, six machine learning classifier models Mv,l can be obtained after the above steps, where l=1, 2, ..., 6 correspond to the six regions, i.e., seats A, B, and C in the back row and aisles A, B, and C, respectively. Each machine learning classifier model Mv,l determines whether a particular candidate target point is a valid target point in each region based on the attributes of the candidate target point, such as a 22 × 1 feature vector consisting of ri,j,k, xi,j,k, zi,j,k, yi,j,k, fi,j,k, snri,j,k, and ampi,j,k,c.

[0220] For example, the first machine learning classifier may perform multiple classifications using one classifier, or may perform multiple hierarchical classifications using fewer classifiers than the number of regions. For example, in a case where six seating regions are to be classified, classifier 1 may be designed to achieve classification of seats and aisles, and then classifiers 2 and 3 may be designed to achieve classification of aisles ABC and classification of seats ABC, respectively.

[0221] For example, oversampling and / or undersampling can be performed when training the first machine learning classifier. Oversampling refers to balancing the quantity difference between different categories in a dataset by increasing the number of copies of minority samples or synthesizing new samples. Adding data to minority samples makes the sample distribution between each category more balanced, thereby improving the model's learning ability for minority samples. Undersampling refers to balancing an unbalanced dataset by reducing the number of majority samples so that the sample amounts between each category are closer. Removing majority samples makes the number of samples in each category equal, thereby reducing the model's priority over the majority.

[0222] When valid point determination is performed by the first machine learning classifier, a particular candidate target point may be determined to belong to multiple regions. For example, if the outputs of classifier models Mv,1 and Mv,2 are both 1 and the outputs of other classifier models are all 0, the candidate target point is determined to be a valid target point for seat A and seat B.

[0223] The second machine learning classifier for determining whether there is a person in each area may be any one of classifiers such as SVM, RF, decision tree, Gaussian mixture model, KNN, hidden Markov, multi-layer perceptron, etc. The training step of the classifier model includes:

[0224] Traverse various areas such as A, B, C seats and A, B, C aisle areas to perform the following operations:

[0225] 1, wl,

number

number

number

number

[0226] 2. Read the region features processed by the frame-by-frame sliding window of other scenes generated by steps 1 to 11 above as input negative sample data, and set all sample labels to 0.

[0227] 3. Combine positive and negative sample data and labels. If the number of positive samples is less than the number of negative samples during the combination, the positive samples can be resampled. The repetition factor is γ=Nn / Np, where Nn and Np are the number of negative and positive samples, respectively.

[0228] 4. Train a second machine learning classifier to obtain a classifier model Mc,l.

[0229] In an application scenario in which six regions, namely, seats A, B, and C in the back row and aisles A, B, and C, are distinguished, six machine learning classifier models Mc,l can be obtained after performing the above steps. Here, l = 1, 2, ..., 6 correspond to the six regions, namely, seats A, B, and C in the back row and aisles A, B, and C, respectively. Each machine learning classifier model Mc,l can determine whether a person is present in region l. When region features of a certain sliding window process are input, if the output of the second machine learning classifier model Mc,l is 0, it is determined that no person is present in region l. If the output of the machine learning classifier model Mc,l is 1, it is determined that a person is present in region l. Note that, based on this processing flow, it is possible to determine whether a scene in which multiple people are present is identified.

[0230] In an exemplary embodiment, empty vehicle scenes may be output as a separate category.

[0231] When valid point determination is performed by the second machine learning classifier, a particular candidate target point may be determined to belong to multiple regions. For example, if the outputs of classifier models Mc,1 and Mc,2 are both 1 and the outputs of other classifier models are all 0, it is determined that a person is present in both seat A and seat B in the scene. If the outputs of all classifier models Mc,l,l=1,2,...,6 are all 0, it is determined that no person is present in the scene.

[0232] Illustratively, the second machine learning classifier can perform multiple classifications using one classifier, or can perform multiple hierarchical classifications using fewer classifiers than the number of regions.

[0233] For example, oversampling and / or undersampling can be performed when training the second machine learning classifier. Oversampling refers to balancing the quantity difference between different categories in a dataset by increasing the number of copies of minority samples or synthesizing new samples. Adding data to minority samples makes the sample distribution between each category more balanced, thereby improving the model's learning ability for minority samples. Undersampling refers to balancing an unbalanced dataset by reducing the number of majority samples so that the sample amounts between each category are closer. Removing majority samples makes the number of samples in each category equal, thereby reducing the model's priority over the majority.

[0234] The method according to the embodiment of the present application differs from the slow time processing between chirps in radar signal processing in the related art. In the embodiment of the present application, multi-frame data is accumulated and inter-frame FFT processing is performed in the form of a sliding window, thereby improving the signal-to-noise ratio between personnel targets and strong static clutter. In the embodiment of the present application, region segmentation is performed using a machine learning model, which makes use of richer information than spatial segmentation and region determination strategies based on predetermined rules. It can be fully automated without human intervention, significantly reducing the pressure of parameter adjustment.

[0235] The embodiments of the present application are not only applicable to the human target detection and location scenario in a vehicle, but also to other similar application scenarios such as indoor personnel detection, personnel detection in a factory building, and the like.

[0236] The above operations may be performed in the CPU or MCU, or in the baseband unit BB, or partly by the CPU or MCU and partly by the BB. For example, 1D-FFT, 2D-FFT, CFAR, DBF and DOA are performed in the BB, and target detection and recognition are performed in the BB.

[0237] 26A and 26B show the results of DBSCAN clustering for all points in multiple experiments in the Seat C scene. The black points in the figures are candidate target points determined as noise by DBSCAN, and the other colored points are different classification clusters. According to the above example, the classification cluster with the largest number of samples (shown in the black frame in the figures) is selected and marked as a valid point in the Seat C area, and the other points are marked as invalid points in Seat C.

[0238] Figures 26C and 26D show the labeling results after DBSCAN clustering considering all points from multiple experiments in the scenes of seats A, B, and C. The classification clusters with the most valid points for seats A, B, and C are shown in the black outlined positions in the figures.

[0239] Compared with the prior art, the method of the present application employs a novel inter-frame accumulation method, performing coherent accumulation at the frequency of human breathing, thereby improving the signal-to-noise ratio between personnel targets and strong static clutter. Furthermore, a clustering method is employed to remove abnormal outliers, enabling self-labeling of radar point cloud samples. By combining machine learning classifiers such as SVM, random forest, and Gaussian mixture models to achieve more precise segmentation than predefined logic, seats corresponding to radar processed point clouds can be accurately classified. By improving the signal-to-noise ratio, self-labeling of unsupervised clustering, and precise machine learning classification, more accurate determination of the presence and location of personnel in enclosed environments such as cabins can be achieved. Furthermore, the model construction and processing flow requires fewer rules and less difficulty in parameter adjustment compared to rule-based area determination, improving the convenience and applicability of the method.

[0240] The division of the steps in the above method is for the purpose of clarity of explanation only, and they may be combined into one step or divided into multiple steps in implementation, as long as the same logical relationship is included, and all are within the scope of protection of this patent. Adding minor modifications to the algorithm or flow or introducing minor designs without changing the core design of the algorithm or process are also within the scope of protection of this patent.

[0241] In an embodiment of the present application, a radar system includes a radio frequency module, an analog signal processing module, and a digital signal processing module connected in series. The radio frequency module is used to generate a radio frequency transmission signal and receive a radio frequency reception signal. The analog signal processing module is used to frequency-down-process the radio frequency reception signal to obtain an intermediate frequency signal. The digital signal processing module is used to analog-to-digital convert the intermediate frequency signal. An integrated circuit is further provided for processing the digital signal according to the target detection method of the embodiment of the present application to achieve target detection. For example, the integrated circuit may be a millimeter-wave radar chip or die. Here, the digital processing module may include a baseband module, a main control module, etc. Each module may be configured to perform a corresponding step of the target detection method of the embodiment.

[0242] In some optional embodiments, the integrated circuit may be an AiP (Antenna-In-Package) chip structure, an AoP (Antenna-On-Package) chip structure, or an AoC (Antenna-On-Chip) chip structure.

[0243] According to some other embodiments of the present application, an electromagnetic wave sensor is further proposed. The electromagnetic wave sensor may include an antenna and the above-described integrated circuit. Here, the integrated circuit is electrically connected to the antenna and is used to transmit and receive electromagnetic wave signals. For example, the electromagnetic wave sensor may include a support, the integrated circuit described in any of the above embodiments, and an antenna. The integrated circuit may be provided on the support. The antenna may be provided on the support, or may be integrated with the integrated circuit as an integrated device and provided on the support (i.e., in this case, the antenna may be an antenna provided in an AiP, AoP, or AoC structure). Here, the integrated circuit is connected to the antenna (i.e., in this case, an antenna is not integrated in the sensing chip or integrated circuit, such as in a conventional SoC), and is used to transmit and receive electromagnetic wave signals. Here, the support may be a printed circuit board (PCB), and the corresponding transmission line may be PCB wiring.

[0244] In this application, electromagnetic waves may include radio waves and light waves. Radio waves include short waves, medium waves, long waves, and microwaves. Microwaves include centimeter waves (i.e., electromagnetic waves of 3 GHz to 30 GHz, such as electromagnetic waves of 3.1 GHz to 10.6 GHz and electromagnetic waves in the 24 GHz frequency band) and millimeter waves (i.e., electromagnetic waves of 30 GHz to 300 GHz, such as electromagnetic waves in the 60 GHz frequency band and electromagnetic waves in the 77 GHz frequency band (77 GHz to 81 GHz, etc.)). Light waves may include ultraviolet light, visible light, infrared light, lasers, etc. Here, the electromagnetic wave frequency band of lasers is (3.846 to 7.895)*10^5 GHz, that is, lasers are included in some frequency bands of ultraviolet light and visible light.

[0245] In an embodiment of the present application, there is provided a device including a device body and the above-mentioned electromagnetic wave sensor provided in the device body, wherein the electromagnetic wave sensor is used for target detection and / or communication to provide reference information for the operation of the device body.

[0246] In some embodiments of the present application, an electronic device is provided that can be represented in the form of a general-purpose computing device. The electronic device assembly can include, but is not limited to, at least one processing unit, at least one storage unit, a bus connecting different system assemblies (including the storage unit and the processing unit), a display unit, etc., wherein the storage unit stores program code. The program code can be executed by the processing unit to cause the processing unit to perform the methods of various exemplary embodiments of the present application described herein. The storage unit can include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) and / or a cache memory storage unit, and can further include a read-only storage unit (ROM).

[0247] The storage unit may include a program / utility having a set (at least one) of programming modules, including, but not limited to, an operating system, one or more application programs, other program modules, and program data, any one or some combination of which may include implementing a network environment.

[0248] The bus may represent one or more of several bus structures, including a storage unit bus or storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any bus structure in multiple bus structures.

[0249] An electronic device may communicate with one or more external devices (e.g., keyboards, pointing devices, Bluetooth devices, etc.), one or more devices that allow a user to interact with the electronic device, and / or any device (e.g., routers, modems, etc.) that allows the electronic device to communicate with one or more other computing devices. This communication may occur via an input / output (I / O) interface. The electronic device may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN) such as the Internet, and / or a public network) via a network adapter. The network adapter may communicate with other modules of the electronic device via a bus. Although not shown, other hardware and / or software modules may be used in conjunction with the electronic device, including, but not limited to, microcode, device drives, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0250] For example, an electronic device according to an embodiment of the present application may further include a device main body and an electromagnetic wave sensor according to any of the above embodiments provided on the device main body, wherein the electromagnetic wave sensor is used to realize functions such as target detection and / or wireless communication.

[0251] Specifically, based on the above embodiment, in any embodiment of the present application, the electromagnetic wave sensor may be provided outside the device body or inside the device body. Also, in any other embodiment of the present application, the electromagnetic wave sensor may be provided partially inside the device body and partially outside the device body. The embodiment of the present application is not limited thereto, and specifically may be determined depending on the situation.

[0252] In any embodiment, the device body may be a component or product applied in the fields of smart cities, smart homes, transportation, smart homes, household appliances, security monitoring, industrial automation, in-cabin detection (such as a smart cockpit), medical equipment, and healthcare, etc. For example, the device body may be a smart transportation device (such as a car, bicycle, motorcycle, ship, subway, or train), a security device (such as a camera), a liquid level / flow rate detection device, a smart wearable device (such as a bracelet or glasses), a smart home device (such as a vacuum cleaner robot, a door lock, a television, an air conditioner, or a smart light), various communication devices (such as a mobile phone or a tablet computer), a barrier gate, a smart traffic light, a smart sign, a traffic camera, various industrial robot arms (or robots), etc. It may also be various instruments used to detect biometric parameters, such as biometric detection in the cabin of a car, indoor occupant monitoring, smart medical equipment, or household appliances, and various devices equipped with such instruments.

[0253] An embodiment of the present application further provides a non-transitory computer-readable storage medium having computer-readable instructions stored thereon, which, when executed by a processor, cause the processor to perform the feeder unequal length compensation method described above.

[0254] As can be easily understood by those skilled in the art from the description of the above embodiments, the exemplary embodiments described herein can be realized by software, or by combining software with necessary hardware. The technical means according to the embodiments of the present application can be embodied in the form of a software product. The software product can be stored in a non-volatile storage medium (such as a CD-ROM, a USB memory, or a mobile hard disk) or a network, and includes a plurality of instructions to enable a computing device (such as a personal computer, a server, or a network device) to execute the above method according to the embodiments of the present application.

[0255] A software product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (not an exhaustive list) of readable storage media include an electrical connection having one or more conductors, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash), fiber optics, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0256] A computer-readable storage medium may include a data signal propagating in baseband or as part of a carrier wave bearing readable program code. This propagated data signal may take multiple forms, including, but not limited to, an electromagnetic signal, an optical signal, or any suitable combination of the above. A computer-readable storage medium may also be any readable medium other than a computer-readable storage medium that transmits, propagates, or transmits a program intended for use by or in connection with an instruction execution system, apparatus, or device. The program code contained in the computer-readable storage medium may be transmitted by any suitable medium, including, but not limited to, wireless, wired, cable, RF, etc., or any suitable combination of the above.

[0257] Program code for carrying out the operations of the present application may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, etc., and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may execute entirely on the user computing device, partially on the user device, as a standalone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet Service Provider).

[0258] The computer-readable medium carries one or more programs, which, when executed by the device, cause the computer-readable medium to implement the functions described above.

[0259] Those skilled in the art will appreciate that the above modules may be distributed within a device according to the description of the embodiment, or may be modified accordingly in one or more devices uniquely different from the embodiment. The modules of the above embodiment may be combined into one module or further divided into multiple sub-modules.

[0260] According to an embodiment of the present application, a computer program including instructions can be implemented by a processor to perform the above method. In an optional embodiment, the integrated circuit can be a millimeter-wave radar chip. The type of digital functional module within the integrated circuit can be determined according to actual needs. For example, in the case of a millimeter-wave radar chip, the data processing module can be used for range-dimensional Doppler transform, velocity-dimensional Doppler transform, false alarm detection, wave direction detection, point cloud processing, etc., to obtain information such as the target's range, angle, velocity, height, micro-Doppler motion characteristics, shape, size, surface roughness, and dielectric properties.

[0261] In addition, wireless devices can realize functions such as target detection and / or communication by transmitting and receiving wireless signals, and can therefore provide detected target information and / or communication information to the device main body, thereby assisting or controlling the operation of the device main body.

[0262] For example, when the above-mentioned device main body is applied to an advanced driver assistance system (ADAS), a wireless device (such as a millimeter-wave radar) as an in-vehicle sensor can support the ADAS system to realize application scenarios such as adaptive cruise control, automatic brake assist (i.e., AEB), blind spot detection warning (i.e., BSD), lane change assist warning (i.e., LCA), reversing assist warning (i.e., RCTA), parking assistance, rear vehicle warning, collision avoidance, pedestrian detection, and in-cabin life detection (i.e., CPD).

[0263] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described, but as long as there is no contradiction in the combinations of these technical features, they shall be deemed to be within the scope described herein.

[0264] The above embodiments merely represent preferred embodiments of the present invention and the technical principles used. Although the descriptions are specific and detailed, they should not be understood as limiting the scope of the invention patent. Various obvious modifications, adjustments, and substitutions are possible for those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments and may include many other equivalent embodiments without departing from the spirit of the present invention. The scope of protection of the invention patent is determined by the appended claims.

[0265] This application is filed based on and claims priority from a Chinese patent application bearing application number "202310913653.8" and filed on July 24, 2023, the entire contents of which are incorporated herein by reference.

Claims

1. FFT processing the echo signals to generate feature vectors of the echo signals; and processing the feature vector based on a machine learning model to achieve target detection in a region of interest. A target detection method comprising:

2. The feature vector comprises characterizing energy features of the echo signal at different range units, or The feature vector includes characterizing energy features of the echo signal in different Doppler dimensions of different range units.

2. The target detection method according to claim 1.

3. The feature vector includes characterizing energy features of the echo signal in different Doppler dimensions of different range units; FFT processing of the echo signal is FFT processing the echo signals to generate a range-Doppler spectrum; performing beamforming processing based on the range Doppler spectrum to obtain a feature vector of the echo signal; 3. The target detection method according to claim 2.

4. FFT processing of the echo signals to generate a range-Doppler spectrum is performing range-dimensional FFT processing on the echo signal; and applying a sliding window to the result obtained by the range-dimension FFT processing, and performing 2D FFT processing on the windowed data, so as to obtain the range Doppler spectrum; The length of the sliding window and the period of the corresponding periodic motion of the target have the same time length.

4. The target detection method according to claim 3.

5. The beamforming process based on the range Doppler spectrum is realized by the following equation: [Equation 1] where P(r, v, b) is the energy feature of the v-th Doppler unit of the r-th range unit based on beam b; sv(c,b) is the steering vector of beam b in virtual channel c; x(c, r, v) is data corresponding to the v-th Doppler unit of the r-th range unit in the virtual channel c in the range-Doppler spectrum; Nc is the total number of beams, 4. The target detection method according to claim 3.

6. The feature vector includes characterizing energy features of the echo signal in different Doppler dimensions of different range units; Performing FFT processing and beamforming processing on the echo signals Range FFT processing of the echo signal; performing beamforming processing on the result obtained by the range FFT processing; and sliding windowing the beamforming result and performing 2D FFT on the windowed data.

3. The target detection method according to claim 2.

7. The feature vector includes characterizing energy features of the echo signal at different range units; Performing FFT processing and beamforming processing on the echo signals Range FFT processing of the echo signal; and performing beamforming processing on the result obtained by the range FFT processing.

3. The target detection method according to claim 2.

8. Beamforming on echo signals is identifying a target direction toward the region of interest and generating a steering vector based on the target direction; beamforming echo signals based on the steering vectors.

8. The target detection method according to claim 3, wherein the target detection method is a method for detecting a target.

9. Identifying a target direction toward the region of interest includes: determining the target direction based on an azimuth angle and an elevation angle of the region of interest relative to a radar; 9. The target detection method according to claim 8.

10. Identifying a target direction toward the region of interest includes: acquiring target measurement data in the region of interest; and performing clustering based on the target measured data, thereby determining the azimuth angle and elevation angle of a clustering center point in each of the regions of interest as the target direction.

9. The target detection method according to claim 8.

11. Generating a steering vector based on the target direction is achieved by the following equation: [Equation 2] where s(c, b) is the steering vector of beam b in virtual channel c, λ is the central wavelength of the radar detection signal, xc and zc are the horizontal and vertical positions of the antenna for virtual channel c, respectively; θb is the azimuth angle corresponding to the target direction, φb is the elevation angle corresponding to the target direction; 9. The target detection method according to claim 8.

12. generating a feature vector of the echo signal, rearranging the energy features obtained by FFT processing and beamforming processing of the echo signals so as to obtain the feature vector according to a predetermined rule; 8. The target detection method according to claim 2, wherein the target detection method is a method for detecting a target.

13. The predetermined rule is: The data having the same first dimension are arranged consecutively according to the gradient direction of the second dimension, and the data are arranged consecutively according to the gradient direction of the first dimension as a whole; Or, the data having the same third dimension and the same fourth dimension are sequentially arranged according to the gradient direction of the fifth dimension, the data having the same third dimension are sequentially arranged according to the gradient direction of the fourth dimension, and the data are sequentially arranged according to the gradient direction of the third dimension as a whole; the first dimension and the second dimension are different dimensions in a beam dimension and a range dimension, respectively; the third dimension, the fourth dimension, and the fifth dimension are different dimensions in a beam dimension, a range dimension, and a Doppler dimension, respectively.

13. The target detection method according to claim 12.

14. generating a feature vector of the echo signal, log-normalizing energy features obtained by FFT processing and beamforming processing of the echo signals; 8. The target detection method according to claim 2, wherein the target detection method is a method for detecting a target.

15. Logarithmically normalizing the energy features obtained by FFT processing and beamforming processing on the echo signals obtaining a logarithm of an energy feature obtained by FFT processing and beamforming processing of the echo signal, and linearly normalizing the logarithm obtained result according to predetermined upper and lower limit constraints.

15. The target detection method according to claim 14.

16. The logarithm is obtained for the energy feature obtained by the FFT processing and the beamforming processing on the echo signal, and the logarithm is linearly normalized according to the predetermined upper and lower limit constraints, as follows: [Equation 3] where Pnorm(r, v, b) is the log-normalization result, P(r,v,b) is the energy signature of beam b in range unit r and Doppler unit v; a0, b0, Pmax, and Pmin are all predetermined parameters.

16. The target detection method according to claim 15.

17. the machine learning model is a deep learning model; before processing the feature vector based on a machine learning model; Obtaining a training feature vector having one-hot encoding labels and characterizing the energy features of the signal in different range units; processing the training feature vectors to obtain a training set using a predetermined data augmentation mode including at least one of randomly increasing white noise, inverting the data along the Doppler dimension, and translating the data along the range dimension; training a deep learning model based on the training set.

8. The target detection method according to claim 2, wherein the target detection method is a method for detecting a target.

18. The result of the target detection is a confidence matrix; The elements of the confidence matrix correspond to at least one of the following information: a confidence that a target will be detected, a confidence that a target exists in the corresponding region of interest, a confidence that an adult target exists in the corresponding region of interest, and a confidence that a child target exists in the corresponding region of interest.

8. The target detection method according to claim 2, wherein the target detection method is a method for detecting a target.

19. The feature vector characterizes energy features of the echo signals at different range units within a predetermined range range and / or a predetermined Doppler range.

8. The target detection method according to claim 2, wherein the target detection method is a method for detecting a target.

20. the machine learning model includes a neural network model; Processing the feature vector based on a machine model includes: Normalizing a predetermined number of frames of data and inputting the normalized results into a pre-trained neural network model; or normalizing a predetermined number of frame data, converting the normalized result into a complex number, and inputting the complex number, or a real part of the complex number, or an imaginary part of the complex number into the neural network model; 2. The target detection method according to claim 1.

21. Normalizing the predetermined number of frame data includes normalizing the data by the following formula: [Equation 4] where Pnorm(r,v,b) is the result obtained after the beam b and range unit r chirp v are normalized, P(r,v,b) is the power of the DBF data of beam b, range unit r chirp v, b0 and a0 are predetermined normalization parameters, Pmax and Pmin are the upper and lower bounds of the given P(r,v,b), respectively.

21. The method of claim 20.

22. b0 is the mean or median of all possible log2(P(r,v,b)) values, a0 is the difference between the 95% quantile and the 5% quantile of all possible log2(P(r,v,b)) values ​​divided by 256; 22. The method of claim 21 .

23. Converting the normalized result to a complex number includes converting it according to the following formula: [Equation 5] where zreal and zimg are the transformed real and imaginary parts, Pnorm(r, v, b) is the result of the normalization process, φ(r,v,b) is the phase of beam b, range unit r chirp v, 21. The method of claim 20.

24. the neural network model is a complex neural network model; and / or The neural network model includes a complex convolutional layer, a complex nonlinear activation layer, a pooling layer, a complex fully connected layer, and an absolute value layer.

24. The method of claim 23.

25. FFT processing the echo signal to generate a feature vector of the echo signal, performing range-dimensional FFT processing based on the echo signal, acquiring 1D-FFT data, and performing multi-frame joint processing on the 1D-FFT data to obtain an RD spectrum as the feature vector; 2. The target detection method according to claim 1.

26. processing the feature vector based on a machine learning model to achieve target detection in a region of interest; inputting target point cloud data including coordinate data into a pre-trained first machine learning classifier so as to obtain a first detection result indicating whether a candidate target detection point is a valid candidate target detection point; the first machine learning classifier is one or more of a support vector machine, a random forest, a decision tree, a Gaussian mixture model, a KNN, a hidden Markov model, and a multi-layer perceptron; 26. The method of claim 25.

27. the data input to the first machine learning classifier further includes one or more of range data, a constant false alarm signal-to-noise ratio, an amplitude value, a Doppler frequency, an azimuth mid-angle, and an elevation angle; 27. The method of claim 26.

28. processing the feature vector based on a machine learning model to achieve target detection in a region of interest; processing the target point cloud data to extract area information of the candidate target detection points, and inputting the area information into a pre-trained second machine learning classifier to obtain a second detection result indicating whether or not a target is present in a predetermined area; The second machine learning classifier is one or more of a support vector machine, a random forest, a decision tree, a Gaussian mixture model, a KNN, a hidden Markov model, and a multi-layer perceptron.

26. The method of claim 25.

29. The radio frequency module, the analog signal processing module, and the digital signal processing module are connected in series; the radio frequency module is used to generate a radio frequency transmission signal and receive a radio frequency reception signal; the analog signal processing module is used for frequency down-processing a radio frequency received signal to obtain an intermediate frequency signal; the digital signal processing module is used to perform analog-to-digital conversion of the intermediate frequency signal and process the digital data obtained by the analog-to-digital conversion according to the target detection method of any one of claims 1 to 28, so as to realize target detection; 1. An integrated circuit comprising:

30. A support; 28. The integrated circuit of claim 27 mounted on the support; an antenna provided on the support or integrated with the integrated circuit as an integral device and provided on the support; the integrated circuit is connected to the antenna for transmitting radio frequency transmission signals and / or receiving radio frequency reception signals; An electromagnetic wave sensor characterized by:

31. A device body, the electromagnetic wave sensor according to claim 28 provided in the device body; The electromagnetic wave sensor is used for target detection and / or communication to provide reference information for the operation of the device body. A terminal device characterized by:

Citation Information

Patent Citations

  • Gesture recognition using sensors

    JP2018516365A

  • Blink detection system, blink detection method

    JP2019030582A

  • Radar device and radar signal processing method thereof

    JP2019086464A

  • Radar system, and radar signal processing method

    JP2023001662A

  • Smart home device using a single radar transmission mode for activity recognition of active users and vital sign monitoring of inactive users

    WO2022060369A1