Radar-based in-vehicle living body detection method, product, device and storage medium

By constructing a three-dimensional tensor using ultra-wideband radar and combining detection and classification branches to create a liveness detection model, the problems of low sensitivity and false alarms/missed alarms in existing in-vehicle liveness detection methods are solved, achieving efficient and reliable detection of liveness inside vehicles.

CN121028025BActive Publication Date: 2026-04-28FENG LEI ARTIFICIAL INTELLIGENCE TECHNOLOGY (SHANGHAI) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
FENG LEI ARTIFICIAL INTELLIGENCE TECHNOLOGY (SHANGHAI) CO LTD
Filing Date
2025-09-11
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing methods for detecting liveness inside vehicles, such as infrared sensors and cameras, suffer from low sensitivity, inability to detect obstructed areas, and a high rate of missed and false alarms.

Method used

The system uses ultra-wideband radar to collect raw data streams inside the vehicle, constructs a three-dimensional tensor, and performs detection using a trained liveness detection model. The model includes detection and classification branches, outputs heatmaps and category probability distributions, and combines the heatmaps and category probability distributions to obtain liveness detection results.

Benefits of technology

It improves the accuracy and reliability of liveness detection, reduces false alarms, enhances the ability to detect liveness inside the vehicle, and can effectively identify occluded liveness and their vital signs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121028025B_ABST
    Figure CN121028025B_ABST
Patent Text Reader

Abstract

The application provides a radar-based in-vehicle living body detection method, program product, electronic device and storage medium, the method comprising: collecting original data stream in the vehicle by using ultra-wideband radar; the original data stream comprises amplitude information and phase information of the received signal at each time point; constructing a three-dimensional tensor based on the original data stream information; the dimensions of the three-dimensional tensor comprise the number of slow time frames, the number of distance gates and the number of channels; inputting the three-dimensional tensor into a trained living body detection model to obtain a model output; the living body detection model comprises a detection branch and a classification branch; obtaining a living body detection result according to a heat map and a category probability distribution. The ultra-wideband radar can effectively suppress most irrelevant environmental interference through distance gates, improving the perception ability of the model to the target micro-motion characteristics. Through multi-dimensional information fusion, the false alarm phenomenon is reduced, and the reliability and practicality of the detection result are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of liveness detection, and more specifically, to a radar-based in-vehicle liveness detection method, program product, electronic device, and storage medium. Background Technology

[0002] To prevent drivers from accidentally leaving pets or children in their vehicles and causing danger, in-vehicle liveness detection methods mainly rely on infrared sensors and cameras. However, these methods have limitations in terms of technology and application: infrared sensors have low detection sensitivity; and cameras cannot detect obstructed areas, resulting in a high number of missed and false alarms. Summary of the Invention

[0003] The purpose of this application is to provide a radar-based in-vehicle liveness detection method, program product, electronic device, and storage medium to improve the above-mentioned problems.

[0004] In a first aspect, embodiments of this application provide a radar-based in-vehicle liveness detection method, comprising: acquiring raw data streams inside the vehicle using an ultra-wideband radar; the raw data streams include amplitude and phase information of the received signals at each time point; constructing a three-dimensional tensor based on the raw data stream information; the dimensions of the three-dimensional tensor include the number of slow-time frames, the number of range gates, and the number of channels; inputting the three-dimensional tensor into a trained liveness detection model to obtain model output; the liveness detection model includes a detection branch and a classification branch; the detection branch is used to output a heatmap representing the spatial location of the liveness; the classification branch is used to output the category probability distribution of the liveness; and obtaining the liveness detection result based on the heatmap and the category probability distribution.

[0005] In the aforementioned implementation process, ultra-wideband radar possesses extremely high range resolution and is highly sensitive to minute movements, enabling it to capture minute displacements caused by vital signs such as breathing and heartbeat. It mitigates the potential information loss risks associated with traditional methods using infrared or temperature sensors, providing a high-fidelity data foundation for subsequent processing and improving the accuracy and reliability of subsequent liveness detection. The three-dimensional tensor displays the information. The in-vehicle environment is complex, with various interferences such as air conditioning airflow, swaying suspended objects, and pedestrian movement outside the vehicle. Ultra-wideband radar can effectively suppress most irrelevant environmental interference through range gating, enhancing the model's ability to perceive the subtle movement features of the target. Through multi-dimensional information fusion (heatmaps and category probability distributions), false alarms are reduced, improving the reliability and practicality of the detection results.

[0006] Optionally, in this embodiment of the application, constructing a three-dimensional tensor based on the original data stream information includes: dividing the original data stream according to the pulse repetition period to obtain a two-dimensional matrix; the rows of the two-dimensional matrix represent all echo signals received after the ultra-wideband radar sends a series of pulse codes; the columns of the two-dimensional matrix represent the range gates; the range gates are used to describe the physical distance from the nearest to the farthest in the echo signal corresponding to a pulse; and obtaining a three-dimensional tensor based on the two-dimensional matrix; the dimensions of the three-dimensional tensor are the number of slow-time frames, the number of range gates, and the number of channels, respectively; the number of slow-time frames is used to capture the micro-motion characteristics of the target over time.

[0007] In the aforementioned implementation process, the three-dimensional tensor greatly enriches the information capacity, and the explicit definition of the range gate directly links the data index to the physical space, enabling the algorithm to perform spatial positioning. Furthermore, by preserving complete channel information, rather than just calculating amplitude, the phase information of the signal is fully retained, improving the detection accuracy of subtle movements of vital signs. By introducing the dimension of slow-time frames, where each "frame" is a snapshot of the radar, multiple consecutive frames record the evolution of the echo signal at each range unit over slow time. This allows the subsequent liveness detection model to utilize its powerful spatiotemporal feature extraction capabilities, clearly observing the periodic phase modulation patterns caused by respiration and heartbeat from this dimension, thereby achieving reliable detection and identification of vital signs and greatly enhancing the system's ability to perceive subtle movements of life.

[0008] Optionally, in this embodiment, the liveness detection model further includes an input layer, a first convolutional layer, and a second convolutional layer; the detection branch and classification branch in the liveness detection model are two branches set in parallel after the second convolutional layer; the input layer is used to receive a three-dimensional tensor; the first convolutional layer is used to capture temporal features of the three-dimensional tensor in the dimension of slow-time frame number; and to capture spatial features of the three-dimensional tensor in the dimension of distance gate number, generating a spatiotemporal feature map; the second convolutional layer is used to extract depth features from the spatiotemporal feature map to obtain a depth feature map; the detection branch is used to receive the depth feature map and calculate the depth feature map to generate a heatmap; the classification branch is used to receive the depth feature map and generate a liveness category probability distribution based on the depth feature map.

[0009] In the above implementation process, the first convolutional layer, through its unique spatiotemporal joint convolution kernel, can effectively extract spatiotemporal local features related to living entities directly from the raw, cluttered radar signals, enhancing and highlighting valuable signal components in the input data while suppressing irrelevant background noise, thus improving the model's initial perception ability of subtle life movement signals. The second convolutional layer, through further refinement and compression of the primary features, generates a deep feature map capable of representing high-level semantic information of the target, improving the model's ability to distinguish between different living entities and complex backgrounds. The detection branch is specifically responsible for parsing the spatial location information contained in the deep feature map and achieving effective localization of living entities inside the vehicle through heatmaps. The classification branch is specifically responsible for interpreting the semantic category information contained in the deep feature map, improving the ability to make fine distinctions.

[0010] Optionally, in this embodiment of the application, the classification branch includes a time-frequency analysis subnet; the time-frequency analysis subnet is used to determine frequency components from the deep feature map, and to determine the category probability distribution of the living organism based on the frequency components and a preset physiological frequency knowledge base.

[0011] In the above implementation process, the time-frequency analysis subnet uses frequency components as the basis for distinguishing liveness categories, enhancing the scientific nature and reliability of classification decisions. It effectively solves the classification ambiguity problem when the target is occluded, has an unusual posture, or is difficult to distinguish based solely on body shape characteristics (e.g., a curled-up child and a small pet dog). By adding a judgment dimension based on physiological and physical laws, it significantly improves the accuracy of fine-grained liveness classification and the system's anti-interference capability.

[0012] Optionally, in an embodiment of this application, the classification branch includes spatial attention; spatial attention is used to generate a weight map based on the heatmap, and the weight map is used to determine the importance of different regions in the depth feature map.

[0013] In the aforementioned implementation process, the spatial attention mechanism introduces prior localization information from the detection branch, guiding the classification branch to concentrate limited computational resources and attention on the region most effective for classification decisions, thus achieving synergy between localization and recognition tasks. Simultaneously, it effectively suppresses the activation of irrelevant features generated by complex in-vehicle backgrounds (such as swaying ornaments or shaking seat covers), greatly reducing the misleading influence of background interference on classification judgments, thereby significantly improving the accuracy and robustness of category discrimination.

[0014] Optionally, in this embodiment of the application, before inputting the three-dimensional tensor into the trained liveness detection model and obtaining the model output, the method further includes: collecting sample data of liveness in different states inside the vehicle using ultra-wideband radar; generating physical adversarial samples through simulation; the simulated scenarios include at least one of air conditioning airflow, swaying of suspended objects, or moving objects outside the vehicle; labeling the sample data and adversarial samples with corresponding labels to obtain labeled data; the labeled data includes location labels and category labels; and training a preset initial model using a multi-task loss function based on the labeled data to obtain a trained liveness detection model.

[0015] In the above implementation process, simulation can generate a massive amount of training samples covering various rare interferences and extreme scenarios at low cost and high efficiency. Generating physical adversarial examples and actively retaining these challenging scenarios during model training forces the model to learn how to ignore interference and focus on genuine liveness features, thereby greatly enhancing the model's adaptability and generalization ability in real complex environments, and significantly improving the system's anti-interference capabilities and overall detection accuracy.

[0016] Optionally, in this embodiment, the heatmap includes the confidence level of the points; obtaining the liveness detection result based on the heatmap and the category probability distribution includes: removing points in the heatmap with confidence levels below a preset threshold to obtain retained points; using a connectivity analysis algorithm, determining connected regions in the heatmap based on the retained points, and using the connected regions as candidate target locations for liveness; obtaining the liveness detection result based on the candidate target locations and the category probability distribution.

[0017] In the aforementioned implementation process, connected component analysis aggregates discrete, fragmented high-confidence points into complete candidate target regions. This effectively distinguishes situations where multiple independent live objects may exist in an image and separates each target, providing crucial technical support for multi-target detection and localization. This step yields an accurate estimate of the number of potential targets and their approximate spatial distribution, laying a solid foundation for subsequent fine-grained classification and verification, and significantly improving detection accuracy in multi-target scenarios.

[0018] Optionally, in this embodiment of the application, obtaining a liveness detection result based on the candidate target location and category probability distribution includes: assigning category probabilities to the candidate target locations based on the category probability distribution to obtain association results; the association results characterize the category probability corresponding to each candidate target location; calculating the comprehensive confidence level of the candidate target locations based on the association results; the comprehensive confidence level is used to characterize the probability that the candidate target location is a live object; and obtaining a liveness detection result based on the comprehensive confidence level.

[0019] In the above implementation process, by fusing independent confidence information from the detection and classification branches, the comprehensive confidence score provides a more reliable evaluation criterion than a single confidence score. It fully utilizes the model's multi-faceted judgment capabilities, effectively filtering out unreliable candidate targets that are "seemingly accurately located but with ambiguous categories" or "with certain category judgments but weak location signals," thereby reducing false detections and false alarms, improving the confidence level of the final output results, and enhancing the overall robustness of the system.

[0020] Optionally, in this embodiment of the application, the use of ultra-wideband radar to collect raw data streams inside the vehicle includes: emitting pulse signals using ultra-wideband radar and receiving reflected signals; and sampling the reflected signals using a high-speed analog-to-digital converter to obtain the raw data streams.

[0021] In the above implementation process, the sampling process performed by the high-speed analog-to-digital converter realizes the digitization of continuous analog signals, which can capture the detailed waveform of nanosecond-level pulse signals, reduce information loss, and provide the most original and most accurate data foundation for all subsequent advanced signal processing algorithms.

[0022] Secondly, embodiments of this application also provide a radar-based in-vehicle liveness detection device, comprising: an acquisition module for acquiring raw data streams inside the vehicle using ultra-wideband radar; the raw data streams include amplitude and phase information of the received signals at each time point; a construction module for constructing a three-dimensional tensor based on the raw data stream information; the dimensions of the three-dimensional tensor include the number of slow-time frames, the number of range gates, and the number of channels; a model detection module for inputting the three-dimensional tensor into a trained liveness detection model to obtain model output; the liveness detection model includes a detection branch and a classification branch; the detection branch is used to output a heatmap representing the spatial location of the liveness; the classification branch is used to output the category probability distribution of the liveness; and a detection result module for obtaining liveness detection results based on the heatmap and the category probability distribution.

[0023] Thirdly, embodiments of this application also provide a computer program product, including computer program instructions, which are executed by a processor to perform the method provided in the first aspect or any implementation thereof.

[0024] Fourthly, embodiments of this application also provide an electronic device, including: a processor and a memory, the memory storing computer program instructions, which are executed by the processor to perform the method provided in the first aspect or any implementation thereof.

[0025] Fifthly, embodiments of this application also provide a computer-readable storage medium storing computer program instructions, which, when executed by a processor, perform the method provided in the first aspect or any implementation thereof.

[0026] This application presents a radar-based in-vehicle liveness detection method, program product, electronic device, and storage medium. Ultra-wideband radar possesses extremely high range resolution and is highly sensitive to minute movements, enabling it to capture minute displacements caused by vital signs such as breathing and heartbeat. It mitigates the potential information loss risk associated with traditional methods using infrared or temperature sensors, providing a high-fidelity data foundation for subsequent processing and improving the accuracy and reliability of subsequent liveness feature analysis. A three-dimensional tensor is used to display and represent information. The in-vehicle environment is complex, with various interferences such as air conditioning airflow, swaying suspended objects, and pedestrian movement outside the vehicle. Ultra-wideband radar can effectively suppress most irrelevant environmental interference through range gating, enhancing the model's ability to perceive the subtle movement features of the target. Through multi-dimensional information fusion (heatmap and category probability distribution), false alarms are reduced, improving the reliability and practicality of the detection results. Attached Figure Description

[0027] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0028] Figure 1 A schematic flowchart of a radar-based in-vehicle liveness detection method provided in an embodiment of this application;

[0029] Figure 2 A schematic diagram of the structure of the radar-based in-vehicle liveness detection device provided in the embodiments of this application;

[0030] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0031] The embodiments of the technical solution of this application will now be described in detail with reference to the accompanying drawings. These embodiments are only used to more clearly illustrate the technical solution of this application and are therefore merely examples, and should not be used to limit the scope of protection of this application.

[0032] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this application.

[0033] In the description of the embodiments of this application, technical terms such as "first" and "second" are used only to distinguish different objects and should not be construed as indicating or implying relative importance or implicitly specifying the number, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly defined.

[0034] To prevent drivers from accidentally leaving pets or children in the vehicle and causing danger, existing vehicles typically use detection methods such as temperature sensors, oxygen concentration detectors, and window lift detectors to detect objects inside the vehicle, or they use cameras to analyze images inside the vehicle and detect objects.

[0035] However, temperature sensors, oxygen concentration detectors, and window regulator detectors are complex to install and inconvenient. Furthermore, due to the inherent characteristics of cameras, they cannot detect obstructions or objects like the trunk, leading to numerous false alarms and missed detections. Infrared sensors have low sensitivity, and camera-based detection algorithms lack privacy considerations.

[0036] This application provides a radar-based in-vehicle liveness detection method using ultra-wideband radar.

[0037] Before introducing the embodiments of this application, let's first introduce the ultra-wideband radar in these embodiments: The ultra-wideband radar in this application refers to UWB radar, which is a radar system that utilizes ultra-wideband (UWB) signals. It differs significantly from traditional narrowband radars (such as automotive radars and weather radars) in its working principle and application scenarios.

[0038] The following is a detailed introduction to UWB radar:

[0039] 1. Definition: Ultra-wideband (UWB) signals refer to radio technology that transmits very low-energy radio signals over a wide frequency range (typically >500 MHz). Its relative bandwidth (bandwidth to center frequency ratio) is typically greater than 20%, or its absolute bandwidth is at least 500 MHz.

[0040] 2. Key characteristics of ultra-wideband (UWB) radar: These include extremely wide bandwidth, low power spectral density, strong penetration capability, high range resolution, and good multipath resolution. Extremely wide bandwidth refers to extremely high temporal resolution. Strong penetration capability means that UWB signals have a certain ability to penetrate non-metallic materials (such as walls, clothing, human tissue, plastics, wood, and soil). This makes UWB radar a potential application in areas such as through-wall imaging, liveness detection, and buried target detection.

[0041] The core principle of UWB radar in detecting living beings inside a vehicle does not rely on optical "seeing," but rather on its ability to transmit extremely high-bandwidth nanosecond-level non-sinusoidal pulse electromagnetic waves that can penetrate non-metallic media and sense subtle movements behind them.

[0042] Please see Figure 1 The illustrated diagram shows a flowchart of a radar-based in-vehicle liveness detection method provided in an embodiment of this application. The radar-based in-vehicle liveness detection method provided in this application can be applied to electronic devices installed inside a vehicle. These electronic devices can be intelligent devices within the vehicle, or physical devices such as standalone PCs, tablets, or smartphones, or virtual devices such as virtual machines or containers. The electronic device can be a single device, a combination of multiple devices, or a cluster of a large number of devices. The radar-based in-vehicle liveness detection method may include:

[0043] Step S110: Use ultra-wideband radar to collect raw data streams inside the vehicle; the raw data streams include amplitude and phase information of the received signal at each time point.

[0044] Step S120: Construct a three-dimensional tensor based on the original data stream information; the dimensions of the three-dimensional tensor include the number of slow-time frames, the number of distance gates, and the number of channels.

[0045] Step S130: Input the three-dimensional tensor into the trained liveness detection model and obtain the model output; the liveness detection model includes a detection branch and a classification branch; the detection branch is used to output a heatmap representing the spatial location of the liveness; the classification branch is used to output the class probability distribution of the liveness.

[0046] Step S140: Obtain the liveness detection results based on the heatmap and category probability distribution.

[0047] In step S110, the ultra-wideband radar transmits non-sinusoidal pulse electromagnetic waves and receives echo signals reflected by objects and living individuals inside the vehicle. The radar receiver uses orthogonal demodulation technology to decompose the echo signals into in-phase and quadrature components, which are then sampled by a high-speed analog-to-digital converter to form a continuous time-series voltage value, i.e., the raw data stream. Each sampling point in this data stream simultaneously contains both signal amplitude and phase shift information. The amplitude information of the raw data stream reflects the target reflection intensity, while the phase information reflects the target's subtle movements.

[0048] As one implementation method, the vehicle-mounted ultra-wideband (UWB) radar module can be installed in the headliner, rearview mirror, or A / B pillars of the vehicle. The UWB radar periodically emits UWB pulses. In another implementation method, if two or more UWB radar modules are installed in the vehicle, these modules will work alternately in a time-sharing manner during liveness detection. When the pulse wave propagates inside the vehicle, it will encounter various obstacles (such as seat backs, blankets) and living targets, such as lying children, adults, or animals.

[0049] UWB pulses can penetrate non-metallic materials (such as fabric and foam) in the front seats and reach the rear seats. When the pulse encounters an interface where the dielectric constant changes (such as from air to human skin or from a blanket to a child's body), some of the energy is reflected back to the radar receiving antenna.

[0050] For example, pulse waves emitted by an ultra-wideband radar penetrate the back of the front seat and illuminate a child curled up in the back. The minute rises and falls of the child's chest due to breathing and heartbeat alter the distance between the radar and the child's chest, causing a slight change in the phase of each pulse echo. These continuous echo signals are recorded in the raw data stream. Therefore, although the radar's "line of sight" is obstructed by the seat, by acquiring a series of temporally continuous echo phase changes (raw data stream) through the ultra-wideband radar, a living person and their vital signs can be detected indirectly but precisely.

[0051] In step S120, firstly, based on the radar's pulse repetition period, the continuous data stream is divided into periods. The sampling points within each period form a row of a matrix, creating a two-dimensional matrix. The rows of the two-dimensional matrix represent all the echo signals received after the ultra-wideband radar transmits a series of pulse codes; the columns of the two-dimensional matrix represent the range gates. The range gates are used to describe the physical distance from the nearest to the farthest in the echo signal corresponding to a pulse. The dimensions of the two-dimensional matrix are the number of pulses and the number of range gates.

[0052] After obtaining the two-dimensional matrix, the in-phase and orthogonal components of each data point are separated to form independent channel dimensions, thus adding channel dimensions to the two-dimensional matrix and forming an initial three-dimensional tensor; the three dimensions of the initial three-dimensional tensor are the number of pulses, the number of range gates, and the number of channels, respectively.

[0053] Finally, for the initial three-dimensional tensor, the pulse dimension is divided into windows of fixed time length, and the number of pulses within each window represents the number of slow-time frames. Through the above processing, the original data stream is transformed into a structured three-dimensional tensor, whose three dimensions represent the number of slow-time frames, the number of distance gates, and the number of channels, respectively. This structured processing explicitly expresses the distance, time, and signal feature information implicit in the original data, which is beneficial for the subsequent efficient extraction of spatiotemporal features by the neural network and improves the model's ability to perceive the subtle movement features of the target.

[0054] In step S130, a pre-trained liveness detection model is used to process the three-dimensional tensor. The three-dimensional tensor can be an end-to-end deep learning model that includes a shared feature extraction network and two parallel task output branches (a detection branch and a classification branch).

[0055] Feature extraction networks automatically learn hierarchical features in data through multi-layer convolution operations. Shallow convolution kernels extract local spatiotemporal features, while deep networks combine simple features to form high-level semantic features.

[0056] The detection branch receives high-level features and generates a heatmap, which is a two-dimensional matrix whose size is usually the same as or similar to the spatial size (slow time × distance) of the input tensor. The value of each pixel on the heatmap is a confidence score between 0 and 1, which intuitively represents the probability of a live target being present at that location; the brighter the pixel, the higher the probability.

[0057] The classification branch analyzes the same feature and outputs the probability distribution of different liveness categories. The category probability distribution includes the category to be classified and the probability corresponding to that category.

[0058] The two output heads (detection branch and classification branch) share feature representations but have different functions and outputs. This design is beneficial for achieving collaborative optimization of spatial location and classification tasks, and improves the accuracy of multi-target detection and classification in complex scenarios.

[0059] In step S140, thresholding can be applied to the heatmap to filter out low-confidence regions, and then a connectivity analysis algorithm can be used to extract potential candidate target locations. For each candidate target location, its spatial location is mapped to a category probability distribution matrix to obtain the category confidence of the corresponding region.

[0060] The system integrates heatmap confidence and classification confidence, and verifies the results by incorporating vital sign frequencies regressed from phase information. Only targets that simultaneously meet the confidence threshold and the reasonableness of the physiological frequency are retained. The final result is a structured detection report containing the number of targets, their spatial location, category attributes, and vital sign parameters.

[0061] In the implementation of the above embodiments: Ultra-wideband radar possesses extremely high range resolution and is highly sensitive to micro-motions, enabling it to capture minute displacements caused by vital signs such as breathing and heartbeat. This mitigates the potential information loss risk associated with traditional methods using infrared or temperature sensors, providing a high-fidelity data foundation for subsequent processing and improving the accuracy and reliability of subsequent liveness detection. A three-dimensional tensor provides a structured representation of radar echo information (the in-vehicle environment is complex, with various interferences such as air conditioning airflow, swaying suspended objects, and pedestrian movement outside the vehicle; ultra-wideband radar can effectively suppress most irrelevant environmental interferences through range gates), enhancing the model's ability to perceive the micro-motion characteristics of the target. By fusing the location information from the heatmap and the semantic information from the classification branches, joint discrimination of live targets is achieved, reducing false alarms and improving the reliability and practicality of the detection results.

[0062] Optionally, in this embodiment of the application, constructing a three-dimensional tensor based on the original data stream information includes:

[0063] The original data stream is divided according to the pulse repetition period to obtain a two-dimensional matrix. The rows of the two-dimensional matrix represent all the echo signals received after the ultra-wideband radar sends a series of pulse codes. The columns of the two-dimensional matrix represent the range gates. The range gates are used to describe the physical distance from the nearest to the farthest in the echo signal corresponding to a pulse.

[0064] Pulse repetition period: When an ultra-wideband radar is operating, it periodically emits extremely short electromagnetic pulses. The time interval between two adjacent pulses is called the pulse repetition period. Its reciprocal is the pulse repetition frequency. A two-dimensional matrix is ​​a matrix composed of rows and columns. In the embodiments of this application, the two-dimensional matrix is ​​an intermediate data form formed after the initial structuring and processing of the one-dimensional raw data stream.

[0065] The raw data output by an ultra-wideband radar is a relatively long, continuous one-dimensional voltage value sequence. First, this one-dimensional sequence is precisely and continuously segmented according to the number of sampling points corresponding to a pre-set pulse repetition period. Each segment extracts a data segment containing the sampled values ​​of all echo signals acquired by the receiver within the entire receiving window after the radar transmits a series of pulse codes. This data segment is used as a row in a two-dimensional matrix. Rows in the two-dimensional matrix represent radar frames, which are a basic cycle of radar operation; for example, one frame may correspond to the transmission and reception of a series of pulse codes. The data segment corresponding to the next pulse cycle is used as the next row, and so on, ultimately assembling the data from a continuous time period into a two-dimensional matrix.

[0066] Each column of the two-dimensional matrix represents the change of the echo signal (pulse sequence) over time at each range cell from the nearest to the farthest. Therefore, the number of columns in the two-dimensional matrix is ​​the number of range gates.

[0067] In this process, a slow time dimension (i.e., radar frames) was constructed. This dimension directly corresponds to the real physical time series, enabling the model to observe and analyze the motion changes of the target within N consecutive frame periods. This provides the possibility of detecting subtle movements of vital signs such as breathing and heartbeat, and is the basis for perceiving the dynamic characteristics of living organisms.

[0068] Range gate: The radar receiver performs high-speed sampling of a single echo signal, and each sampling point is called a range gate. Each range gate corresponds to a specific time delay, which, after conversion using the speed of light, corresponds to a specific physical distance. The number of range gates determines the radar's detection range and range resolution.

[0069] For each data segment (containing N frames) obtained in the above steps, each frame is further processed. The echo data of a single frame (i.e., a complete CIR channel impulse response) is arranged into a vector according to the order of the sampling points. The length of this vector is the range gate number. This establishes the second dimension of the three-dimensional tensor: the range gate number. It maps the echo signal onto a defined physical space, enabling the model to accurately perceive the target reflection intensity at different range units, thereby achieving spatial localization of live targets and providing a physical basis for distinguishing targets in different positions inside the vehicle.

[0070] A receiving antenna is a physical antenna on a radar system used to receive reflected signals. A system can have one or more receiving antennas, each providing an independent signal receiving channel. The signal received by each receiving antenna is decomposed into two baseband signals, I (In-phase) and Q (Quadrature-phase), by an orthogonal demodulator. A complex number (I + jQ) can fully characterize the amplitude and phase of this signal.

[0071] The data for each range gate originates from a specific receiving antenna and contains both I and Q components. Therefore, the total number of channels = the number of receiving antennas × 2.

[0072] In the implementation of the above embodiments: the three-dimensional tensor greatly enriches the information capacity, and the explicit definition of the range gate directly links the data index to the physical space, enabling the algorithm to perform spatial positioning. Furthermore, by preserving complete channel information, rather than just calculating amplitude, the phase information of the signal is fully preserved, improving the detection accuracy of subtle movements of vital signs. By introducing the dimension of slow-time frames, where each "frame" is a snapshot of the radar, multiple consecutive frames record the change history of the echo signal at each range unit over slow time. This allows the subsequent liveness detection model to utilize its powerful spatiotemporal feature extraction capabilities to capture the phase changes of the echo signal over continuous time. From this dimension, the periodic phase modulation patterns caused by respiration and heartbeat can be clearly seen, thereby achieving reliable detection and identification of vital signs and greatly enhancing the system's ability to perceive subtle movements of life.

[0073] Optionally, in this embodiment, the liveness detection model further includes an input layer, a first convolutional layer, and a second convolutional layer; the detection branch and classification branch in the liveness detection model are two branches set in parallel after the second convolutional layer;

[0074] The input layer receives the 3D tensor. It receives the 3D tensor constructed in the preprocessing stage, verifies whether the input data conforms to the shape specification, and then feeds it into the subsequent first convolutional layer for computation.

[0075] The first convolutional layer is used to capture temporal features of the three-dimensional tensor in the slow-time frame dimension and spatial features of the three-dimensional tensor in the distance gate dimension, generating a spatiotemporal feature map.

[0076] Temporal characteristics refer to the regularity or pattern of a signal's changes over time, which manifests in radar data as periodic micro-movements caused by vital signs such as breathing and heartbeat. Spatial characteristics refer to the regularity of a signal's spatial distribution, which manifests in radar data as the distribution of target reflection intensity at different range cells, i.e., the target's shape, size, and position. The spatiotemporal feature map is the output of the first convolutional layer; it is a multi-channel two-dimensional feature map, where each channel represents the primary spatiotemporal features extracted from the input data after filtering by a specific convolutional kernel.

[0077] The process of spatiotemporal feature map extraction is as follows: The first convolutional layer uses multiple (e.g., 8) convolutional kernels of size (5,3). The height (5) of the convolutional kernel corresponds to the slow time dimension, designed to cover a continuous pulse period to capture temporal features (such as periodicity). The width (3) of the convolutional kernel corresponds to the distance dimension, designed to cover several adjacent distance units to capture spatial features (such as target contours). The core of the first convolutional layer is spatiotemporal joint perception. The number and size of the above convolutional kernels can be determined according to the actual situation.

[0078] The convolutional kernel slides across the slow-time frame-distance gate plane of the 3D tensor with a certain stride (e.g., 2). At each rest position, the weights inside the convolutional kernel are multiplied element-wise with the data in their covered region, summed, and a bias term is added to form an output value. This computation process traverses the entire input plane, generating a 2D feature map, i.e., a spatiotemporal feature map.

[0079] The second convolutional layer is used to extract deep features from the spatiotemporal feature map, obtaining a deep feature map. The deep feature map is the output of the second convolutional layer. The second convolutional layer has more channels, but its spatial dimensions (slow temporal dimension and distance dimension) may be smaller. Therefore, the deep feature map contains more abstract and higher-level semantic information than the first layer feature map.

[0080] The process of deep feature map extraction is as follows: a deeper convolution operation is applied to the deep feature map, typically using a smaller convolution kernel (e.g., 3x3) and more output channels (e.g., 16). The operation is similar to the first convolution layer, except that the first convolution layer extracts features based on the original data stream, while the second convolution layer extracts features based on the deep feature map.

[0081] The detection branch receives depth feature maps and performs calculations on them to generate heatmaps. The detection branch is the branch in the model responsible for the target spatial localization task.

[0082] Heatmap generation process: 1x1 convolutional layers are used to operate on the depth feature map. The main function of the 1x1 convolution is to fuse feature information from all channels (e.g., weighted combination of features from 16 channels) and project it onto a single channel. An upsampling operation is then performed to enlarge its size, matching it to the spatial dimensions of the original input, thus generating the final high-resolution heatmap.

[0083] As one implementation method, the Sigmoid activation function can also be used to constrain the output value to the range of [0,1] to form the confidence probability value of the pixel or heatmap point.

[0084] The classification branch receives the deep feature map and generates a probability distribution for the liveness category based on the deep feature map. The classification branch is the branch in the model responsible for the target identification task.

[0085] In one implementation, the obtained deep feature map is input into one or more fully connected layers. The fully connected layers perform a global weighted combination of all input features, and through a non-linear transformation, ultimately map it to the class space. Finally, the output of the fully connected layers is converted into a normalized class probability distribution using the Softmax activation function.

[0086] In the implementation of the above embodiments: the first convolutional layer, through its unique spatiotemporal joint convolution kernel, can effectively extract spatiotemporal local features related to living entities directly from the raw, cluttered radar signals, enhancing and highlighting valuable signal components in the input data while suppressing irrelevant background noise, thus improving the model's initial perception ability of micro-motion signals. The second convolutional layer, through further refinement and compression of the primary features, generates a deep feature map that can represent the high-level semantic information of the target, improving the model's ability to distinguish between different living entities and complex backgrounds. The detection branch is used to analyze the spatial location information contained in the deep feature map and achieves effective localization of living entities inside the vehicle through heatmaps. The classification branch is responsible for interpreting the semantic category information contained in the deep feature map and outputting the probability distribution. The two branches achieve collaborative optimization through sharing a feature extraction layer, improving the ability to make fine distinctions.

[0087] Optionally, in this embodiment, the classification branch includes a time-frequency analysis subnet; the time-frequency analysis subnet is used to determine frequency components from the deep feature map, and to determine the category probability distribution of the living organism based on the frequency components and a preset physiological frequency knowledge base.

[0088] Frequency components refer to the various vibrational modes of different frequencies contained in a signal. In the context of liveness detection, it specifically refers to the periodic micro-movements of specific frequencies hidden in the radar echo phase caused by vital activities such as breathing and heartbeat. The physiological frequency knowledge base is a database built upon prior medical and biological knowledge. It includes reasonable ranges of vital sign frequencies for different categories of living organisms, such as the typical respiratory rate range for adults, children, and common pet respiratory and heart rate ranges.

[0089] The time-frequency analysis subnet receives deep feature maps and processes them using layers with Fast Fourier Transform (FFT) or layers with one-dimensional convolutional kernels sliding along the slow time dimension to learn frequency features, thereby obtaining frequency components. The goal is to transform the sequence features (i.e., the change pattern of the target over time) in the slow time dimension from the time domain to the frequency domain.

[0090] The calculated frequency components are compared with predefined standard ranges for each category in a physiological frequency knowledge base. For example, if the calculated dominant frequency falls within the respiratory frequency range of "adults," evidence favorable to the "adult" category is generated. This frequency-based evidence is then weighted and fused or logically decided upon with the category probabilities calculated by the main pathway of the classification branch (through analysis of features such as shape and size via fully connected layers). The fusion method can be a simple weighted average or a more complex attention-based fusion, ultimately influencing and correcting the final category probability distribution output.

[0091] In the implementation of the above embodiments: the time-frequency analysis subnet extracts frequency components from the deep feature map through fast Fourier transform and compares them with a preset physiological frequency knowledge base as the basis for distinguishing the liveness category, thereby enhancing the scientificity and reliability of the classification decision. It effectively solves the classification ambiguity problem when the target is occluded, has an unusual posture, or is difficult to distinguish based solely on body shape characteristics (e.g., a curled-up child and a small pet dog). By adding a judgment dimension based on physiological and physical laws, it significantly improves the accuracy of fine-grained liveness classification and the system's anti-interference capability.

[0092] Optionally, in this embodiment, the classification branch includes spatial attention; spatial attention is used to generate a weight map based on the heatmap, and the weight map is used to determine the importance of different regions in the depth feature map.

[0093] Spatial attention is a resource allocation strategy in neural networks that simulates human attention behavior, allowing the network to automatically and dynamically focus on the most relevant parts of the input information while suppressing irrelevant background information. The weight map is a matrix identical in spatial dimensions (height and width) to the input feature map. Each value in the matrix is ​​a weight coefficient (typically between 0 and 1); a larger value indicates greater importance of that spatial location in subsequent calculations.

[0094] For example, affine transformations, scaling, normalization, or even a small neural network can be used to generate a weight map based on the heatmap. The generation logic of the weight map is to map regions with high confidence in the heatmap to high weights (close to 1) and regions with low confidence (background or noise) to low weights (close to 0).

[0095] The generated weight map is then spatially multiplied element-wise with the depth feature map. For example, all channel feature values ​​at each spatial location in the depth feature map are multiplied by the corresponding weight coefficient on the weight map. After weighting, the feature values ​​of important regions in the depth feature map are enhanced, while the feature values ​​of minor and background regions are weakened or even reduced to zero.

[0096] In one implementation, the weighted deep feature map is fed into fully connected layers following the classification branch for final classification. The classification branch can simultaneously consider the importance of different regions in the deep feature map while generating the probability distribution of the liveness category.

[0097] In the implementation of the above embodiments: the spatial attention mechanism, by introducing prior localization information from the detection branch, guides the classification branch to concentrate limited computational resources and attention on the region most effective for classification decisions, thus achieving synergy between localization and recognition tasks. Simultaneously, it effectively suppresses the activation of irrelevant features generated by complex in-vehicle backgrounds (such as swaying ornaments or shaking seat covers), greatly reducing the misleading influence of background interference on classification judgments, thereby significantly improving the accuracy and robustness of category discrimination.

[0098] Optionally, in this embodiment of the application, before inputting the three-dimensional tensor into the trained liveness detection model and obtaining the model output, the method further includes:

[0099] Ultra-wideband radar was used to collect sample data of living beings in different states inside the vehicle. These different states refer to living beings in different postures (e.g., sitting, lying down, curled up), different types (e.g., adults, children, pets), and located in different positions within the vehicle (driver's seat, rear seats, floor mats). The sample data consists of raw digital signal sequences obtained from radar acquisition and analog-to-digital conversion, containing amplitude and phase information of the echoes.

[0100] For example, systematic data acquisition is conducted in a real vehicle environment. First, a test scenario is set up, with different types of live subjects positioned in designated locations within the vehicle under various preset conditions. Then, the ultra-wideband radar system is activated, operating continuously at a fixed pulse repetition frequency and sampling rate, emitting electromagnetic pulses and receiving echoes. The radar receiver converts the echo signals into in-phase and quadrature digital signals using quadrature demodulation technology, which are then recorded and saved by a data acquisition card or embedded processor, forming the initial sample dataset.

[0101] Considering that real-world data collection cannot cover all extreme situations (such as an infant completely covered by a thick blanket), and is extremely costly and time-consuming, especially in vehicle-based liveness detection, which involves many specific interferences, such as air conditioning vents or swaying suspended objects inside a moving vehicle, radar may detect subtle movements that affect the judgment of liveness. Therefore, this application's embodiments generate physical adversarial examples through simulation, enabling the model to learn the impact of these interferences on liveness detection during the training phase. This allows the model to make accurate judgments in real-world detection scenarios, greatly improving the model's robustness and generalization ability.

[0102] Physical adversarial examples are generated through simulation. The simulated scenarios include at least one of the following: air conditioning airflow, swaying of suspended objects, or moving objects outside the vehicle. Physical adversarial examples do not refer to simple digital noise, but rather signal data generated based on radar sensing physics principles (such as electromagnetic wave propagation, scattering, and interference models) to simulate complex interference and extreme scenarios in the real world.

[0103] First, accurate physical models are established for various interference scenarios. For example, for air conditioning airflow interference, the physical model can be a broadband, time-varying incoherent noise field. During simulation, a Gaussian random noise field is calculated based on parameters such as air velocity and air outlet direction, and then superimposed on the radar echo signal of a real living object.

[0104] For suspended object swaying (such as interior decorations in a car): the physical model is one or more point scatterers with period, sway amplitude, and radar cross section. During simulation, the distance delay and Doppler frequency shift over time are calculated based on the sway pattern, generating the corresponding echo signal, which is then superimposed with the echo signal of a real living object.

[0105] For moving objects outside the vehicle (such as pedestrians on the roadside): the physical model is a strong scatterer located outside the radar detection range but still within the beam illumination range. During the simulation, the object is simulated to move at a certain speed, generating a strong signal with a significant Doppler frequency shift, which is received simultaneously at the receiver and the signal inside the vehicle.

[0106] When generating physical adversarial examples, the above physical model is used to modulate and superimpose the real sample data to create adversarial examples with various interferences and different signal-to-noise ratios.

[0107] Labeling is performed on both the sample data and the adversarial examples, resulting in labeled data. The labeled data includes location labels and category labels. Labeling can be a manual, semi-automatic, or automatic process of assigning truth values ​​to the data.

[0108] Annotation tools are used to label the radar data, displaying multiple views (such as range-time plots and range-Doppler spectra). First, the presence and location of live targets are determined based on energy concentration and micro-Doppler characteristics in the data. For each frame containing a live target, the location label is determined by drawing a rectangle tightly enclosing the target region on the range-time 2D image; the category label is manually selected by the annotator based on experimental settings (or feature-based judgment). For labels of physical adversarial examples, the labels are consistent with the labels of the original, clean sample data upon which they are based, as the addition of interference does not change the location or category of the live target itself. All labels undergo multiple verifications to ensure accuracy.

[0109] Based on labeled data, a pre-defined initial model is trained using a multi-task loss function to obtain a trained liveness detection model. The multi-task loss function simultaneously considers and combines the losses from multiple learning tasks (such as object detection and object classification).

[0110] The training process is an iterative optimization loop. In each iteration, forward propagation takes a batch of sample data from the labeled data and inputs it into the initial model. The data is then processed sequentially through each layer of the model, ultimately yielding heatmap predictions from the detection branch and class probability predictions from the classification branch.

[0111] The model calculates the loss by comparing its predicted values ​​with the actual label data (true values) corresponding to the samples. A multi-task loss function is enabled, which typically consists of a weighted sum of two parts: a localization loss (such as mean squared error loss or Focal Loss) that calculates the difference between the predicted heatmap and the actual location labels, and a classification loss (such as cross-entropy loss) that calculates the difference between the predicted class probabilities and the actual class labels.

[0112] Then, backpropagation and parameter updates are performed: the gradient of the total loss with respect to all model parameters is calculated. Using the gradient descent algorithm, the gradient information is backpropagated to each layer, and the model's weights and biases are fine-tuned based on the gradient direction and learning rate to reduce the total loss.

[0113] The model continues until the loss converges to a low value. At this point, the model has learned to accurately extract features from the input data and complete the localization and classification tasks, thus obtaining a well-trained liveness detection model.

[0114] In the implementation of the above embodiments: through simulation, a massive amount of training samples covering various rare interferences and extreme scenarios can be generated at low cost and high efficiency. Physical adversarial examples are generated, and these challenging scenarios are actively retained during the model training phase, forcing the model to learn how to ignore interference and focus on true liveness features. This greatly enhances the model's adaptability and generalization ability in real complex environments, and significantly improves the system's anti-interference capabilities and overall detection accuracy.

[0115] Optionally, in this embodiment, the heatmap includes the confidence level of the points; based on the heatmap and the category probability distribution, the liveness detection result is obtained, including:

[0116] Points in the heatmap with confidence levels below a preset threshold are removed, leaving only the retained points. The preset threshold is a preliminary threshold for judging the authenticity of a signal; this threshold can be adjusted and optimized based on the model's performance on the validation set. The retained points are the set of pixels remaining on the heatmap after threshold filtering, with confidence levels higher than or equal to the preset threshold.

[0117] Connectivity analysis algorithms are used to determine connected regions in a heatmap based on preserved points, and these connected regions are used as candidate locations for live targets. Connectivity analysis is an image processing algorithm used to identify all interconnected foreground pixel blocks in a binary image. Connection methods typically include 4-connectivity (considering only vertical and horizontal adjacency) or 8-connectivity (additionally considering diagonal adjacency).

[0118] Starting with an unlabeled pixel, the algorithm searches for all adjacent preserved pixels and groups all found connected pixels into the same region, assigning each region a unique identifier. This process continues until all preserved pixels in the image have been visited and categorized. The algorithm then calculates the minimum bounding rectangle for each connected region, or the centroid coordinates of all pixels within that region. This bounding rectangle or centroid coordinates is defined as a candidate target location.

[0119] One implementation method is to compare connected regions with preset region thresholds. If a connected region is smaller than the minimum region threshold or larger than the maximum region threshold, it is discarded and not considered as a candidate target location. This filters connected regions based on their area, excluding regions that are not live.

[0120] In the implementation of the above embodiments: connected component analysis aggregates discrete, fragmented high-confidence points into complete candidate target regions. This effectively distinguishes situations where multiple independent live objects may exist in an image and separates each target, providing key technical support for multi-target detection and localization. This step obtains an accurate estimate of the number of potential targets and their approximate spatial distribution, laying a solid foundation for subsequent fine-grained classification and verification, and significantly improving the detection accuracy in multi-target scenarios.

[0121] Optionally, in this embodiment of the application, obtaining the liveness detection result based on the candidate target location and category probability distribution includes:

[0122] Based on the category probability distribution, category probabilities are assigned to candidate target locations to obtain association results; the association results represent the category probability corresponding to each candidate target location.

[0123] First, obtain the spatial coordinate range of each candidate target location in the original data. Then, map these spatial coordinates to the same region on the category probability distribution map output by the classification branch. For each candidate target location, extract the category probability information within its corresponding region. This extraction can be done by calculating the average category probability of all pixels within the region, or by directly taking the probability value of the center point of the region. Finally, assign a complete category probability vector to each candidate target location. This vector represents the association result for that location, indicating the probability that the target at that location belongs to each category.

[0124] Based on the association results, the overall confidence level of the candidate target location is calculated; the overall confidence level is used to characterize the probability that the candidate target location is a living organism.

[0125] The calculation of the overall confidence score is an information fusion process. It requires two types of input information: one is the location confidence score from the detection branch, which reflects the probability of the existence of the target (i.e., the average or maximum response value of the candidate target area in the heatmap); the other is the classification confidence score from the classification branch, which reflects the certainty of the category judgment (i.e., the maximum value of the probability vector in the association result, i.e., the probability of the most likely category).

[0126] The two confidence scores can be combined into a single comprehensive score using a weighted geometrical mean or a weighted arithmetic mean. For example, the comprehensive confidence score = location confidence score × classification confidence score. This multiplicative relationship means that the comprehensive confidence score will only be high if both the target and the classification scores are high. Different weights can also be assigned to the two scores, and these weights can be dynamically adjusted based on the specific scenario.

[0127] Liveness detection results are obtained based on the overall confidence level. For example, a final confidence threshold can be set, and each candidate target and its overall confidence level are compared with this final threshold. For candidate targets with an overall confidence level higher than the final threshold, liveness detection results are generated based on the category (the category with the highest probability in the association results), spatial location information, vital sign data (if available), and its overall confidence level.

[0128] In the implementation of the above embodiments: by fusing independent confidence information from the detection branch and the classification branch, the comprehensive confidence provides a more reliable evaluation criterion than a single confidence. It fully utilizes the model's multi-faceted judgment capabilities, effectively filtering out unreliable candidate targets that are "seemingly accurately located but with ambiguous categories" or "with certain category judgments but weak location signals," thereby reducing false detections and false alarms, improving the confidence level of the final output results and the overall robustness of the system.

[0129] Optionally, in this embodiment of the application, the use of ultra-wideband radar to collect the raw data stream inside the vehicle includes:

[0130] Ultra-wideband (UWB) radar emits pulse signals and receives reflected signals. The radar's pulse generator, under trigger conditions, produces an electromagnetic pulse conforming to UWB specifications. This pulse is radiated into the vehicle's interior via a transmitting antenna. The electromagnetic wave propagates at the speed of light within the vehicle and is reflected when it encounters surfaces with different dielectric constants. The radar's receiving antenna is specifically designed to capture these significantly attenuated reflected signals returning from various directions. The receiving antenna converts the captured weak electromagnetic wave energy into a corresponding voltage change signal, which is then fed into a low-noise amplifier for preliminary amplification to increase signal strength for subsequent processing.

[0131] The reflected signal is sampled by a high-speed analog-to-digital converter (ADC) to obtain the raw data stream. A high-speed ADC is an electronic component that rapidly converts continuously changing analog voltage signals into discrete digital values ​​within extremely short time intervals. The analog signal is immediately fed into the high-speed ADC. The core of the high-speed ADC is a high-precision comparator and encoding circuit that instantaneously measures the input analog voltage at a very high sampling rate. This process continues continuously, ultimately outputting a long, chronologically ordered sequence of discrete digital values—the raw data stream. Each number in this data stream uniquely corresponds to the amplitude of the radar-received signal at a specific moment.

[0132] In the implementation of the above embodiments: the sampling process performed by the high-speed analog-to-digital converter realizes the digitization of continuous analog signals, can capture detailed waveforms of nanosecond-level pulse signals, reduce information loss, and provide the most original and most accurate data foundation for all subsequent advanced signal processing algorithms.

[0133] This application embodiment simultaneously outputs the spatial location heatmap and category probability distribution of the liveness detection system in an end-to-end manner, and finally fuses the two results to achieve high-precision in-vehicle liveness detection, eliminating the complex manual feature extraction steps in traditional solutions. "End-to-end" means that the most raw data is directly input into the model, and the model can output the final required result. All the complex feature extraction and transformation steps in between are automatically completed by the model without human intervention.

[0134] Before deployment, the trained liveness detection model can undergo model compression and optimization. This includes using knowledge distillation techniques to guide a lightweight student network to learn with a large teacher network, and performing 8-bit integer quantization and weight sparsification operations to compress the model size. The optimized model can be deployed in an in-vehicle microcontroller unit, achieving lower single inference time, better meeting real-time performance requirements, and satisfying the stringent requirements of automotive-grade hardware for computing power, power consumption, and reliability.

[0135] Please see Figure 2 The diagram shown is a structural schematic of a radar-based in-vehicle liveness detection device provided in an embodiment of this application; this embodiment of the application provides a radar-based in-vehicle liveness detection device 200, including:

[0136] The acquisition module 210 is used to acquire raw data streams inside the vehicle using ultra-wideband radar; the raw data streams include amplitude and phase information of the received signal at each time point;

[0137] Module 220 is used to construct a three-dimensional tensor based on the original data stream information; the dimensions of the three-dimensional tensor include the number of slow-time frames, the number of distance gates, and the number of channels;

[0138] The model detection module 230 is used to input a three-dimensional tensor into a trained liveness detection model and obtain the model output. The liveness detection model includes a detection branch and a classification branch. The detection branch is used to output a heatmap representing the spatial location of the liveness object. The classification branch is used to output the probability distribution of the liveness object's category.

[0139] The detection result module 240 is used to obtain the liveness detection result based on the heat map and category probability distribution.

[0140] It should be understood that this device corresponds to the radar-based in-vehicle liveness detection method embodiment described above, and is capable of performing the various steps involved in the above method embodiment. The specific functions of this device can be found in the description above, and detailed descriptions are omitted here to avoid repetition. The device includes at least one software functional module that can be stored in memory or embedded in the device's operating system (OS) in the form of software or firmware.

[0141] Please see Figure 3 The diagram shows a structural schematic of an electronic device provided in an embodiment of this application. An electronic device 300 provided in this application includes a processor 310 and a memory 320. The memory 320 stores machine-readable instructions executable by the processor 310. When the machine-readable instructions are executed by the processor 310, the method described above is performed.

[0142] Figure 3 The components shown can be implemented using hardware, software, or a combination thereof. Electronic device 300 may be a physical device, such as a server or PC, or a virtual device, such as a virtual machine or virtualization container. Furthermore, electronic device 300 is not limited to a single device; it can be a combination of multiple devices or a cluster of numerous devices.

[0143] This application also provides a storage medium storing a computer program, which is executed by a processor to perform the above-described method.

[0144] The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0145] This application also provides a computer program product, including computer program instructions, which are executed by a processor to perform the method described above.

[0146] It should be understood that the disclosed apparatus and methods can also be implemented in other ways, given the several embodiments provided in this application. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0147] In addition, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0148] The above description is only an optional implementation of the embodiments of this application, but the protection scope of the embodiments of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the embodiments of this application should be covered within the protection scope of the embodiments of this application.

Claims

1. A radar-based in-vehicle liveness detection method, characterized in that, include: Use ultra-wideband radar to collect raw data streams inside the vehicle; The raw data stream includes amplitude and phase information of the received signal at each time point; A three-dimensional tensor is constructed based on the original data stream information; the dimensions of the three-dimensional tensor include the number of slow-time frames, the number of distance gates, and the number of channels. The three-dimensional tensor is input into the trained liveness detection model to obtain the model output; the liveness detection model includes a detection branch and a classification branch; the detection branch is used to output a heatmap representing the spatial location of the liveness; the classification branch is used to output the category probability distribution of the liveness. Based on the heatmap and the category probability distribution, the liveness detection results are obtained; The liveness detection model further includes an input layer, a first convolutional layer, and a second convolutional layer; the detection branch and the classification branch in the liveness detection model are two branches set in parallel after the second convolutional layer; The input layer is used to receive the three-dimensional tensor; The first convolutional layer is used to capture temporal features of the three-dimensional tensor in the dimension of slow time frame number; and to capture spatial features of the three-dimensional tensor in the dimension of distance gate number, generating a spatiotemporal feature map; The second convolutional layer is used to extract deep features from the spatiotemporal feature map to obtain a deep feature map; The detection branch is used to receive the depth feature map and perform calculations on the depth feature map to generate a heat map; The classification branch is used to receive the depth feature map and generate the category probability distribution of the liveness based on the depth feature map; The classification branch includes a time-frequency analysis subnetwork; the time-frequency analysis subnetwork is used to determine frequency components from the depth feature map, and to determine the probability distribution of the category of the living organism based on the frequency components and a preset physiological frequency knowledge base; the classification branch also includes spatial attention; the spatial attention is used to generate a weight map based on the heatmap, and the weight map is used to determine the importance of different regions in the depth feature map.

2. The method according to claim 1, characterized in that, The construction of the three-dimensional tensor based on the original data stream information includes: The original data stream is divided according to the pulse repetition period to obtain a two-dimensional matrix; the rows of the two-dimensional matrix represent all the echo signals received after the ultra-wideband radar sends a series of pulse codes; the columns of the two-dimensional matrix represent the range gate number; the range gate number is used to describe the physical distance from the nearest to the farthest in the echo signal corresponding to a pulse; Based on the two-dimensional matrix, the three-dimensional tensor is obtained; the dimensions of the three-dimensional tensor are the number of slow-time frames, the number of distance gates, and the number of channels, respectively; the number of slow-time frames is used to capture the micro-motion features of the target over time.

3. The method according to claim 1, characterized in that, Before inputting the three-dimensional tensor into the trained liveness detection model and obtaining the model output, the method further includes: Ultra-wideband radar was used to collect sample data of living organisms in different states inside the vehicle; Physical adversarial scenarios are generated through simulation; the simulated scenarios include at least one of air conditioning airflow, swaying of suspended objects, or moving objects outside the vehicle. The sample data and the adversarial sample are labeled with corresponding tags to obtain labeled data; the tag data includes location tags and category tags. Based on the labeled data, a preset initial model is trained using a multi-task loss function to obtain the trained liveness detection model.

4. The method according to claim 1, characterized in that, The heatmap includes the confidence level of the points; based on the heatmap and the category probability distribution, the liveness detection result is obtained, including: Points in the heat map with a confidence level lower than a preset threshold are removed to obtain retained points; Using a connectivity analysis algorithm, connected regions in the heatmap are determined based on the reserved points, and these connected regions are used as candidate target locations for living organisms. Based on the candidate target location and the category probability distribution, the liveness detection result is obtained.

5. The method according to claim 4, characterized in that, Based on the candidate target location and the category probability distribution, a liveness detection result is obtained, including: Based on the category probability distribution, category probabilities are assigned to the candidate target locations to obtain association results; the association results represent the category probability corresponding to each candidate target location. Based on the association results, the overall confidence level of the candidate target location is calculated; the overall confidence level is used to characterize the probability that the candidate target location is a living organism. Based on the comprehensive confidence level, the liveness detection result is obtained.

6. A computer program product, characterized in that, It includes computer program instructions that are executed by a processor to perform the method as described in any one of claims 1 to 5.

7. An electronic device, characterized in that, include: A processor and a memory, the memory storing computer program instructions that, when executed by the processor, perform the method as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions that, when executed by a processor, perform the method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Image recognition method and device based on deep learning, equipment and storage medium

    CN112767366A

  • In-vehicle living body detection method and device using phase matching

    CN114527463A

  • Millimeter wave radar in-vehicle living body detection method based on decision fusion

    CN117214850A