Method for pre-processing radar data for further processing by a machine learning model to generate a control signal for controlling a device
A radar and spiking neural network combination addresses the limitations of camera-based smart doorbells by providing reliable human presence detection with low power consumption and extended battery life, enhancing user experience.
Patent Information
- Application Number
- PCT/EP2025/073214
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-16
- Filing Date
- 2025-08-13
- Publication Date
- 2026-02-19
AI Technical Summary
Current smart doorbell solutions rely heavily on camera-based motion detection, leading to false positives and increased power consumption, which negatively impacts battery life and user experience, and lack reliable human presence detection due to reliance on error-prone proxies.
A radar system combined with a spiking neural network for pre-processing radar data to detect human presence, allowing for low-power, motion-resistant detection with a low rate of false alarms, using a processing pipeline that includes generating data from transmitted and received signals, calculating differences in frequency range bins, and applying machine learning models like SNNs.
The system achieves robust human presence detection with high accuracy (>99%) and significantly extends battery life by operating at ultra-low power, maintaining detection latency without compromising aesthetics or buildability, and enabling flexible power management.
Smart Images

Figure EP2025073214_19022026_PF_FP_ABST
Abstract
Description
Method for pre-processing radar data for further processing by a machine learning model to generate a control signal for controlling a deviceTECHNICAL FIELD
[0001] This disclosure generally relates to pre-processing of radar data for further processing by a machine learning model for efficient automated target recognition, which may be used to generate a control signal to control a device, such as (but in no way limited to) a smart doorbell. More specifically, but not exclusively, the disclosure relates to a system and method for pre-processing of received radar signals for human presence detection recognition using a machine learning model with temporal dynamics. Upon human presence detection, a control signal may be generated that may be transferred to a device, such as a smart doorbell, which may e.g. be used to generate a notification that a person is present in the vicinity.BACKGROUND
[0002] A smart doorbell is a device which monitors an entrance area, e.g. a front door area, of a building, typically using an infrared (IR) camera, and alerts the owner if a person is present in the vicinity. Functionality can be extended by inclusion of a microphone and speaker for communication between the homeowner and a visitor, and security-improving features such as turning light on and producing alert sound to discourage intruders. With more extensive Al use, facial recognition can be used to recognize authorized people and unlock the door for them.
[0003] Current smart doorbell solutions rely heavily on camera-based “motion” detection as a surrogate for human presence. In result, false positives are triggered by events such as change in lighting conditions, wind-triggered movements, rain, and dust on the lens. Detections are also triggered by insects. Coincidentally, IR cameras, which are used to enable monitoring in poor lighting conditions, attract insects.
[0004] Mitigation strategies include use of multiple sensors, especially passive infrared (PiR) and radar, which rely on temperature and proximity, respectively. Overall, this leads to an increase in the Bill of Materials and power while still relying on extremely basic, error- prone proxies for human presence.
[0005] Lack of detection reliability has a major negative impact on user experience. This impact is mostly reported as unnecessary, false alerts. However, it further extends to filling up data storage with irrelevant footage, and reduction of battery life, as discussed next.
[0006] Smart doorbells can be wired or battery-operated. Battery-operated devices have a much reduced barrier to adopt, especially in rental properties.
[0007] Smart doorbell users report much lower battery lifetimes than those advertised. Battery life of a device is estimated based on the predicted workload, including how often detections are likely to happen for a given application. The current rate of false alarms has a user-perceptible negative impact on the battery longevity. Charging batteries becomes another chore, and is particularly inconvenient for devices which need to be taken down for charging.
[0008] Smart doorbells commonly rely on internet connectivity. This reliance is likely to grow further with the increasing use of Al in the cloud. Rural areas may suffer from unreliable and slow internet connection. Lack of reliability would result in poor user experience, decreasing the market adoption. While Al models can be run locally, the feasibility of this solution relies heavily on the power efficiency of their implementation, especially for battery-operated smart doorbells.
[0009] The existing smart doorbell solutions are commonly triggered by movement. This leads to decreased reliability of human presence detection due e.g. to an increase in false positives.
[0010] Radar uses radio waves to determine the distance (ranging), angle (azimuth and elevation), and radial velocity of objects relative to the radar system site. Radar is typically used to detect and track flying aircraft, spacecraft, missiles, etc. and to map weather formations and terrain. Recently small radar systems have also been used to detect object movements such a hand gestures for hands-free control of devices such as televisions.
[0011] A radar system consists of a transmitter producing electromagnetic waves in the radio or microwaves domain, a transmitting antenna, a receiving antenna (the same antenna can be used for transmitting and receiving) and a receiver and processor to determine properties of the objects. Radio waves (pulsed or continuous) from the transmitter reflect off the objects and return to the receiver, giving information about the objects’ locations and speeds.
[0012] Automated target recognition using machine learning models have been used for gesture recognition for hands-free control of devices. In previously proposed systems, the pre-processing of data for input to such machine learning models has involved the generation of spectrograms of data derived from the received radar signals. Examples of spectrograms used for these applications include range-Doppler, micro-Doppler, range-angle and angle- Doppler. These spectrograms comprise a rectangular map of the relevant features against a time axis.
[0013] In case of range-Doppler, creation of a spectrogram requires computation of range profiles and Doppler (radial velocity) profiles to generate a range-Doppler surface comprising a large set of data representing variation in features such as range and velocity with time. This involves the storing and processing of a constant number of samples into a joint 2D representation of the selected features. This spectrogram-based approach requires a large computational and memory load which remains constant regardless of the input data, imposing a constant drain on the memory and power of the data pre-processing system. The same limitations also apply when extending from 2D to multi-dimensional representation, e.g. range-Doppler-angle.
[0014] The document ZHANG BO ET AL: "Long-Range Real-Time Gesture Recognition for Millimeter Wave Radar", 2022 2ND INTERNATIONAL CONFERENCE ON FRONTIERS OF ELECTRONICS, INFORMATION AND COMPUTATION TECHNOLOGIES (ICFEICT), IEEE, 19 August 2022, pages 298-303, XP034234546, DOI: 10.1109 / ICFEICT57213.2022.00061 discloses a long-range millimeter wave radar gesture recognition method, which can be deployed on radar for real-time gesture recognition.Firstly, the moving target indication (MTI) is used to eliminate the influence of static clutters. Then the biaxial projection method is used to find the foreground images. The proposed method uses the histogram of oriented gradients (HOG) algorithm to extract features. Finally, support vector machine (SVM) is used to leam and classify features.
[0015] The SVM is an example of a non-temporal machine learning model. A non-temporal machine learning model may be used to classify radar data. However, such models operate on static input representations and lack internal memory mechanisms for modelling temporal relationships between successive data points. The inventors of the present application have realized that a major downside of these non-temporal machine learning models is that they treat each input independently and are unable to capture time-dependent patterns or motiondynamics that unfold across multiple frames or time steps. This limits their effectiveness in applications where temporal evolution of radar signals is a key feature, such as in detecting micro-movements, identifying gesture sequences, or monitoring vital signs.
[0016] The document US 2021 / 190902 Al discloses techniques and apparatuses that implement a smart-device-based radar system capable of performing symmetric Doppler interference mitigation. The radar system employs symmetric Doppler interference mitigation to filter interference artifacts caused by the vibration of the radar system or the vibration other objects. This filtering operation incorporates the interference artifact within the noise floor, without significantly attenuating reflections from a desired object.SUMMARY
[0017] To address the above discussed drawbacks of the prior art, the present invention is proposed. The attached claims determine the scope of protection of this patent.
[0018] The invention provides a radar and a processor system implementing a processing pipeline for processing the radar data, the processing pipeline including a machine learning model, which allows for detection of signatures of human presence regardless of movement and at low power and data storage requirements. The invention achieves this by preprocessing temporal radar data and feeding the resulting data to a machine learning model, such as a spiking neural network. Spiking Neural Networks (SNNs) excel at extracting patterns from spatio-temporal data. Each neuron and synapse contains time-weighted traces of the past samples, with the network being trained to combine them to efficiently turn sensor data into information and action.
[0019] The combination of a radar and a machine learning model, preferably a spiking neural network, enables realization of an ultra low power Al customer solution. The system itself operates with a low mW power envelope. However, significant power benefits are realized at the systems level, where these additional components more than pay for themselves. The radar and processing system may remain always on, scanning the environment for human presence. While the radar and processing system keeps the watch, more power hungry components of e.g. the doorbell, including the camera, can be put into sleep. Upon detection of human presence, the output of the machine learning model can be used to coordinate the desired actions such as waking up other system components, alerting the home owner and triggering camera recording.
[0020] The proposed setup significantly extends the battery life of the doorbell thanks to the low average power combined with robust, motion-resistant human presence detection with a low rate of false alarms and, consequently, a lower frequency of waking up other system components. Thanks to the always on mode, these power benefits are realized without compromising detection latency.
[0021] Unlike the PiR sensor, radar does not have a protruding lens; it functions well even when covered with a wide range of materials. This, together with a small footprint of the radar and machine learning models chips, empowers the OEM (original equipment manufacturer) to deliver a technologically cutting edge solution without compromising the aesthetics or buildability of the design.
[0022] In one aspect, a method is provided for pre-processing data for further processing by a machine learning model. The pre-processing comprises generating data from a transmitted signal and a received signal over a time period and allocating the sample data to a plurality of frequency range bins. The differences between radar data at different time steps are then calculated for each of the frequency range bins, and the calculated differences are provided to a machine learning model.
[0023] Each frequency range bin may comprise sample data for a respective subrange of the frequency range of the transmitted signal. In this way, the range bins may include the sample data arranged in accordance with the range to which the sample data relates.
[0024] In another aspect, a system is provided for pre-processing data for further processing by a machine learning model with temporal dynamics. The system is arranged to acquire radar data from a transmitted signal and a received signal over a time period. The system comprises a processor which allocates the radar data to a plurality of frequency range bins according to the radar frequency of radar data of the received signal. The processor is arranged to then calculate a difference between the radar data at two or more different time steps for each of the plurality of range bins, and provide the calculated difference to a machine learning model with temporal dynamics for further processing.
[0025] The processor may comprise a microprocessor, ASIC, FPGA, or combination of these and other circuits. The term "processor" may encompass multiple processors, each designed to perform distinct functions. For example, one processor might be dedicated to executing Fast Fourier Transforms (FFT) for signal analysis, while another processor is specialized in implementing machine learning models for advanced data interpretation. Theseprocessors can operate within the same module or be distributed across separate modules, depending on the system's architecture and design requirements. This modular approach allows for the integration of various processing tasks, enhancing the overall functionality and flexibility of the radar system.
[0026] The system may also comprise a memory, which may be implemented as a block of memory or as distributed memory elements. The system may further comprise a sensor comprising one or more transmit antennas and one or more receive antennas, the sensor generating the receive signal from the one or more receive antennas. The system may further comprise a processor implementing a machine learning model. The system including the sensor and machine learning model may be integrated in a single semiconductor chip.BRIEF DESCRIPTION OF DRAWINGS
[0027] Embodiments will now be described, by way of example only, with reference to the accompanying schematic drawings in which corresponding reference symbols indicate corresponding parts, and in which:
[0028] FIG. 1 is an example of a spiking neural network processor.
[0029] FIG. 2 is a schematic diagram of an example the method for pre-processing radar data for a machine learning model, the machine learning model being given in this case by a preferred embodiment comprising a spiking neural network, the output of which is used to control any of a variety of devices.
[0030] FIG. 3 is a flow diagram showing an example of radar data pre-processing for a machine learning model.
[0031] FIG. 4 is a schematic diagram showing an embodiment of a system comprising a radar of which the data is pre-processed for further processing by a machine learning model, the result of which is used to generate a control signal sent to a device to be controlled.
[0032] The figures are intended for illustrative purposes only, and do not serve as restriction of the scope or the protection as laid down by the claims.DESCRIPTION OF EMBODIMENTS
[0033] Hereinafter, certain embodiments will be described in further detail. It should be appreciated, however, that these embodiments may not be construed as limiting the scope of protection for the present disclosure.
[0034] This invention introduces a robust human detection solution intended as an always on wake-up system for the traditional camera-based smart doorbell. The solution comprises for example of a radar and a spiking neural network, for example a 60 GHz Continuous Wave Frequency Modulated (FMCW) radar, and a Spiking Neural Processor T1 or Pulsar created by Innatera. The radar enables acquisition of high-resolution data describing the environment in 3D. From this data, one can extract information about static and dynamic characteristics of the environment and targets within it such as target types, locations, sizes and speeds of movement of various targets. Other radar systems can also be used. The T1 or Pulsar Spiking Neural Processor is designed to be sufficient as the first and only chip a sensor talks to. The T1 or Pulsar incorporates Neural Processing Units and a lightweight RISC-V CPU to provide application developers with a heterogeneous platform for sensor data processing and action coordination. Other neural processing chips can also be used.
[0035] FIG. 1 shows an example of a spiking neural network processor that can be used in the present invention, namely the T1 or Pulsar Spiking Neural Processor.
[0036] Here, the term “always on” means that the sensor and Al work at a frequency appropriate for the application and / or delays are not perceptible to a human user. Both radar and microprocessor may be put into sleep when not actively in use to minimise power draw. However, the duration of sleep may be much shorter than that of e.g. a camera linked to the system, and the frequency with which the radar and Al wake up will be much higher, so as to frequently monitor the space, while camera and other systems would wake up e.g. only when a person is detected.
[0037] FIG. 2 is a schematic diagram of an example of the method 200 for pre-processing radar data for a machine learning model, the machine learning model being represented by a neural processor 203, preferably comprising a spiking neural network, the output of which is used to control any of a variety of devices 208.
[0038] A radar 201 emits radio waves, receives their reflections, and optionally pre- processes the returned signals to provide real-time sensor data.
[0039] The signal transmission is performed by generating chirps. A chirp is a short burst of radio energy transmitted by the radar 201. This signal sweeps across a range of frequencies over a brief time. These signals are emitted into the environment using one or more antennas.
[0040] The emitted waves travel outward and when they encounter an object (e.g., a person, wall), part of the energy is reflected back toward the radar.
[0041] The radar’s receiver antenna captures the returning echoes. Because the signals reflect off targets at different distances, they return at different time delays, allowing the radar to infer range (distance to target).
[0042] In one embodiment, a system comprising multiple spatially separated radar sensors or antennas is configured to determine the Angle of Arrival of an incoming wavefront emitted by a source. By leveraging a combined near-field and far-field model, the system is positioned such that it can resolve the angular orientation of the source relative to the sensor array. As a plane wave impinges upon at least a pair of sensors in the array, a measurable time delay arises between the signals received at each sensor which may be expressed as a shift in phase. This phase shift is indicative of the angle of incidence of the incoming wavefront. Using trigonometric relationships, the time delay can be translated into an Angle of Arrival estimate. The system applies cross-correlation techniques to estimate this delay with high precision, taking into account that the data is sampled and not continuous. Based on the derived Angle of Arrival, the system may determine either the angular position of the source relative to the array or estimate the source's absolute position, depending on the known geometry and distance assumptions.
[0043] The received analog signal may be converted into digital samples in real time using an Analog-to-Digital Converter (ADC). Each analog or digital sample contains information about the amplitude and phase of the signal. This information is stored as complex numbers (real and imaginary components).
[0044] The resulting real-time sensor data 202 is then submitted to the neural processor 203. The neural processor 203 performs e.g. pattern detection 204, pattern identification 205, signal processing 206 and / or data fusion 207 with data from other sensor streams being processed within the same processing pipeline.
[0045] The output of these neural processor is used to control any of a variety of devices 208.
[0046] The radar and processing system solution relies on detecting the signatures of human presence regardless of movement. This robust human presence detection is achieved because Spiking Neural Networks (SNNs) excel at extracting patterns from spatio-temporal data. Each neuron and weight contains time-weighted traces of the past samples, with the network being trained to combine them to efficiently turn sensor data into information and action. The exemplary T1 or Pulsar Spiking Neural Processor provides a highly efficient hardwareimplementation of the SNNs. The Machine Learning pipeline was trained and extensively tested on a large body of data capturing real world application scenarios. The dataset was carefully crafted to examine optimization of low-latency true human presence detection with a low rate of false positives. Consequently, the inventors observed that the pipeline delivers >99% accuracy on traditionally error-prone scenarios containing non-human movement.
[0047] Human presence is detected even if a person stands still. Conversely, non-human movement, such as bushes moving in the wind, may not trigger a detection.
[0048] The combination of the radar and neural processing chip, for example the T1 or Pulsar microprocessor enables realization of an ultra-low power Al customer solution. The system itself operates with a low mW power envelope. However, significant power benefits are realized at the systems level, where these additional components more than pay for themselves. The radar and processing system may remain always on, scanning the environment for human presence.
[0049] While the radar and processing system keeps the watch, more power-hungry components of the doorbell, including the camera, can be put into sleep. Upon detection of human presence, the neural processing chip, for example the T1 or Pulsar RISC-V can be used to coordinate the desired actions such as waking up other system components, alerting the homeowner and triggering camera recording.
[0050] The proposed setup significantly extends the battery life of the doorbell thanks to the low average power combined with robust, motion-resistant human presence detection with a low rate of false alarms and, consequently, a lower frequency of waking up other system components. Thanks to the always on mode, these power benefits are realized without compromising detection latency.
[0051] Unlike the PiR sensor, radar does not have a protruding lens; it functions well even when covered with a wide range of materials. This, together with a small footprint of the radar and neural processing chip, empowers the OEM to deliver a technologically cutting- edge solution without compromising the aesthetics and buildability of the design. Because of the small and convenient foodprint, the system can be built in any location.
[0052] The radar and processing system solution includes a user-adjustable embedded Al pipeline. Parameters such as detection range and sensitivity are exposed, and can be adjusted without the need for network retraining, and without the requirement for Al or embeddedsystems knowledge. Customers requiring a high level of customisation can adjust the pipeline and retrain the Al model using a Software and Embedded Development Kit.
[0053] With the growing use of sensors and Al in smart home products like doorbells, the cost of the electronic components as a proportion of the price keeps increasing. Profit margins can be improved by reusing the same components in more than one role. Such a reuse also decreases the risk and cost of component integration.
[0054] The combination of radar and neural processing chip such as the T1 or Pulsar Spiking Neural Processor is particularly suitable for such a reuse. Radar is an active sensor, with adjustable operating range and resolution. Thanks to RISC-V, T1 or Pulsar can be programmed to execute more than one application. These applications can use the same or different radar settings; switching between the settings and applications can be based on the Al results. Further, T1 or Pulsar is compatible with a wide range of other sensors, including sensor fusion. Any other neural processing chip that is reconfigurable and reusable in this way can also be used.
[0055] Accurate and robust human presence detection not only enhances the value of many existing applications but also paves the way for new ones. Some applications of the present invention include: complex power management of devices, devices such as TVs and personal computers can rapidly be put to sleep when no one is present, depth of sleep can be controlled based on human proximity to provide balance between power saving and response latency; smart buildings, person’s comfort can be maximized while bills and environmental impact minimized by adjusting temperature, airflow and lighting based on which areas are occupied; sound optimization, sound can be optimized for the area of human presence to improve the listener's audio experience and reduce noise pollution outside the identified zone; automotive in and outside the cabin, detection of human presence is required for such key applications like Rear Occupant Alert System and Child Presence Detection, it preserves health and life when used for alerting the driver about people left in a locked car, reminding passengers about seatbelts, and providing location information that can be used to adjust airbag deployment parameters, car security can be enhanced through intruder detection systems; interactive displays, screens can turn on and change in response to human proximity, touch free interface can be realized thanks to the already present radar and processor chip.
[0056] In a market where demands for smart products clash with the need for green transformation, the present invention provide a solution. The combination of a versatile radarsensor and a powerful microprocessor is tailored for delivering real-time data analysis and Al-driven actions in a low mW power envelope. This joint design empowers OEMs to join the edge Al revolution quickly and without compromising on meeting the environmental targets. The radar and processing system solution for human presence detection increases the value proposal for products such as smart doorbells, TVs, and many others.
[0057] The embodiments described hereafter comprise embodiments of a method for preprocessing radar data for a machine learning model, the output of which may be used to control a device. The embodiments of said method described hereafter may be used in the operation of a system such as described previously, including smart doorbell systems and other examples.
[0058] FIG. 3 is a flow diagram showing an example of radar data pre-processing method 300 for a machine learning model.
[0059] As a first step, radar data acquisition 301 may be performed. The radar data acquisition step involves capturing raw time-domain radar signals from one or more receiving antennas. The radar system transmits electromagnetic pulses and records the reflected signals returning from objects within the radar’s field of view. The acquired data typically consists of complex samples organized into frames, with each frame representing a sequence of radar sweeps or chirps.
[0060] The type of radar used should enable gathering of data on both static and moving targets. Typically, this would be a Frequency Modulated Continuous Wave Radar, but can apply to Ultra-Wide Band radar. The method can be applied to a radar with any number of receiving antennas. The method can be applied to a radar with any number of transmitting antennas, assuming additional computation steps applied to obtain range-based information are used. For example, appropriate signal separation or beamforming techniques can be used to resolve contributions from each transmit antenna and reconstruct range-dependent information accordingly.
[0061] Radar frequency used would typically be around 60 GHz. Use of other frequencies is possible with the proposed method, and can include e.g. 24 GHz, or Ultra-Wide Band range, which typically is between 3 and 10 GHz.
[0062] Radar data may be acquired in an analog or digital domain. Data dimension is described by the number of frames F, chirps CH, and samples S, where the F, CH and S are non-negative integers.
[0063] A chirp can be defined as a single radar transmission that sweeps across a range of frequencies over a short time, used to measure target range and velocity.
[0064] A sample can be defined as an individual data point, possibly digitized from the received radar signal, representing the reflected signal strength at a specific time during a chirp.
[0065] A frame can be defined a complete set of radar data captured over a specific observation period, typically consisting of multiple chirps across all transmit-receive antenna combinations.
[0066] Next, a calculation 302 of range-based information may take place. Radar data is transformed to a format which represents radar information per range or range bin. This is typically achieved using Fourier Transform (FT), e.g. Fast Fourier Transform. The FT parameters such as the number of points can be selected as required. Windowing (e.g. but not limited to Blackman, Hann and Hamming) and zero-padding can be applied to the data before the FT calculation. The data e.g. undergoes a (Fast) Fourier Transform ((F)FT) to extract frequency domain representations. This transforms the time-domain radar signals into range profiles by computing the frequency content corresponding to the delay of the reflected signals, which correlates with object distance. The output remains in complex form, preserving both amplitude and phase information for further processing.
[0067] Acquired radar data comprises amplitude and phase data, which may be expressed by complex numbers comprising a real and an imaginary component.
[0068] A range can be defined as the distance between the radar sensor and a target, calculated based on the time delay between signal transmission and echo reception.
[0069] A range bin can be defined as a discrete interval of distance within which radar reflections are grouped, representing the resolution element for separating targets by range.
[0070] Operating in a range domain is a benefit that the method provides for the application. It allows the end user of a device implementing the proposed method to specify range-based requirements such as the range(s) of interest and range-based weighting. User requirements can be easily met by applying the steps of the proposed method or additional modifications (such as weighting) as required to the selected range.
[0071] In some embodiments, a selection is made of one or more frequency range bins that form a range of interest to a user. This selection may be performed manually, for example via a configuration interface allowing the user or installer to specify a fixed distance range withinwhich detection is to be performed, such as a zone immediately in front of a door. In other embodiments, the selection may be made dynamically by analysing the radar data over time to identify which range bins provide the most useful information for the current detection task. Such dynamic selection may be based on statistical measures such as signal-to-noise ratio, variance, or activity level in each range bin, and may adapt to changes in the environment or in the movement patterns of objects. The selection process may also incorporate range-based weighting, wherein each frequency range bin is assigned a weight indicating its relative importance for the machine learning model’s decision-making. This enables the system to focus computational resources on the most informative portions of the radar data while ignoring or down- weighting less relevant ranges.
[0072] Next, a (complex) subtraction step 303 takes place. A temporal difference in radar data is calculated in the complex domain, i.e. by separately calculating the difference in the real and imaginary components of the radar data.
[0073] Calculating a temporal difference in the complex domain, i.e., by separately evaluating changes in the real and imaginary components of radar data, is important because it preserves both the amplitude and phase information of the signal, which may be important for e.g. accurately detecting motion, phase shifts, and other changes over time.
[0074] Phase changes over time encode information about motion (e.g., Doppler shifts from moving targets). Calculating differences in the complex domain ensures phase differences are not distorted or lost, which would happen if only magnitudes were compared. A real-valued difference (e.g., on amplitude only) would lose key information that could indicate motion direction or subtle target behaviour. Complex-domain differencing retains the signal’s vector nature, enabling high-sensitivity change detection. Computing differences post-magnitude extraction introduces non-linearities (e.g., square roots), which can mask or distort small but meaningful variations.
[0075] Calculating the difference may consist of a simple subtraction of one complex number from the other, where the two aforementioned complex numbers represent the radar data at a first time step and a last time step. The first and last time steps are the delimiting time steps for the time interval for which the difference is calculated. Time steps can correspond to frames or chirps that are consecutive or non-consecutive.
[0076] The difference can be calculated using complex subtraction. We write Real Signal(flrst) to refer to the real part of the signal for the first time step, Real Signal(last) torefer to the real part of the signal for the last time step, Imaginary Signal(flrst) to refer to the imaginary part of the signal for the first time step, and Imaginary Si gnat (last) to refer to the imaginary part of the signal for the last time step. Time steps can correspond to frames or chirps that are consecutive or non-consecutive. The difference can be calculated for more than one pair of time steps. The difference can be calculated as given below, or using another method for calculating complex subtraction: R = Real Signal(flrst)- Real Signal(last)I = Imaginary Signal(flrst)- Imaginary Signal(last)
[0077] This step contributes to the benefits realised with the method we propose. Experimental results show that calculation of the difference in the complex domain is beneficial as, among other things, it improves representation of movements, sensitivity to micro movements and vital signs. Ability to calculate the difference for more than one time step size allows a straightforward extension of the method to multi-timescale representation.
[0078] In one embodiment, the system is configured to calculate differences between radar data using multiple time step sizes, dynamically selected based on contextual factors such as the type of motion to be detected or the current state of the environment. For instance, when the system detects the presence of a subject in the scene, it may employ shorter time step sizes to enhance sensitivity to fast or fine-grained movements, such as hand gestures or tremors. Conversely, in the absence of detected movement, or when monitoring slow physiological patterns like gestures or walking, the system may apply longer time intervals between compared time steps to better capture gradual changes. The choice of time step size may also be informed by scene conditions or prior detections, such as when larger body motion is visible or when multiple moving subjects are present. This adaptive, multitimescale differencing approach enables the system to flexibly tune its sensitivity to different motion profiles and improves the robustness and richness of the extracted temporal features.
[0079] The method may thus comprise calculating, in the complex domain, one or more differences between complex radar data acquired at a first time step and a second time step, wherein the difference is determined by separately subtracting the real components and the imaginary components of the radar data at the respective time steps. The first and second time steps may be selected to have one or more different time step sizes, enabling the method to calculate differences corresponding to both short and long time intervals. The selection of time step size may be dynamically based on one or more criteria, including whether thepresence of a target has been detected, whether fast or slow movement is to be captured, or based on characteristics of the scene as determined from the radar data.
[0080] This is because it enables the system to capture and analyse temporal dynamics at different rates or granularities. First of all, different motion patterns occur at different timescales. Fast movements (like vibrations and vital signs) and slow movements (like gestures or walking) manifest differently in radar data. By computing differences over short and long time intervals, both rapid and gradual changes can be detected and distinguished. In order to distinguishing between fast and slow events, one typically looks at the chirp frequency, which would respectively be in the order of hundreds of Hz and in the order of a couple of Hz. In any case, there is not a well-defined boundary between fast and slow events.
[0081] Using multiple time step sizes allows the system to maintain sensitivity to short-term transients while also recognizing long-term trends or patterns. This improves robustness across a variety of targets and behaviours.
[0082] The main benefit can be that including this step of complex subtraction of data allows the network to generalise better; for example, the network removes static signals coming from the surroundings from the data, and leaves the signal of interest.
[0083] Note that while the pre-processing may include the step of subtraction (preferably in the complex domain), and thus while it relies on differences between radar data at different time steps, the human detection performed using the pre-processed radar data is not movement based in a conventional manner, i.e. it does not rely on macro movements from humans or using any delta (with or without a decision threshold) as a proxy for human presence.
[0084] The radar sensor can operate at an exemplary bandwidth of up to 10 GHz, for example centred around any value between 55-65 GHz. In principle any radar sensor can be used. The larger the bandwidth the better the range resolution. The number of processed frames (e.g. produced subtracted frames) per second can for example be in the order of 1-100 Hz, preferably 2-30 Hz. The number of subtracted frames per second can be equal to the inference rate of the neural processor.
[0085] Known solutions are triggered using macro movements or using a delta as a proxy for human presence. For example, when lights go off and the light is triggered by a movement sensor, one needs to go up or wave hand to turn them back on. It is also common for devices to be triggered by any type of movement / change, with no discriminationregarding what kind of movement it is. For "human presence" applications, radar-based solutions indicate "presence" when anything changes in the signal (above a certain threshold), thermal solutions are triggered by changes in temperature caused by the sun, camera solutions are triggered by change of lighting or any moving object.
[0086] For the present solution, it is not sufficient for movement to be present, and it is not required for a macro-movement to be present. Even without macro-movements, there is a lot of motion present in the internal organs, micro-motions from not being able to be truly perfectly still, and, importantly for our networks, there is also a history of the spatio-temporal data at that location. Sensor type, sensor settings and signal processing can extract and preserve these for the application pipeline to be able to benefit from their presence.
[0087] As a next step, the two-dimensional data comprising the amplitude and phase data in complex number form may then be converted 304 to the real-number domain. A typical operation applied here is calculation of the signal magnitude as follows: ^l(R2+ I2).Depending on the computational requirements and hardware support, an alternative version, e.g. omitting applying the square root, can be used. This operation yields the magnitude squared of the signal, which may reflect the magnitude of the change of the reflection.
[0088] As a next step, data can be aggregated 305 e.g. by calculating sum or average of a range bin per frame. This is an optional but beneficial step which reduces data dimensionality and improves signal to noise ratio.
[0089] As a next step, data can be normalised 306 based on static or dynamic parameters. Parameters include mean, standard deviation, min and max and interquartile ranges of the dataset or its subset, e.g. data describing a particular prediction class. Parameters can also include statistics describing relations between classes, e.g. value half way between the means of the classes, or other value at which separation between the classes is the largest. Static parameters can be calculated based on the statistics of the data gathered in advance. Dynamic parameters can be fully or partially adjusted during the operation of the pipeline. Examples include calculation of mean with a time-weighted contribution of the current and past mean samples. The further in the past samples were collected, the lower their contribution. The desirable discounting function can be applied, including linear, exponential and step function.
[0090] Normalisation parameters may be calculated and applied per range bin. Signals in further ranges tend to have smaller magnitude (as radar signals travel farther, they weaken due to free-space path loss and target reflectivity). By adjusting the scaling for each range binindependently, the radar system can bring weaker distant signals into a similar dynamic range as closer ones. This range-based treatment improves the use of the available data range, ensuring that valuable signal content is preserved and comparably weighted across all distances. It also results in a more equal contribution of data from each range to the neural network inputs in the later steps of the pipeline. Range-based weighting is therefore not a side effect of the signal loss over distance. Range-based weighting remains a possible result of later steps, including but not limited to network training, which maximises accuracy and therefore realises benefits at the application level, and user specification, which brings the benefit of application being user-customisable.
[0091] An example of the normalisation step includes division of the data by a parameter like the max for the corresponding range, then clipping it within bounds to avoid overflows if the sample comes from a distribution not captured by the assumed max. Further, this can be extended to separate treatment of positive and negative values, with max being used for positive values and (negative) minimum for negative values; the clipping bounds can similarly be defined separately for positive and negative value ranges. Normalisation is applied to improve the utilisation of the available data range available on the hardware.
[0092] As a next step, data quantisation 307 can be applied to the data in a way that results in a linear or non-linear transformation being applied to the data. Quantisation can be implemented using a lookup table. Quantisation is applied to improve the utilisation of the available data quantisation available on the hardware.
[0093] The lookup table may comprise a list of value ranges and a representative value of that value range, and a value of the data may be compared with the list of value ranges and quantised by setting it equal to the representative value of the value range wherein the value lies.
[0094] A linear transformation includes a simple mapping of the value range onto the quantisation range, and changing data representation to the quantised one with a lower resolution than the original one.
[0095] A non-linear transformation can include assigning more quantisation bins to data ranges which contain a large proportion of the sample values and / or more information, e.g. data ranges associated with class overlap. Conversely, fewer quantisation bins can be assigned to data ranges that:i. Contain a small proportion of the sample values and / or; ii. Contain mostly background data, e.g. values around the noise level.
[0096] In linear quantisation, the entire range of signal values is evenly divided into a fixed number of quantisation levels. This means that each quantisation bin represents an equal portion of the value range, regardless of how often values occur within that range. It is a straightforward technique that simplifies data by reducing resolution while preserving proportional relationships between values. This method is computationally efficient but may waste precision in areas of the data that are more informative or densely populated.
[0097] In contrast, non-linear quantisation allocates quantisation levels unevenly, based on the distribution and relevance of the data. More quantisation bins are assigned to regions that contain a high density of sample values or important features, such as those where different classes (e.g., target types or motions) may overlap, because preserving more detail in these regions supports better detection or classification performance. Conversely, fewer bins are allocated to data regions that are sparsely populated or dominated by background noise, where less information is typically lost by aggressive compression. This adaptive strategy improves efficiency and representation fidelity in the parts of the data that matter most.
[0098] In other words, the bounds of the quantisation bins may be defined as a function of data distribution, i.e. where for each unit of data range, a range with fewer samples with respect to other ranges is represented using fewer quantisation bins; conversely, a range with more samples is represented using more quantisation bins.
[0099] As a next step, data encoding 308 is an optional but beneficial step, and whether it is required depends on the type and hardware implementation of the neural network used in the next step. Encoding is necessary for networks expecting a “spiking” data representation as an input, where “spiking” means representing data in a sparse, event-based format, using both timing and value to encode the original value.
[0100] An example of encoding is thermometer encoding. When using this method, a time window for representing an input is defined. Multiple inputs can be processed in parallel, typically using time windows of the same length, and with the same start / end time. The time window is composed of time steps, where each time step can accommodate a spike event. The larger the input value, the more consecutive time steps contain spikes. This resembles filling of a scale in an analog thermometer. In the method, the encoding range corresponds to the quantisation range. The encoding step contributes to the benefits realised using theproposed method. It creates a sparse, event-based data representation therefore reducing the amount of memory and network activations required.
[0101] As a next step, the pre-processed data is transferred to a machine learning model and processed 309 using the machine learning model. The method is not network type or architecture-specific, but benefits are more fully realised when the selected network has a spatio-temporal data representation. This includes for example recurrent networks, long- short-term memory networks and spiking neural networks. Further benefits are incurred when the machine learning model has a sparse activation pattern, where integration and / or activation does not have to be calculated for all neuronal units at each forward pass. This includes spiking neural networks. A preferred implementation is realised using spiking neural networks.
[0102] Spiking Neural Networks (SNNs) are a type of neural network that more closely mimics the way biological neurons communicate. Unlike traditional neural networks, where neurons output continuous values, SNNs operate using discrete spikes or events. These spikes represent neuronal firing and can encode temporal information. SNN implement memory capabilities in various ways. SNNs can encode and process temporal patterns directly through the timing of spikes. This allows SNNs to remember the sequence and timing of events. Spike-Timing Dependent Plasticity is an exemplary learning rule used in SNNs where the timing difference between spikes of connected neurons affects the strength of their synaptic connections. This mechanism helps SNNs retain and learn temporal patterns over time. A further example that can be used is a gradient-based method with surrogate gradients to train the SNN. An example of a gradient-based method applied to SNNs is Backpropagation Through Time. Finally, some SNNs can maintain a state of activity for a certain period, which allows them to retain information for short durations even after the initial stimulus has passed.
[0103] Recurrent Neural Networks (RNNs) are designed to recognize patterns in sequences of data by utilizing internal states to process sequences. They have connections that loop back, allowing them to maintain a form of memory about previous inputs. RNN’s memory capabilities can be implemented in various ways. For example, RNNs use their internal state (or hidden state) to remember information from previous time steps. This state is updated as new data is processed, which allows the network to maintain temporal context. RNNs are capable of learning dependencies across different time steps in the sequence, making themsuitable for tasks where context from earlier in the sequence influences the current output. Finally, traditional RNNs can suffer from vanishing and exploding gradient problems, which affect how well they can retain information over long sequences. Specialized variants like LSTMs and GRUs (Gated Recurrent Units) are developed to address these issues.
[0104] Long Short-Term Memory Networks (LSTMs) are a specific type of RNN designed to overcome the limitations of traditional RNNs, particularly the vanishing gradient problem. LSTMs have a more complex architecture with specialized units that help them retain information over longer periods. LSTMs have a cell state that runs through the entire network and is modified by various gates. This cell state acts as a memory buffer that can retain information for long periods. LSTMs use three types of gates — input gate, forget gate, and output gate — to control the flow of information into, out of, and within the cell state. These gates regulate how much information is retained or forgotten at each time step. Due to their design, LSTMs can capture long-term dependencies and maintain context over extended sequences, making them effective for tasks requiring the model to remember information across many time steps.
[0105] The machine learning model can thus be configured with memory capabilities to the radar data, wherein the memory capabilities are provided by at least one of a Spiking Neural Network (SNN), a Recurrent Neural Network (RNN), or a Long Short-Term Memory (LSTM) network.
[0106] The pipeline can be trained to recognise any kind of spatio-temporal signature. It can for example include different kinds of human activity, animals, adult vs child, or objects. The machine learning model needs to be trained with labelled data. The pipeline can be used for different applications by adjusting the hyperparameters and network training.
[0107] As a next step, a decoding operation 310 can be performed on the output of the machine learning model. The decoding operation which may be required depends on the type and hardware implementation of the neural network used in the previous step. If the preferred approach was chosen and therefore a spiking network was used, the spiking output is converted into a class prediction. Methods to perform this operation include mapping output neurons IDs to class IDs and then applying max decoding or first-to-spike decoding. In max decoding, the winning class is determined based on the ID of the neuron which spiked the most in a defined time window. In first-to-spike decoding, the winning class is determined based on the ID of the neuron which spiked first in a defined time window.
[0108] Optional benefits can be realised by performing additional post-decoding operations after the decoding step. Example of operations applied to the decoded output include majority vote in a window containing decoding results from more than one time step. Another example is a threshold-based vote where a certain number of predictions of a given type has to be reached in a time window described in the previous point. This threshold can be set separately for each class. Another example is a debouncing mechanism 311 where a certain number of consecutive predictions has to be reached before the predicted class changes. This number can be set separately for each class. The count can be continuous, or can reset when a class changes.
[0109] The output of the above method, either from the decoding step or the post-decoding step described in the previous point, can also be combined with the outputs from other pipelines. These can include use of other neural network pipelines, statistical or heuristic methods.
[0110] A machine learning model configured with memory capabilities is a model designed to retain and process information over time, rather than treating each input independently. These models possess temporal dynamics, meaning they can recognize patterns that unfold across sequences of data. Examples include Recurrent Neural Networks (RNNs), Long Short- Term Memory networks (LSTMs), and architectures like Spiking Neural Networks (SNNs). These architectures allow the system to understand time-dependent features, such as rhythmic motion, periodic breathing, gestures, or changes in presence and activity patterns — capabilities essential for interpreting radar data that naturally evolves over time.
[0111] In the context of the proposed radar pre-processing method, the use of a model with memory means the system can exploit the temporal structure of the radar signal. Instead of feeding the raw radar data directly into the model, the system first calculates differences between frames for each frequency range bin, effectively highlighting changes over time. These differences emphasize movement and other time-varying phenomena while suppressing static or background components. By structuring the input in this way, the preprocessing step produces features that align well with the temporal learning capabilities of memory-based models. For example, sudden spikes in difference magnitude may indicate motion, which can trigger temporal sequences that the model leams to associate with specific activities or events.
[0112] To interface effectively with such models, particularly SNNs or other architectures sensitive to temporal information, specific encoding and decoding schemes may be necessary. For SNNs, radar differences might be encoded as spike trains, where the presence, frequency, or timing of spikes represents the magnitude or significance of a radar return. Alternatively, in LSTM-based models, the input could be formatted as sequences of real- valued vectors, where each vector corresponds to the difference across all bins for a given time step. The choice of encoding must preserve temporal coherence and emphasize the most informative aspects of motion or presence in the scene.
[0113] This approach enables capabilities, such as motion classification, presence detection, activity recognition, or even vital sign monitoring, all in a manner that is robust to noise and adaptable to different temporal resolutions. Moreover, by applying the difference step in preprocessing and structuring the input accordingly, the system ensures that the machine learning model operates on temporal features rather than raw, noisy data. This reduces computational complexity and improves interpretability and generalization. The combination of (multi-timescale) differencing and temporally aware machine learning enables more efficient, accurate radar-based sensing.
[0114] In one embodiment, the method operates in an always-on mode in which the radar continuously acquires data from its environment and the pre-processing pipeline continuously computes the differences between allocated radar data at successive time steps. The calculated differences are then provided without interruption to the machine learning model. The model is configured to process the incoming data stream in real time and to produce classification outputs, such as the detection of human presence, immediately as new data is available. This configuration enables continuous monitoring of the target region without requiring explicit triggering events, thereby reducing latency and ensuring that transient or brief occurrences are captured and classified. The always-on operation may be implemented in low-power embedded hardware so as to maintain energy efficiency while preserving responsiveness.
[0115] In another embodiment, the pre-processing pipeline described herein is designed to be agnostic to the specific use case for which the radar and machine learning system is deployed. The steps of acquiring radar data, allocating to frequency range bins, calculating temporal differences, and optionally applying further pre-processing such as magnitude computation, normalisation, and encoding, are performed identically regardless of theintended application. Preferably, the pipeline does not extract application-specific features from the radar data, this is done by the neural network to which the pre-processing pipeline is coupled. The specific use case, for example human presence detection, gesture recognition, object classification, or movement pattern analysis, is addressed solely through the configuration or training of the machine learning model. As such, to adapt the system to a different application, only the parameters or architecture of the machine learning model are modified, while the pre-processing pipeline remains unchanged. This modularity simplifies system design, reduces development time for new applications, and allows for a consistent hardware and software pre-processing implementation across multiple deployment scenarios.
[0116] The invention comprises a method for device control based on the output of a data processing pipeline, which includes a neural network, which detects human presence-related context from radar data. The neural network may be the main component which determines the human presence-related context. The method may be implemented on the sensor edge, where data travels between the sensor, the microcontroller or microprocessor through a physical connection or a method like Bluetooth, but preferably without the use of cloud. Similarly, the data processing and recognition of the human presence-related context may preferably take place locally and not in the cloud.
[0117] The “human presence-related context” includes in its simplest form recognition of presence or absence of at least one human. It can include recognition of: a. Range with regards to the sensor; b. Location within the space including presence in a particular spatial zone; c. The person’s speed of movement; d. Performed activity, including sleeping; e. Age indication, in particular adult vs child; f. The number of people, the previously listed can be applied to each person or group; g. A non-human moving target.
[0118] The method is not a simple movement detector, unlike most of the existing solutions. It: a. Distinguishes a human from a non-human moving object such as a plant moving in the wind; b. Recognises human presence even when the person remains still.
[0119] Radar is the source of data, but does not have to be the exclusive source. Devices controlled in such a way can include: a. Smart doorbell, Detection of a person can be used to turn on a camera, control camera settings (pan, tilt, focus), start data recording or notify the owner; b. Security camera, Actions as above; c. Smart tv, Streaming can be paused / resumed based on human presence; d. Interactive display, Can be turned on / off; e. Car alarmindicating passengers not using a seat belt; f. Car alarm indicating a presence of a person in a locked car, or a child in the absence of an adult.
[0120] The example implementation is human presence detection using a 60 GHz radar and T1 or Pulsar microcontroller from Innatera. Other suitable radars and microcontrollers can be used. Data processing takes place on the sensor chip or the microcontroller, for example the T1 or Pulsar. Neural network is a spiking neural network on T1 or Pulsar. The network produces inference indicating whether at least one person is present within a certain radius. Additional operations such as majority vote in a time window can be applied to the network output. When human presence is detected, a camera may be turned on. The proposed method is sensitive yet robust, does not rely on movement, and as a result reduces the number of false positives.
[0121] T1 or Pulsar Spiking Neural Processor from Innatera and a 60 GHz radar have exemplary been combined to enable reliable human presence detection in a mW power envelope. The solution is particularly well-suited for applications such as smart doorbells, where robustness, power efficiency, and low latency are critical. T1 or Pulsar is a highly programmable microprocessor. It combines a neural network engine and a nimble RISC-V processor core to form a single-chip solution for processing sensor data quickly and efficiently.
[0122] FIG. 4 is a schematic diagram showing an embodiment of a system 400 comprising a radar of which the data is pre-processed for further processing by a machine learning model, the result of which is used to generate a control signal sent to a device to be controlled.
[0123] In the embodiment illustrated in FIG. 4, the system 400 provides a process for detecting human presence in a target region using radar data and controlling a device based on the detection result.
[0124] A radar sensor 401 transmits electromagnetic signals towards a target region and receives reflected signals from objects within that region. The radar sensor may for example be configured for frequency-modulated continuous-wave (FMCW), or pulsed operation. The reflected signals are obtained and may be converted into digital radar data, for example as complex samples representing the amplitude and phase of the received waveforms over time.
[0125] The acquired radar data is processed by a signal processing stage 402 implemented in hardware and / or software. The signal processing stage may perform one or moreoperations such as fast Fourier transformation (FFT), background subtraction, conversion to magnitude, aggregation across range or Doppler bins, data normalisation, clipping, quantisation, and encoding. These operations transform the raw radar data into a form suitable for input to a neural processing stage, while reducing noise and removing irrelevant information.
[0126] The pre-processed radar data is provided to a neural processing unit 403 configured to execute a machine learning model. In one embodiment, the neural processing unit implements a spiking neural network trained to detect the presence of a human in the target region. The neural processing unit outputs a detection signal, for example a classification score or spike-rate code, indicative of whether the observed radar data corresponds to human presence.
[0127] An action determination module 404 evaluates the detection signal from the neural processing unit according to a predetermined decision criterion. If the criterion is met, such as the detection probability exceeding a threshold or the spike-rate exceeding a set value for a defined duration, the action determination module generates a control signal. The control signal indicates that a human presence has been detected.
[0128] The control signal is transmitted to a device to be controlled 405, such as a smart doorbell with an integrated camera. The device may respond by activating the camera, initiating recording, sending a notification to a user device, or performing other predetermined actions. In some embodiments, the control signal is also fed back to one or more earlier stages of the method, such as the radar sensor 401, the signal processing stage 402, and / or the neural processing unit 403, to adapt their operation dynamically. Such feedback may adjust radar operating parameters, modify processing thresholds, or alter the neural inference schedule to optimise detection performance in varying environmental conditions.
[0129] In some embodiments, the radar sensor 401, signal processing stage 402, neural processing unit 403, and action determination module 404 are all implemented within a single physical device 408. For example, these components may be integrated in or near the radar sensor 401 itself, such that the radar front-end and the associated processing circuitry are colocated. In a preferred implementation, the signal processing stage 402, neural processing unit 403, and action determination module 404 form part of a radar processing chip, optionally realised as a system-on-chip (SoC) device. This integration reduces latencybetween data acquisition and decision output, minimises interconnect complexity, and enables a compact and power-efficient design suitable for embedded applications such as smart doorbells, presence detectors, or other loT devices. In such configurations, the control signal produced by the action determination module 404 may be output directly from the radar device to the device to be controlled 405, without requiring separate external processing hardware.
[0130] In some embodiments, the radar sensor 401 and signal processing stage are implemented within a single physical device 406.
[0131] In some embodiments, the signal processing stage 402, neural processing unit 403, and action determination module 404 are all implemented within a single physical device 407.
[0132] Device as used herein may refer to any physical implementation of one or more functional components, including but not limited to a chip, die, package, module, printed circuit board assembly, system-on-chip (SoC), multi-chip package, computing unit, or any other hardware arrangement configured to perform the described functions. A device may be a standalone unit or integrated with other devices, and may include associated memory, interconnects, and supporting circuitry.
[0133] In one embodiment, a doorbell system is provided that is arranged to detect the presence of a human in a target region, such as an area in front of a building entrance. The system comprises a radar device for acquiring radar data. The radar device includes a transmitter configured to emit electromagnetic signals towards the target region, and a receiver configured to detect the electromagnetic signals reflected from objects located within that region. The transmitter and receiver may operate according to a frequency-modulated continuous-wave (FMCW) or pulsed radar scheme. A processing unit is coupled to the receiver and configured to extract radar data from the detected reflected electromagnetic signals, for example by digitising the signals and organising them into a sequence of complex-valued samples representing the received in-phase and quadrature components over time.
[0134] The doorbell system further comprises a processing system for processing the radar data by a machine learning model, preferably as described above. The machine learning model may be configured to receive pre-processed radar data as input and generate an output indicative of the likelihood that a human is present in the target region. In a preferredembodiment, the machine learning model comprises a spiking neural network that is trained to detect patterns characteristic of human motion or presence from radar returns. The doorbell system uses the output of the machine learning model to generate a control signal. When the output of the model satisfies a predetermined detection criterion, the control signal indicates that the radar device has detected a human presence in the target region. This control signal is then used to initiate further actions.
[0135] In one example, the further action comprises generating a notification. The doorbell system may include a notification system such as a wireless transmitter configured to send an alert to a user device, such as a smartphone or a home automation system. The notification may be delivered via a mobile application, push notification, text message, or an audio / visual alert generated locally at the doorbell unit. This enables a user to be informed in real time when a person is detected near the entrance, even when the user is remote from the premises.
[0136] In another example, the further action comprises activating a camera integrated into the doorbell system. In such an embodiment, the camera remains in a low-power or standby state until human presence is detected by the radar system. Upon activation, the camera may capture still images or video of the target region, optionally in combination with audio recording. This approach conserves power and processing resources while ensuring that relevant visual data is captured only when necessary.
[0137] In a preferred embodiment, the doorbell system additionally comprises an activation sensor separate from the radar device. The activation sensor may be an acoustic sensor, a motion sensor, and / or an optical sensor, and is configured to generate a detection signal in response to detecting activity in the target region. An activation processor receives the detection signal and determines whether the camera should be activated, based on whether the signal satisfies a predetermined activation condition. The activation condition may be defined to reduce false activations, for example by requiring a certain signal strength, pattern, or persistence before activation is triggered. By combining radar-based human presence detection with one or more auxiliary sensors, the system can achieve improved detection reliability while maintaining energy efficiency.
[0138] The following paragraphs lists a number of clauses, detailing embodiments of the invention.
[0139] Clause 1. A method for pre-processing radar data for further processing by a machine learning model configured with memory capabilities, the method comprising:
[0140] acquiring radar data from a received signal of a radar and preferably also from a transmitted signal of a radar, wherein the transmitted signal was transmitted by the radar into an environment of the radar, and wherein the received signal comprises the transmitted signal reflected by the environment of the radar and received by the radar, wherein the radar data comprises a series of frames over time;
[0141] allocating the radar data to a plurality of frequency range bins according to the radar frequency of radar data of the received signal;
[0142] calculating one or multiple difference between the allocated radar data at two different time steps for each of the plurality of frequency range bins, wherein the time steps correspond to the frames over time of the radar data; and
[0143] providing the calculated differences for each of the plurality of frequency range bins to the machine learning model with temporal dynamics for further processing.
[0144] Clause 2. The method of clause 1, wherein the radar data comprises amplitude and phase data which is representable in two-dimensional data format and wherein the difference is calculated separately for each of the two data dimensions of the two-dimensional data format; preferably wherein the difference is calculated in the complex domain.
[0145] Clause 3. The method of clause 1 or 2, wherein the radar data originates from a Frequency Modulated Continuous Wave Radar, or a Ultra Wide-Band radar; and / or
[0146] wherein the radar data originates from a radar operating at a frequency between 1 GHz and 100 GHz, preferably between 3 and 10 GHz or between 20 and 70 GHz, more preferably the radar operates at a frequency of 24 or 60 GHz; and / or
[0147] wherein data is acquired in an analog or a digital domain; and / or
[0148] wherein the radar data is derived from the transmitted signal and the received signal by the following steps:
[0149] mixing the received signal with a reference signal derived from the transmitted signal, thereby generating an intermediate frequency signal;
[0150] filtering the intermediate frequency signal using a low-pass filter to remove high- frequency components;
[0151] digitizing the filtered intermediate frequency signal using an analog-to-digital converter to produce a digital signal;
[0152] processing the digital signal to determine radar data, wherein the radar data preferably includes at least one of amplitude, phase, range, velocity, and / or angle data.
[0153] Clause 4. The method of any one of the preceding clauses, wherein the machine learning model has a spatio-temporal data representation, preferably recurrent networks, long-short-term memory networks and spiking neural networks; and / or a sparse activation pattern, where integration and / or activation is not required to be calculated for all neuronal units at each forward pass; and / or preferably wherein the machine learning model is a spiking neural network; and / or wherein the machine learning model is configured with memory capabilities with respect to the radar data, wherein the memory capabilities are provided by at least one of a Spiking Neural Network (SNN), a Recurrent Neural Network (RNN), or a Long Short-Term Memory (LSTM) network.
[0154] Clause 5. The method of any one of the preceding clauses, wherein allocating the radar data to a plurality of frequency range bins according to the radar frequency of radar data of the received signal comprises transforming radar data to a format which represents radar information per range or range bin, preferably wherein in allocating the radar data to a plurality of frequency range bins according to the radar frequency of radar data of the received signal comprises using an integral transformation, preferably a Fourier Transform, more preferably a Fast Fourier Transform and / or wherein preferably before allocating the radar data to a plurality of frequency range bins windowing and zero-padding is applied to the radar data.
[0155] Clause 6. The method of any one of the preceding clauses, wherein a selection is made of one or multiple frequency range bins forming a range of interest to a user, and wherein only the selection one or multiple frequency range bins are further processed by the machine learning model; and / or wherein range-based weighting is applied per frequency range bin.
[0156] Clause 7. The method of any one of the preceding clauses, wherein the difference is calculated in the complex domain or an equivalent representation using the allocated radar data, preferably obtained using the integral transformation of clause 5, and wherein the difference is calculated using complex subtraction at the two or more different time steps.
[0157] Clause 8. The method of clause any one of the preceding clauses, wherein the calculated difference for each of the plurality of frequency range bins is converted to the real domain, preferably by a calculation of the signal magnitude, or the squared signal magnitude.
[0158] Clause 9. The method of any one of the preceding clauses, the method further comprising data aggregation of the radar data, preferably by calculating sum or average of a range bin per frame.
[0159] Clause 10. The method of any one of the preceding clauses, the method further comprising a step of data normalisation comprising the application of static and / or dynamic normalisation parameters which are calculated and applied per frequency range bin, the normalisation parameters including mean, standard deviation, min and max and interquartile ranges, or mean with a time-weighted contribution of the current and past mean samples, wherein the static parameters are calculated based on the statistics of the radar data acquired, and wherein the dynamic parameters are fully or partially adjustable during the operation of the method; wherein the normalisation parameters are calculated and applied per frequency range bin, preferably normalisation is performed by division of the radar data in a frequency range bin by the maximum value of the corresponding frequency range bin and performing clipping if the sample comes from a distribution not captured by the maximum value.
[0160] Clause 11. The method of any one of the preceding clauses, the method further comprising a step of data quantisation comprising a linear and / or non-linear transformation, preferably implemented using a lookup table;
[0161] wherein the linear transformation preferably includes a mapping of the allocated radar data onto the quantisation range; and
[0162] wherein the non-linear transformation preferably includes assigning more or less quantisation bins to frequency range, wherein more quantisation bins are assigned to a frequency range if the frequency range comprises a larger proportion of the sample values and / or more information and wherein less quantisation bins are assigned to a frequency range if the frequency range comprises a smaller proportion of the sample values and / or mostly background data such as values around a noise level.
[0163] Clause 12. The method of any one of the preceding clauses, wherein an additional Constant False Alarm Rate operation is applied after the allocation or calculation of the difference step.
[0164] Clause 13. The method of any one of the preceding clauses, wherein after the step of allocation, range-based data can be used to calculate Angle of Arrival, and further used to determine a feature or class of interest.
[0165] Clause 14. A method of processing radar data by a machine learning model, comprising:
[0166] the method for pre-processing the radar data according to any one of clauses 1-13, resulting in pre-processed calculated differences;
[0167] encoding the pre-processed calculated differences into an input having a representation that can be used by the machine learning model;
[0168] processing the input using the machine learning model, thus obtaining an output from the machine learning model, wherein the memory capabilities of the machine learning model enable the machine learning model to retain and utilize temporal information from previous radar data frames to enhance analysis;
[0169] decoding the output into a classification prediction, preferably wherein the classification prediction predicts human presence based on the radar data.
[0170] Clause 15. The method of clause 14, wherein the encoding step creates as input a sparse, event-based data representation; preferably wherein an encoding scheme used during the encoding step is thermometer encoding, comprising:
[0171] defining a time window for representing an input, composed of one or multiple time steps wherein each time step can accommodate a spike event;
[0172] processing multiple frequency range bins in parallel, concurrently or sequentially, preferably using time windows of the same length, and with the same start and end times;
[0173] obtaining time steps containing spikes, wherein the larger the input value, the more consecutive time steps contain spikes.
[0174] Clause 16. The method of clauses 14 or 15, wherein the decoding step comprises converting a spiking output as the output into the class prediction, preferably by mapping output neurons IDs of the machine learning model, preferably a spiking neural network, to class IDs and
[0175] thereafter applying max decoding, wherein the winning class is determined based on the ID of the neuron which spiked the most in a defined time window, or first-to-spike decoding, the winning class is determined based on the ID of the neuron which spiked first in a defined time window.
[0176] Clause 17. The method of any one of clauses 14-16, wherein the method further comprises a post-decoding step, the post decoding step comprising: a majority vote in a window containing decoding results from more than one time step; a threshold-based votewhere a certain number of predictions of a given type has to be reached in the time window of clause 16, preferably wherein this threshold can be set separately for each class; and / or debouncing mechanism where a certain number of consecutive predictions has to be reached before the predicted class changes, preferably wherein this number of consecutive predictions is set separately for each class and / or wherein the count is continuous or is reset when a class changes.
[0177] Clause 18. The method of any one of clauses 14-17, wherein the machine learning model uses the difference to detect a human presence.
[0178] Clause 19. A system for pre-processing radar data for further processing by a machine learning model, the system comprising a processor configured to perform the method steps of any one of clauses 1-13.
[0179] Clause 20. A system for processing radar data by a machine learning model, the system comprising a processor implementing the machine learning model and configured to perform the method steps of any one of clauses 14-18; preferably wherein the processor implements the machine learning model by using dedicated analog components wherein each dedicated analog component is identified with e.g. a node or connection between the nodes of the machine learning model or preferably wherein the processor implements the machine learning model by using digital components.
[0180] Clause 21 The system of clause 20, wherein the output of the machine learning model is used to generate a control signal that is transferred to an external device, generally leading to a change of state of the device; and / or
[0181] wherein the processor is integrated with a device and the output of the machine learning model is used to generate a control signal that is transferred to the device, generally leading to a change of state of the device.
[0182] Clause 22. The system of clause 20 or 21, further comprising a radar device for acquiring radar data, comprising:
[0183] a transmitter configured to transmit the transmitted signal;
[0184] a receiver configured to detect the received signal; and
[0185] a processing unit configured to extract the radar data from the received signal and preferably also from the transmitted signal.
[0186] Clause 23. The system of clause 22, wherein the system further comprises:
[0187] an activation sensor, preferably a acoustic, motion and / or optical sensor, which generates a detection signal; and
[0188] an activation processor which is configured to receive the detection signal and determines based on the detection signal whether the radar device and the processor need to be activated, wherein the determination is made such that the radar device and the processor operate when a predetermined activation condition is detected by the activation sensor;
[0189] and wherein the radar device and the processor are configured to acquire and process radar data to obtain the classification prediction.
[0190] Clause 24. The system of clause 23, wherein the activation processor is configured to generate the activation signal based on a threshold level or specific pattern identified in the detection signal from the activation sensor.
[0191] Clause 25. The system of any one of clauses 20-24, wherein the device is a smart doorbell; a smart tv and / or computer; an interactive display; a sound system; a building management system, managing e.g. the temperature, airflow, and / or lighting in a building; and / or a vehicle such as a car.
[0192] Clause 26. A doorbell system arranged to detect a human presence in a target region, the doorbell system comprising:
[0193] a radar device for acquiring radar data, comprising:
[0194] a transmitter configured to emit electromagnetic signals towards the target region;
[0195] a receiver configured to detect reflected electromagnetic signals from the target region; and
[0196] a processing unit configured to extract radar data from the detected reflected electromagnetic signals;
[0197] a system for processing the radar data by a machine learning model, preferably the system according to any one of clauses 20-24;
[0198] wherein the doorbell system is configured to use the output of the machine learning model to generate a control signal that indicates that the radar device has detected a human presence in the target region, and wherein the control signal is used to take further action based on the human presence in the target region.
[0199] Clause 27. The doorbell system of clause 26, wherein the doorbell system further comprises a notification system, and wherein the further action comprises generating a notification upon detection of the human presence.
[0200] Clause 28. The doorbell system of clause 26 or 27, wherein the doorbell system further comprises a camera, and wherein the further action comprises activating the camera.
[0201] Clause 29. The doorbell system of clause 28, wherein the doorbell system further comprises:
[0202] an activation sensor, preferably an acoustic, motion and / or optical sensor, which generates a detection signal; and
[0203] an activation processor which is configured to receive the detection signal and determines based on the detection signal whether the camera needs to be activated, wherein the determination is made such that the camera operates when a predetermined activation condition is detected by the activation sensor.
Claims
-35-CLAIMS1. A method for pre-processing radar data for further processing by a machine learning model configured with memory capabilities, the method comprising: acquiring radar data from a received signal of a radar and preferably also from a transmitted signal of a radar, wherein the transmitted signal was transmitted by the radar into an environment of the radar, and wherein the received signal comprises the transmitted signal reflected by the environment of the radar and received by the radar, wherein the radar data comprises a series of frames over time; allocating the radar data to a plurality of frequency range bins according to the radar frequency of radar data of the received signal; calculating one or multiple difference between the allocated radar data at two different time steps for each of the plurality of frequency range bins, wherein the time steps correspond to the frames over time of the radar data; and providing the calculated differences for each of the plurality of frequency range bins to the machine learning model with temporal dynamics for further processing.
2. The method of claim 1, wherein the radar data comprises amplitude and phase data which is representable in two-dimensional data format and wherein the difference is calculated separately for each of the two data dimensions of the two-dimensional data format; preferably wherein the difference is calculated in the complex domain.
3. The method of claim 1 or 2, wherein the radar data is derived from the transmitted signal and the received signal by the following steps: mixing the received signal with a reference signal derived from the transmitted signal, thereby generating an intermediate frequency signal; filtering the intermediate frequency signal using a low-pass filter to remove high- frequency components; digitizing the filtered intermediate frequency signal using an analog-to-digital converter to produce a digital signal; processing the digital signal to determine radar data, wherein the radar data preferably includes at least one of amplitude, phase, range, velocity, and / or angle data.-36-4. The method of any one of the preceding claims, wherein the machine learning model has a spatio-temporal data representation and is configured with memory capabilities with respect to the radar data, wherein the memory capabilities are provided by at least one of a Spiking Neural Network (SNN), a Recurrent Neural Network (RNN), or a Long Short-Term Memory (LSTM) network.
5. The method of any one of the preceding claims, wherein allocating the radar data to a plurality of frequency range bins according to the radar frequency of radar data of the received signal comprises transforming radar data to a format which represents radar information per range or range bin using an integral transformation, preferably a Fourier Transform, more preferably a Fast Fourier Transform.
6. The method of any one of the preceding claims, wherein a selection is made of one or multiple frequency range bins forming a range of interest to a user, and wherein only the selected one or multiple frequency range bins are further processed by the machine learning model; and / or wherein range-based weighting is applied per frequency range bin.
7. The method of any one of the preceding claims, wherein the difference is calculated in the complex domain using the allocated radar data, preferably obtained using the integral transformation of claim 5, and wherein the difference is calculated using complex subtraction at the two or more different time steps.
8. The method of claim any one of the preceding claims, wherein the calculated difference for each of the plurality of frequency range bins is converted to the real domain, preferably by a calculation of the signal magnitude, or the squared signal magnitude.
9. The method of any one of the preceding claims, the method further comprising data aggregation of the radar data, preferably by calculating sum or average of a range bin per frame.
10. The method of any one of the preceding claims, the method further comprising a step of data normalisation comprising the application of static and / or dynamic normalisationparameters which are calculated and applied per frequency range bin, the normalisation parameters including mean, standard deviation, min and max and interquartile ranges, or mean with a time-weighted contribution of the current and past mean samples, wherein the static parameters are calculated based on the statistics of the radar data acquired, and wherein the dynamic parameters are fully or partially adjustable during the operation of the method; wherein the normalisation parameters are calculated and applied per frequency range bin, preferably normalisation is performed by division of the radar data in a frequency range bin by the maximum value of the corresponding frequency range bin and performing clipping if the sample comes from a distribution not captured by the maximum value.
11. The method of any one of the preceding claims, the method further comprising a step of data quantisation comprising a linear and / or non-linear transformation, wherein the linear transformation includes a mapping of the allocated radar data onto the quantisation range; and wherein the non-linear transformation includes assigning quantisation bins to a frequency range, wherein the number of quantisation bins assigned to the frequency range is dependent on a proportion of sample values in the frequency range, an amount of information in the frequency range and / or an amount of noise in the frequency range.
12. The method of any one of the preceding claims, wherein an additional Constant False Alarm Rate operation is applied after the allocation or calculation of the difference step.
13. The method of any one of the preceding claims, wherein after the step of allocation, range-based data is used to calculate Angle of Arrival, and further used to determine a feature or class of interest.
14. The method of any one of the preceding claims, wherein the providing of the calculated differences to the machine learning model is performed in an always-on manner such that the machine learning model continuously processes incoming radar data to generate classification outputs in real time.
15. The method of any one of the preceding claims, wherein the pre-processing pipeline is independent of a specific use case such that, for different use cases, only parameters of the machine learning model are adapted while the pre-processing steps remain unchanged.
16. A method of processing radar data by a machine learning model, comprising: the method for pre-processing the radar data according to any one of claims 1-15, resulting in pre-processed calculated differences; encoding the pre-processed calculated differences into an input having a representation that can be used by the machine learning model; processing the input using the machine learning model, thus obtaining an output from the machine learning model, wherein the memory capabilities of the machine learning model enable the machine learning model to retain and utilize temporal information from previous radar data frames to enhance analysis; decoding the output into a classification prediction, preferably wherein the classification prediction predicts human presence based on the radar data.
17. The method of claim 16, wherein the encoding step creates as input a sparse, event-based data representation; and wherein an encoding scheme used during the encoding step is thermometer encoding, comprising: defining a time window for representing an input, composed of one or multiple time steps wherein each time step can accommodate a spike event; processing multiple frequency range bins in parallel, concurrently or sequentially; obtaining time steps containing spikes, wherein the larger the input value, the more consecutive time steps contain spikes.
18. The method of claims 16 or 17, wherein the decoding step comprises converting a spiking output as the output into the class prediction, preferably by mapping output neurons IDs of the machine learning model, preferably a spiking neural network, to class IDs and thereafter applying max decoding, wherein the winning class is determined based on the ID of the neuron which spiked the most in a defined time window, or first-to-spike decoding, the-39- winning class is determined based on the ID of the neuron which spiked first in a defined time window.
19. The method of any one of claims 16-18, wherein the method further comprises a postdecoding step, the post decoding step comprising: a majority vote in a window containing decoding results from more than one time step; a threshold-based vote where a certain number of predictions of a given type has to be reached in the time window of claim 16, preferably wherein this threshold can be set separately for each class; and / or debouncing mechanism where a certain number of consecutive predictions has to be reached before the predicted class changes, preferably wherein this number of consecutive predictions is set separately for each class and / or wherein the count is continuous or is reset when a class changes.
20. The method of any one of claims 16-19, wherein the machine learning model uses the difference to detect a human presence.
21. A system for pre-processing radar data for further processing by a machine learning model, the system comprising a processor configured to perform the method steps of any one of claims 1-15.
22. A system for processing radar data by a machine learning model, the system comprising a processor implementing the machine learning model and configured to perform the method steps of any one of claims 16-20; wherein the processor implements the machine learning model by using dedicated analog components wherein each dedicated analog component is identified with a node or connection between the nodes of the machine learning model or wherein the processor implements the machine learning model by using digital components.23 The system of claim 22, wherein the output of the machine learning model is used to generate a control signal that is transferred to an external device, generally leading to a change of state of the device; and / or-40- wherein the processor is integrated with a device and the output of the machine learning model is used to generate a control signal that is transferred to the device, generally leading to a change of state of the device.
24. The system of claim 22 or 23, further comprising a radar device for acquiring radar data, comprising: a transmitter configured to transmit the transmitted signal; a receiver configured to detect the received signal; and a processing unit configured to extract the radar data from the received signal and preferably also from the transmitted signal.
25. The system of claim 24, wherein the system further comprises: an activation sensor, preferably a acoustic, motion and / or optical sensor, which generates a detection signal; and an activation processor which is configured to receive the detection signal and determines based on the detection signal whether the radar device and the processor need to be activated, wherein the determination is made such that the radar device and the processor operate when a predetermined activation condition is detected by the activation sensor; and wherein the radar device and the processor are configured to acquire and process radar data to obtain the classification prediction.
26. The system of claim 25, wherein the activation processor is configured to generate the activation signal based on a threshold level or specific pattern identified in the detection signal from the activation sensor.
27. The system of any one of claims 22-26, wherein the device is a smart doorbell; a smart tv and / or computer; an interactive display; a sound system; a building management system, managing e.g. the temperature, airflow, and / or lighting in a building; and / or a vehicle such as a car.
28. A doorbell system arranged to detect a human presence in a target region, the doorbell system comprising:-41- a radar device for acquiring radar data, comprising: a transmitter configured to emit electromagnetic signals towards the target region; a receiver configured to detect reflected electromagnetic signals from the target region; and a processing unit configured to extract radar data from the detected reflected electromagnetic signals; a system for processing the radar data by a machine learning model, preferably the system according to any one of claims 22-26; wherein the doorbell system is configured to use the output of the machine learning model to generate a control signal that indicates that the radar device has detected a human presence in the target region, and wherein the control signal is used to take further action based on the human presence in the target region.
29. The doorbell system of claim 28, wherein the doorbell system further comprises a notification system, and wherein the further action comprises generating a notification upon detection of the human presence.
30. The doorbell system of claim 28 or 29, wherein the doorbell system further comprises a camera, and wherein the further action comprises activating the camera.
31. The doorbell system of claim 30, wherein the doorbell system further comprises: an activation sensor, preferably an acoustic, motion and / or optical sensor, which generates a detection signal; and an activation processor which is configured to receive the detection signal and determines based on the detection signal whether the camera needs to be activated, wherein the determination is made such that the camera operates when a predetermined activation condition is detected by the activation sensor.
Citation Information
Patent Citations
System and method for radio-assisted sound sensing
EP4102247A1
Intelligent door lock system with reduced door bell and camera false alarms
US20160189502A1
Smart-Device-Based Radar System Performing Symmetric Doppler Interference Mitigation
US20210190902A1
Cited By
A data processing method and system for geological disaster early warning
CN122245078A