Pilot signal suppression for acoustic doppler motion tracking

By constructing a composite time domain window to suppress the pilot signal energy, the method effectively addresses the dominance issue in acoustic Doppler tracking, enabling accurate detection of slow movements in smartphones.

GB2636812APending Publication Date: 2025-07-02REVIVA SOFTWORKS LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
GB2023019856
Authority / Receiving Office
GB · GB
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-21
Publication Date
2025-07-02

AI Technical Summary

Technical Problem

Acoustic Doppler motion tracking in smartphones is hindered by the dominant energy of the pilot signal, making it difficult to detect weak Doppler shifts from nearby movements due to the proximity of the microphone to the speaker and interference from stationary surfaces.

Method used

A method involving the construction of a composite time domain window to suppress the dominant pilot signal energy by combining samples offset relative to each other, allowing for accurate Doppler shift detection through Discrete Fourier Transform analysis.

Benefits of technology

This approach enables sensitive and stable motion estimation, particularly for slow movements, by clearly observing Doppler shifts through frequency spectrum analysis, enhancing the accuracy of motion tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Apparatus for performing a pilot signal suppression technique in which the dominant energy of the pilot signal, when performing acoustic Doppler motion tracking, is mitigated. The apparatus for tracki
Need to check novelty before this filing date? Find Prior Art

Description

Summary The present invention relates to tracking motion of a target ensonified by an acoustic pilot signal. Particularly, but not exclusively, the pilot signal is an ultrasonic pure tone. The signal reflected by the target is analysed, to estimate motion of the target from a Doppler shift to the frequency of the pilot signal. The present invention relates particularly to the construction of a composite time domain window, designed to eliminate the dominant energy of the pilot signal, and enable motion estimation from Doppler shifts close to the pilot frequency. Technical Background Smartphones, and other devices that feature speakers and microphones, have the capability to detect motion near the device by analysing the reflections of an acoustic signal emitted from the device’s speaker. This technology can be used for a myriad of useful applications, such as sleep-tracking, gaming, security systems, and control of musical instruments, as explored in the Inventor’s British Patent No. GB 2609061 entitled “Motion Tracking using Pure Tones”. To date, however, acoustic Doppler motion tracking has not been widely used in smartphones despite the ubiquity of devices with sufficient hardware to support it. Doppler motion sensing devices often use radio waves that propagate at the speed of light at mega- or giga-Hertz frequencies, with optimised hardware that is designed to measure the velocity and / or distance of remote objects. Acoustic Doppler applications for common hardware like smartphones, for which the techniques disclosed herein are designed, have much more restrictive operational limits, propagating undirected, low-energy acoustic waves that travel at the speed of sound, restricted to a narrow band of frequencies and sampling intervals, with receivers (microphones) that are not well isolated from the transmitter (speaker), that are intended to detect slow movements over a distance of several metres. These operational limits combine to create the central problem that the presented techniques aim to solve; namely that the acoustic signal detected by the microphone is dominated by the energy of the pilot signal, making it difficult to observe weaker Doppler shifts from nearby movements with frequencies close to the pilot frequency. The solution presented herein is both highly effective and computationally simple. The presented techniques are designed for (but not limited to) a single frequency, continuous wave pilot signal (pure tone). Alternative acoustic motion-sensing systems have been demonstrated that use frequency-modulated continuous waves (chirping) or acoustic pulses to measure the distance of objects by measuring the time of flight of reflections. Pure tone Doppler systems, by contrast, emit longer duration acoustic waves at a single frequency to measure Doppler shifts caused by the reflections from moving objects. The reflected signal has a similar frequency to the pilot signal, offset by the Doppler shift, wherein movements towards the microphone will cause energy reflections at higher frequencies to the pilot frequency, and movements away from the microphone will cause energy reflections at lower frequencies. A core advantage of pure tone Doppler systems over pulsed or chirping methods is that the pilot signal creates a constant acoustic pressure in the environment that is less perceptible to users than a pulsing or chirping signal that causes fluctuations in acoustic pressure. Pure tone Doppler systems are thus useful in contexts where it is desirable that the user should not be disturbed by the motion tracking signal, for example for sleep tracking, or to control motion-responsive visualisations intended to relax the user. A key technical challenge to overcome in smartphone pure tone systems is that, because the microphone is typically located close to the speaker that is emitting the pilot signal and not acoustically isolated, the energy of the signal received by the microphone is dominated by the energy of the transmission, (referred to herein as the “direct path signal”) and of reflections from stationary surfaces at the same frequency, which, combined, are typically several orders of magnitude larger than the energy of Doppler-shifted reflected signals from moving surfaces. When the signal is analysed with a Discrete Fourier Transform (DFT) using a windowing function such as Hann, strong frequencies such as the pilot frequency present as lobes, with tapered energy (spectral leakage) in the frequency bins surrounding the peak extending over several energy bins (e.g. up to +1- 50 Hz). User movements commonly have speeds less than 0.2m / s, resulting in Doppler shifts less than 20Hz, which fall within the range of the bins dominated by the energy at and around the pilot frequency. Due to the energy of the direct path signal and reflections from stationary surfaces typically being orders of magnitude greater than the reflected Doppler shifted energy, Doppler shifts of such slow movements do not present as distinctive peaks in the DFT, but rather as complex interference patterns that alter the shape of the main lobe. Doppler shifts at frequencies close to the pilot frequency are thus not readily measurable by DFT analysis without employing techniques for mitigating the dominant energy of the pilot signal. British Patent No. GB2609061 discloses techniques for estimating Doppler shifts from slow movements, based on measuring the time variance of energy between consecutive DFT frames. Since the energy of the pilot tone is largely constant between consecutive analysis frames, analysing energy differentials between frames provides estimates of the Doppler energy and can be used to construct an effective system for motion tracking. The present invention relates to a new pilot signal suppression technique in which the issue of the dominant pilot signal energy is addressed differently via manipulation of the time domain signal received by the microphone. Time domain signals from offset analysis windows are combined to produce a composite window in which the pilot signal energy is suppressed or substantially eliminated, which is then processed in subsequent frequency domain analysis. The techniques require use of specific pilot signal frequencies selected for their periodicity with respect to the microphone sampling rate, in a manner to be described below. Summary of invention According to an aspect of the present invention, there is provided an apparatus for tracking motion of a target, comprising a speaker configured to transmit a periodic acoustic pilot signal, a microphone configured to receive a signal reflected by a target ensonified by the pilot signal, a storage means, and processing means configured to control the storage means to store a plurality of consecutive samples of the signal received by the microphone, according to a predetermined sampling rate, determine, for each of a plurality of consecutive first samples received by the microphone at respective first sampling times, a linear combination of the first sample and a respective second sample, received by the microphone at a respective second sampling time, wherein the difference between the first sampling time and the respective second sampling time is an integer multiple, m, of the half-period of the pilot signal, control the storage means to store a composite window containing the linear combination determined for each first sample, estimate motion of the target based on analysis of a Discrete Fourier Transformation, DFT, of a composite window containing the linear combinations, wherein the DFT is represented by a plurality of energy values in each of a respective plurality of frequency bins, the apparatus further comprising output means configured to output a representation of the estimated motion. In embodiments, the frequency of the pilot signal is an integer multiple of the predetermined sampling rate divided by 2", for integer n >2. The pilot signal may have an ultrasonic frequency. In embodiments, the linear combination is an addition of the first sample to the respective second sample where m is an odd integer, and a difference between the first sample and the respective second sample where m is an even integer. The linear combination may be a weighted combination in which a first weight is applied to the first sample and a second weight is applied to the respective second sample. The second samples may be non-consecutive. In embodiments, the plurality of consecutive first samples are received by the microphone during a first sampling window and each respective second sample is received by the microphone during a second sampling window, wherein a portion of the first sampling window may overlap in time with a portion of the second sampling window. The processing means may be arranged to control a user interface in accordance with one or more gestures determined from the estimated motion. The output means may be configured to output an alarm or notification according to the estimated motion. According to a second aspect of the present invention, there is provided a musical instrument apparatus comprising the apparatus according to the first aspect, wherein the output means is configured to output a sound according to one or more gestures determined from the estimated motion. According to a third aspect of the present invention, there is provided a gaming apparatus comprising an apparatus according to the first aspect, wherein the processing means is configured to control a game character according to one or more gestures determined from the estimated motion. According to a fourth aspect of the present invention, there is provided a sleep-tracking apparatus comprising the apparatus of the first aspect, and means for performing sleep analysis for a target based on the estimated motion of the target. The output means may be configured to output an alarm within a predetermined time of the means for performing sleep analysis determining arousal or light sleep of the target. In embodiments, the output means is configured to display a visual representation of the estimated motion. According to a fifth aspect of the present invention, there is provided a method for tracking motion of a target, comprising: transmitting a periodic acoustic pilot signal, receiving a signal reflected by a target ensonified by the pilot signal, storing a plurality of consecutive samples of the received signal, according to a predetermined sampling rate, determine, for each of a plurality of consecutive first samples received at respective first sampling times, a linear combination of the first sample and a respective second sample, received at a respective second sampling time, wherein the difference between the first sampling time and the respective second sampling time is an integer multiple, m, of the half-period of the pilot signal, storing a composite window containing the linear combination determined for each first sample, estimating motion of the target based on analysis of a Discrete Fourier Transformation, DFT, of a composite window containing the linear combinations, wherein the DFT is represented by a plurality of energy values in each of a respective plurality of frequency bins, outputting a representation of the estimated motion. According to a sixth aspect of the present invention, there is provided a computer program which, when executed by one or more processors, is arranged to cause the method of the fifth aspect to be performed. Embodiments of the present invention represent a motion estimation technique which is highly sensitive, particularly to slow motions, and stable over time. This arises from the ability to observe clearly a Doppler shift to a pilot signal frequency, caused by motion of a target, through performing frequency spectrum analysis on a pilot-suppressed set of signal samples. Detailed                                                                  description Embodiments of the present invention will now be described by way of example only, with reference to the accompanying drawings, in which: Figure 1 shows an example of a device which may be used for tracking motion of a target according to an embodiment of the present invention; Figure 2 illustrates a schematic illustration of apparatus for estimating motion a target, according to embodiments of the present invention; Figure 3 is a flow chart illustrating a motion estimation method according to embodiments of the present invention; and Figure 4 is a comparison between a DFT obtained using the method of Figure 3, and a DFT of a single sample window of a received signal. Principle of operation The relationship between motion of a target ensonified by a signal, and the Doppler shift which is present in the frequency of the signal reflected by the target, is well known, and extensive description of this principle is not repeated here, in the interests of conciseness. Embodiments of the present invention use a device equipped with a speaker and a microphone, such as (but not limited to) a smartphone to transmit a periodic acoustic signal. The signal is referred to herein as a pilot signal. In embodiments, the pilot signal has a fixed frequency, and is referred to as a pilot tone. The pilot signal may be amplitude-modulated in some embodiments, by a switching signal, and may comprise more than one frequency component. A frequency component of the pilot signal provides the reference from which Doppler shifts are measured. Figure 1 illustrates an example of such a device 1, according to embodiments of the present invention, to be used for tracking motion of a target 2, comprising a speaker 3 and a microphone 4. In the illustrated embodiment, the device 1 is a smartphone, and the target 2 is a user’s hand making gestures to control the smartphone. The device comprises processing means (not shown) configured to control the speaker 3 and the microphone 4, and to execute computer-readable instructions to perform the method set out in the flow chart in Figure 3, described in detail below, in order to track motion of the target 2. The present embodiment is described in relation to use of a sinusoidal signal 5 of a single fixed frequency, referred to herein as an acoustic pilot tone. The microphone 4 receives a signal reflected by the target 2, in addition to signals reflected by other surfaces which are not necessarily of interest to a particular motion-tracking application (such as walls, furniture and other objects). Further, the microphone 4 receives a signal which has not been reflected by the target 2 or any surface, and which has reached the microphone 4 directly from the speaker 3. In a device such as a smartphone 1, the relative proximity of the microphone 4 and the speaker 3 is such that this ‘direct path’ signal dominates the energy of the overall signal which is received. The microphone 4 is controlled by a processor, such that the received signal is sampled at a particular sampling rate. Typical sampling rates commonly used in the audio-processing field include 44.1 kHz, 48 kHz and 96 kHz, and these frequencies are particularly suitable for use in embodiments of the present invention, since the on-board hardware within a device such as a smartphone 1 will usually already have a native configuration for use with the same sampling rates. Embodiments of the present invention are, however, not restricted to use of any particular sampling rate. As above, the received signal is a combination of the direct path pilot tone, and a reflected signal, containing reflections from stationary surfaces and the moving target. The components of the received signal interfere with each other, and the interference varies with time, as the direct path pilot tone and reflection from stationary surfaces, and the reflected signal, have different frequencies from each other due to the Doppler shift associated with the reflected signal. Assuming the transmission of a sinusoidal pilot tone, and the reflection of the sinusoidal signal and a Doppler-shifted sinusoidal signal, the interference is fully constructive at instants in time at which the maximum amplitudes of the sinusoidal signals interfere, and the energy of the combined signal is a maximum at such instants. The interference is fully destructive where the sinusoidal signals are out of phase with each other by 180o, or k, when expressed in radians. Pilot signal suppression, a principle which underlies embodiments of the present invention, is based on eliminating, or substantially reducing, redundant information associated with the received signals. The majority of such redundant information derives from the energy associated with the direct path pilot signal and reflections from stationary surfaces. Pilot signal suppression is achieved by acquisition, by the microphone, of two windows of samples of the received signal which are offset relative to each other such that the pilot signals of each sample window are in phase or anti-phase, and identification of the difference between the samples in these windows. Two windows having a short offset between them are likely to have similar, but not identical, Doppler-shifted components, that will not have the same degree of phase and amplitude alignment as the pilot signal. As such, if the differences between samples in two windows are determined, the redundant energy can be substantially eliminated, such that it is possible to expose the Doppler-shifted frequency component, which then presents as a distinct lobe in the DFT. This significantly improves the accuracy of the motion estimation, as will be described below. Specific pilot signal frequencies are used in order to ensure that the interval between combined samples is an integer multiple of a half period of the pilot signal, so that the samples can be combined in the time-domain in a manner which will enable the pilot signal to be suppressed or substantially eliminated. Embodiments of the present invention also employ particular optimisations in the implementation of the approach outlined above, which maximises performance for a particular application. Hardware Figure 2 is a schematic illustration of an apparatus 100 for estimating motion of a target, according to embodiments of the present invention. The apparatus 100 comprises a processor 400, which controls the operation of a microphone 200 and a speaker 300, and which comprises a motion estimation module 600 for handling operations associated with estimating motion of a target. The motion estimation result is output by an output module 700. The processor executes computer-executable instructions which are stored in storage module 800. The processor 400 controls the microphone 200 to sample a received signal and to store the samples in the storage module 800. In embodiments, the storage module 800 is implemented by a plurality of discrete storage modules, with volatile memory, such as a buffer used to store sampled signals, and non-volatile memory, such as a read-only memory, storing computer-executable instructions. The processor 400 performs one or more preliminary operations associated with preprocessing information for use in motion estimation. Specifically, in embodiments of the invention, such pre-processing involves combining signal samples from two sampling windows, offset in time, for provision to a DFT module 500 for performing a Discrete Fourier transform (DFT) on sampled data. In alternative embodiments, a Fast Fourier transform (FFT) is performed. The implementation of a DFT or FFT is in accordance with a manner known in the art. Motion estimation is performed by the motion estimation module 600 by interpreting results of the DFT module 500, again in a manner known the art. Examples of the operation of DFT module 500 and motion estimation module are to be found in British Patent No. GB2609061. Where the apparatus 100 is a device such as smartphone, the components of Figure 2 may correspond to on-board components of the smartphone. For example, the output module 700 may correspond to display, audio output or communication signal output components of the smartphone, while the microphone 200 and speaker 300 are those already used by the smartphone for making and receiving calls. In yet further embodiments, the microphone and speaker may be off-the-shelf components coupled to a personal computer. A user interface module 900 is provided to enable interaction with the apparatus by a user, and may implement a GUI or button-based interface for provision of controls or input of configuration settings. Although the DFT module 500 and motion estimation module 600 are shown as separate components in Figure 2 they may, in alternative embodiments, be considered as a single component represented by processor 400. In yet further embodiments, such a single component may be separate from, or contained within a central processing unit of the device embodying Figure 2. The specific configuration to be used will depend on the application for the motion estimation technique, and examples of such applications are described below. It will also be appreciated that the components of Figure 2 may be embodied in hardware, software, or a combination of both. For example, the processor 400, DFT module 500 and motion estimation module 600 may correspond to functional modules, such as sections of software code, executed by one or more processors in the mobile phone. It will be appreciated that other mobile devices, such as tablets, or personal computers may also support the apparatus shown in Figure 2. The speaker 300 of the apparatus 100 transmits the pilot signal, and the microphone 200 of the apparatus 100 receives the signal reflected from the user. An application running on the apparatus 100 and executed by the processor controls the operation of the microphone 200, speaker 300, DFT module 500 and motion estimation module 600 periodically. In embodiments, the operations are performed through the night to enable estimation of a user’s sleep patterns by outputting motion estimation to an analysis engine, or for output to a dashboard for analysis by a user, or for triggering a smart alarm to wake a user when determined to be in light sleep based on detected motion within a predetermined time range or period. The inaudible transmission of the pilot signal prevents disturbance of the user while sleeping. In alternative embodiments, the apparatus 100 represents an emulation of a musical instrument, in which a particular musical sound is produced by the speaker 300 in dependence upon a particular motion of the user that corresponds to a gesture used when playing a musical instrument. In alternative embodiments, the apparatus 100 is implemented as a gaming device. Here, a particular gesture of a user may enable contactless interaction with the game, via provision of a particular command to a game character. In alternative embodiments, the apparatus 100 is implemented as a user interface, in which estimated gestures by a user’s hand, for example, control one or more aspects of a user interface object, such as a cursor, scroll bar, or game character. Such embodiments are shown in Figure 1. In alternative embodiments, the apparatus 100 is implemented as a security apparatus configured to trigger an alarm in response to detection of motion. In each embodiment, the nature of the output provided by output module 700 is dependent on the specific context in which the embodiment operates. In some embodiments, the output module represents a display unit which illustrates detected motion visually so that a user can take particular action. In other embodiments, the output module represents an alarm, whether sound and / or light, or some other notification such as an email or message, or operating system notification. In yet further embodiments, the output module represents the packaging of data representing motion detection, whether instantaneous, or whether historical, measured over a period of time, which can be synchronised with an external device such as a personal computer, tablet, mobile phone or server, including cloud-based devices, in order for motion data to be uploaded and analysed further by such an external device. Signal sampling In embodiments of the present invention, a sinusoidal pilot tone is transmitted by the speaker 300, having a time-varying amplitude x at time t, and frequency fo, such thatx(t) = sm(2zfof). A proportion of this pilot tone is reflected to the microphone 200 by a target. In the simplest case, excluding, for ease of description, the influences of reflections from walls and other stationary surfaces, the reflected signal has a time-varying amplitude y at time t, and Doppler-shifted frequency fo (where the Doppler shift is fo - fo), such that y(t) = sin(2^fo0. At the microphone 200, a signal, s(t) is received which represents the interference between the direct-path x(t) and y(Q. The processor 400 controls the microphone 200 to digitally sample the amplitude of s(Q at a predetermined sampling rate, fmic such as 48 kHZ, as set out above. The result is a timesequence of digitised samples, over the duration of a sampling period, / sample, represented by a vector s of samples. Each column of the sample vector is associated with a sampling index z, such thats = {si, S2,.....sz,..., s}, fori <z< / samples, where / = fmic* Sample. After sampling, the following signal processing operations are performed by processor 400, in embodiments of the present invention. In embodiments of the present invention, samples associated with two sampling windows are weighted by applying first and second window functions to samples of s(f), such as a Hann function, or other weighting function known in the art. The first sampling window is such that it contains the sample vector g = {Sj, Sj+i, Sj+2, ..., Sj+k}, namely a plurality (k+1) of consecutive samples of s(f), beginning with sample index j, where y <( / -k-1) and concluding with sample index k, where k <i. The second sampling window is such that it contains sample vector g = {s«, s^, sz, ..., s^}, for a<(i- k-1) and $<i. The number of samples in the second sample vector g is the same as the number of samples in the first sample vector g, such that a sequence of pairs of corresponding samples, having the same index in their respective sample vectors, can be derived, namely [sj, sa], [Sj+i, sp], [Sj+2, sy], ...., [Sj+k, s^]. However, the sequence {s„, s# sz, ..., s^ need not contain consecutive samples of s(f), such that sa+ / # s^, for example. What is required is that the spacing between two samples in a given pair, follows the mathematical relationship set out below. Having applied the sampling windows, a linear combination of g and g is performed, of the form Pg + Qg, for scalars P and Q. The output of the linear combination is stored in a vector, c = {ci, C2, cz, ..., Cj+k} referred to herein as a ‘composite window’, to which spectral analysis is applied, as described in more detail below. The combination is performed in a pairwise manner, such that Ci = PSj + Qsa, C2 = Psj+i + Qsp, and so on. The purpose of the linear combination is to eliminate redundant common energy between the sample vectors g and g This is achieved by destructively combining the sample vectors, such that either a subtraction if performed, where the phase, with respect to s(0, of the combined samples of g and g is aligned, or an addition is performed, if the combined samples are of opposite phase. Considering the period, To, of the pilot tone where To = 1 / fo, a phase of a particular sample with index z can be considered to be (2k*z) / (WTo), when expressed in radians. One phase cycle of samples therefore corresponds to samples which are spaced apart by a time period of To, which corresponds to a sample index spacing of (fmic*To). Therefore, for sample pair [s„ sj, for example, if the sampling interval between s and s, represents a time interval which is an integer multiple of T, samples having indices yand a will be in phase with each other with respect to the pilot signal frequency. To combine g and g destructively, a subtraction of Pg - Qg_is performed. In cases in which both g and g contain consecutive samples of s(t), the choice of reference sample, from which the other sample is subtracted, is not of significance. Conversely, if the sampling interval between samples s, and s, represents a time interval which is an odd-numbered integer multiple of the ratio L / 2, sample indices j and a will be out of phase with each other, with respect to the pilot signal frequency. In such cases, to combine g and g destructively, an addition of Pjo + Qg is performed, with addition of out-of-phase signals being considered analogous to subtraction of in-phase signals. As such, the two linear operations set out above can be generalised as Pg + Qg, with addition performed if the time interval between Sj and s« represents an integer multiple m of half of the period of the pilot tone, To, for odd m, and with subtraction performed if m is even. The definitions provided above apply to the offset between sj and s«. In embodiments of the present invention, it is required that there is an offset of m half periods of the pilot tone signal To, between the sample pairs which are combined to produce each term in the composite window. This therefore determines a set of pilot tone frequencies which can be used for a given sample pair configuration. It is not necessary for m to have the same value for each term in the composite window, as long as the subtraction or addition rule above is followed. Therefore, if the samples in the second vector g are consecutive, m will be the same for each term of c, but if the samples of the second vector g are non-consecutive, m will increase as the spacing between corresponding samples in vectors g and g increases with the sampling index. Composite window c contains an expression of the time-difference of the combined pilot tone and the Doppler shifted reflected signals, in which the energy of the pilot signal has been substantially eliminated via the destructive combination described above. The energy of the Doppler signal is also affected by the combination of the samples, but in practice is not fully destroyed due to amplitude differences and imperfect phase alignment of g and g. In some embodiments, motion of a target towards or away from the microphone 200 will cause a change in the size of the reflecting surface of the target, as perceived by the microphone 200, such that the energy of a reflected signal increases as the target moves closer to the microphone 200 and presents a larger surface. Spectral analysis is performed on composite window c by DFT module 500 to identify its constituent frequency components. Such analysis is performed using a Fast Fourier Transform or a Discrete Fourier Transform, in a manner known in the art, for example as set out in British Patent GB 2609061. In embodiments, DFT module 500 may be referred to as a Fast Fourier Transform (FFT) module 500. The majority of the energy content of the frequency spectrum derived from such analysis is representative of the motion of the target. Background noise and surface reflections will contribute to energy being seen in each of a plurality of frequency bins of the spectrum, but the spectrum will demonstrate a peak energy level in a particular frequency bin. Based on the shift from that particular frequency bin, relative to the frequency of the pilot tone, the speed and direction of motion of the target can be estimated by the motion estimation module 600, for provision to the output module 700. The resolution of motion estimation is therefore dependent on the width of a frequency bin, which is dependent on the size of the composite window to which spectral analysis is applied. Typically having more data samples will lead to a more precise frequency estimation, but collecting more data samples leads to higher latency, and the skilled person will be able to adopt a configuration that takes into account the requirements associated with a particular application. Frequency selection It is desirable, in applications such as sleep tracking, that the motion tracking system is inaudible and not disturbing to the user. Single, pure tone systems have advantages over chirping or multiple tone systems in this respect, as chirps and tone combinations can create fluctuations in acoustic pressure and beat patterns that are perceptible and disturbing for users. To minimize perceptibility in such embodiments of the present invention, it is optimal to use a single high frequency tone for the acoustic pilot signal that is above the hearing range of the user. Most adults are unable to hear single tone frequencies above 16 kHz, although children and sensitive adults can detect frequencies up to 20 kHz. Many current devices are able to produce and sample audio at a theoretical maximum frequency of 48 kHz, enabling frequencies up to a maximum frequency of 24 kHz (half the sampling rate) to be analyzed using FFT. Whilst it is desirable to employ a tone frequency above 20 kHz for inaudibility, the selection of frequency is often constrained by the hardware’s ability to produce a stable tone at higher frequencies, with some devices struggling to produce stable tones above 18 kHz. It is hoped that future hardware improvements will enable frequencies above 20 kHz to be employed more frequently. As set out above, the set of sample frequencies to be used is determined by the offset between the sampling windows, in order to ensure the correct phase alignment of the samples used for pilot tone suppression. In the embodiments of the present invention, the pilot tone has a frequency that can be expressed, with respect to the sampling rate, as a fraction of two integers, in which the denominator has an upper limit of 8192. For example, with respect to a sample rate of 48,000 Hz, a pilot tone of frequency 18,000 Hz can be expressed as the integer ratio %, such that three complete phase cycles occur in every eight samples, whilst three half phase cycles occur in every four samples. Therefore, linear combinations of vectors £ and g are possible in embodiments of the present invention, in which g and g, are offset from each other by an integer multiple of four sample indices, because this will ensure that there is always a whole number of half phase cycles represented in g and g. Whilst a 18,000 Hz pilot frequency, in the context of an 48,000 Hz sample rate, provides a large set of viable offsets, other ultrasonic frequencies provide smaller sets of viable offsets, whilst frequencies whose periodicity is not aligned with the sample rate may have no viable offsets at all. Although inaudible to most adults, 18,000 Hz tones are often audible to children, and so it is desirable to identify higher frequencies for general use. Additional suitable frequencies can be identified by seeking integer multiples of the sample rate divided by a number which is 2", for integer n >2, which corresponds to the number of samples with which an integer number of periods of the acoustic pilot tone is aligned. For example, moving from an eight-sample period (23) to the next power of two in the geometric sequence, 24, a pilot tone frequency of 21,000 Hz completes seven half-phase cycles in eight samples. 21,000 Hz would be an ideal tone frequency for many applications due to its inaudibility and flexibility for tone subtraction. However, many current generation smartphones can struggle to produce a stable tone at that frequency. Moving further through geometric sequence, a pilot tone frequency of 19,500 Hz completes 13 half-periods every 16 samples, and 18,750 kHz completes 25 half-periods every 32 samples. Both are pilot tone frequencies within the output capabilities of more devices than 21,000 Hz. Selecting frequencies with large sets of viable offsets provides greater system flexibility, although in practice optimal performance is achieved by using offsets between 256-2048 samples in order to preserve Doppler components, without introducing excessive latency. Seeking frequencies whose ratio, with respect to the sampling rate, has a denominator which is 2", yields frequencies with superior properties in the DFT analysis phase, as the pilot frequency will also correspond exactly to the frequency of a DFT bin computed using the efficient radix-2 FFT algorithm. When analysed with FFT using a window function such as Hann, the pilot signal produces a symmetrical lobe centered on the pilot frequency. A frequency bin distribution which is symmetrical about the frequency of the pilot tone thus enables the frequency spectra to be ‘anchored’ such that in the absence of motion of the target, the energy peak and centroid will be at the pilot frequency, and the motion estimation process will interpret this as zero motion. In contrast, a pilot frequency that does not match a DFT bin frequency typically results in a non-symmetrical energy distribution, with peak and centroidal frequencies that differ from the pilot frequency, which can make it more difficult to determine slow motions and identify the neutral zero-motion state. It is beneficial, in some embodiments of the present invention, for scalars P and Q, in the linear combination, to be different from each other by a small amount. For example, P may be 1 and Q may be 0.9. By adopting this approach, the pilot tone is not fully eliminated, but is suppressed so that it does not dominate the frequency spectra. However, preserving a percentage of pilot tone energy provides a reference value such that only energy shifts that are large enough in comparison to the preserved pilot tone energy will lead to a shift in peak energy away from the frequency of the pilot tone. This avoids situations where small energy fluctuations from external noises that are not Doppler shifts, or that are distant reflections, would otherwise cause a shift in peak energy away from the pilot frequency. The specific choice of values for P and Q depends on a particular application, taking into account the desired level of sensitivity. An additional benefit of an asymmetric weighting of sample vectors, based on P and Q being non-equal, is that it is possible to give more weight to the samples which are more recent in time, and thus provide an estimation of target motion which is more current. The embodiments described above are based on the combination of two sample vectors. It is possible, in further embodiments of the present invention, to combine more than two sample vectors. For example, a composite window can be generated from the most recent sampling window, and a weighted combination of two previous sampling windows. The weighted combination can be based on applying a stronger weighting to the more recent of the two previous sampling windows, and in this manner, recent movements persist in the estimation in diminishing magnitude. For example, a composite window can be generated using 50% weighting of the most recent sampling window, and 50% weighting of a combination of previous sampling windows, weighted in the ration of 2:1 in favour of the more recent of the previous sampling windows. Of course, it will be appreciated that different weighting ratios may be used, including equal weightings, and more than three sample windows may be combined. Typical sizes of the composite window c, for use with the DFT are 2048, 4096 and 8192, with longer lengths providing higher resolution frequency estimates. Motion-estimation method Figure 3 is a flow chart illustrating a motion estimation method according to embodiments of the present invention. At step S10, a pilot signal is transmitted towards the target from the speaker 300. At step S20 a sound signal, is received by the microphone 200 and digitally sampled at a sampling frequency. The microphone 200 is controlled by a processor 400 of the apparatus 100, which causes the microphone 200 to record for a predetermined time period, such that a particular number of samples of the received signal is collected based on the sampling frequency, the samples arranged in a plurality of sampling windows. At least a first sampling vector p and as second sampling vector g are stored in the storage means 800. At step S30, a linear combination of the samples in the two sample vectors, is performed, as described above, based on the expression Pp + Qq, for weighting coefficients P and Q. The linear combination is stored as a composite window c in storage means 800. At step S40 a DFT is performed by DFT module 500, on the composite window c The frequency spectrum is provided to the motion estimation module 600. In step S50, motion estimation is performed by identifying the frequency of the peak energy of the frequency spectrum, and correlating this with an estimated size and direction of motion associated with a Doppler shift corresponding to the frequency of the peak energy. The motion estimation result is output in one or more forms, as described above. For example, in some embodiments, the output is an acoustic alarm, alert or notification. In other embodiments, the output is visual, or a control signal for a user interface on a mobile device. Example of Performance Figure 4 illustrates a comparison of the energy of a frequency spectrum 6 obtained by DFT module 500 when performing the method of Figure 3, and a frequency spectrum 7 when performing spectral analysis of a single sample of a reflected signal. The frequency spectra 6, 7 are obtained by DFT module 500 in the context of estimation of motion of a target towards the microphone 200, and each sampling window contains 4096 samples. For the two-window time-differenced frequency spectrum 6, there is a separation of 2048 samples such that 50% of the two sampling windows overlap. The pilot signal which is transmitted by speaker 300 is an acoustic pure tone of frequency 18,000 Hz. In the reference example of performing a DFT on a single set of samples, the frequency spectrum 7 has a centroidal magnitude at 17,999 Hz. This is very close to the pure tone frequency, and is not illustrative of motion of the target. As such, it can be seen that the energy of the pure tone frequency dominates the expected Doppler shift caused by motion of the target, to the extent that the Doppler shift cannot be observed clearly. Additionally, in this example, the centroidal magnitude is actually below the tone frequency, rather than above, which is a consequence of complex interference patterns between the Doppler-shifted signal and the pilot tone, such that the direction of the frequency shift is not what would be expected in the case of motion towards the microphone. In contrast, when the method of Figure 3 is performed, it can be seen that the centroidal magnitude of frequency spectrum 6 has shifted away from the tone frequency (as shown by the arrow in Figure 4), at 18,011 Hz. The effect of pilot tone suppression is clearly seen in Figure 4, as the DFT curve is not dominated in magnitude by the 18,000 Hz tone. The effect on the frequency spectrum of motion of the target can be clearly observed. It will be appreciated that a variety of implementations will fall within the scope of the claims, the specific implementation depending on at least the desired nature of the motion detection result, the desired motion sensitivity, the motion detection environment, the anticipated motion of the target, and the available hardware. Compatible functions of the embodiments described above may therefore be combined as required in order to form new embodiments, as desired for a particular application. In particular, the frequency or frequencies of the pilot signal, the sampling rate of the microphone, the length of the sample windows and sample offsets, and the weighting of the combination of samples may be controlled. Generally, the longer the sample windows, and the greater the offset between the sample windows, the longer will be the time taken to acquire the required samples, leading to greater system latency between the target motion and the estimated motion, and lower noise levels. Shorter sample windows and smaller offsets may capture dynamic movements more effectively and with lower latency, but with greater noise levels. A combination of sample window lengths and offsets can be selected based on these trade-offs.

Claims

1. An apparatus for tracking motion of a target, comprising:a speaker configured to transmit a periodic acoustic pilot signal;5 a microphone configured to receive a signal reflected by a target ensonified by the pilotsignal;a storage means; andprocessing means configured to:control the storage means to store a plurality of consecutive samples of the10 signal received by the microphone, according to a predetermined sampling rate;determine, for each of a plurality of consecutive first samples received by the microphone, a linear combination of the first sample and a respective second sample, received by the microphone at an integer multiple, m, of the half-period of the pilot signal from the microphone receiving the first sample;15 control the storage means to store a composite window containing the linearcombination determined for each first sample;estimate motion of the target based on analysis of a Discrete Fourier Transformation, DFT, of a composite window containing the linear combinations, wherein the DFT is represented by a plurality of energy values in each of a respective >0 plurality of frequency bins,the apparatus further comprising output means configured to output a representation of the estimated motion;wherein the linear combination is an addition of the first sample to the respective second sample where m is an odd integer, and a difference between the first sample and the 25 respective second sample where m is an even integer.

2. An apparatus according to claim 1, wherein the frequency of the pilot signal is an integer multiple of the predetermined sampling rate divided by 2”, for integer n >2.30 3. An apparatus according to claim 2, wherein the pilot signal has an ultrasonic frequency.

4. An apparatus according to any one of the preceding claims, wherein the linearcombination is a weighted combination in which a first weight is applied to the first sample and a second weight is applied to the respective second sample.

355. An apparatus according to any one of the preceding claims, wherein the second samples are not consecutive.

6. An apparatus according to any one of the preceding claims, wherein the plurality of 5 consecutive first samples are received by the microphone during a first sampling window and each respective second sample is received by the microphone during a second sampling window, wherein a portion of the first sampling window overlaps in time with a portion of the second sampling window.10 7. A musical instrument apparatus comprising the apparatus according to any one of thepreceding claims, wherein the output means is configured to output a sound according to one or more gestures determined from the estimated motion.

8. A gaming apparatus comprising an apparatus according to any one of claims 1 to 6, 15 wherein the processing means is configured to control a game character according to one or more gestures determined from the estimated motion.

9. An apparatus according to any one of claims 1 to 6, wherein the processing means is arranged to control a user interface in accordance with one or more gestures determined from >0 the estimated motion.

10. A sleep-tracking apparatus comprising the apparatus of any one of claims 1 to 6, comprising means for performing sleep analysis for a target based on the estimated motion of the target.2511. A sleep-tracking apparatus according to claim 10, wherein the output means is configured to output an alarm within a predetermined time of the means for performing sleep analysis determining arousal or light sleep of the target.30 12. An apparatus according to any one of claim 1 to 6, wherein the output means isconfigured to output an alarm or notification according to the estimated motion.

13. An apparatus according to any one of the preceding claims, wherein the output meansis configured to display a visual representation of the estimated motion.

14. A method for tracking motion of a target, comprising:transmitting a periodic acoustic pilot signal;receiving a signal reflected by a target ensonified by the pilot signal;storing a plurality of consecutive samples of the received signal, according to a predetermined sampling rate;determine, for each of a plurality of consecutive first samples, a linear combination of 5 the first sample and a respective second sample, received at an integer multiple, m, of the half-period of the pilot signal from the microphone receiving the first sample;storing a composite window containing the linear combination determined for each first sample;estimating motion of the target based on analysis of a Discrete Fourier Transformation, 10 DFT, of a composite window containing the linear combinations, wherein the DFT is represented by a plurality of energy values in each of a respective plurality of frequency bins, andoutputting a representation of the estimated motion,wherein the linear combination is an addition of the first sample to the respective 15 second sample where m is an odd integer, and a difference between the first sample and the respective second sample where m is an even integer.

15. A computer program which, when executed by one or more processors, is arranged to cause the method of claim 14 to be performed by an apparatus comprising:>0 the one or more processors;a speaker to transmit the periodic acoustic pilot signal;a microphone to receive a signal reflected by a target ensonified by the pilot signal;a storage means to store the plurality of consecutive samples and the composite window; and25 an output means configured to output the representation of the estimated motion.

Citation Information

Patent Citations

  • Detecting physiological movement from audio and multimodal signals

    EP3515290B1

  • Motion tracking using pure tones

    GB2609061A

  • Motion tracking using pure tones

    WO2022185025A1