Physical change detection based on active acoustic wave sensing related application data

EP4646139A4Pending Publication Date: 2026-03-11HUAWEI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-04-24
Publication Date
2026-03-11

Smart Images

  • Figure CN2023090305_31102024_PF_FP_ABST
    Figure CN2023090305_31102024_PF_FP_ABST
Patent Text Reader

Abstract

Methods, devices, and processor-readable media for detecting a physical change within an environment(100), including acquiring a first set of sound frames representing sound waves reflected in the environment(100), applying frequency filtering to extract, from the first set of sound frames, a corresponding first set of filtered frames limited to a predefined frequency band that corresponds to a frequency band of a predefined cyclic acoustic probing signal, computing instantaneous amplitudes for the filtered frames in the first set of filtered frames, evaluating, based on the computed instantaneous amplitudes and predefined reference parameters that correspond to a reference state of the environment(100), whether the first set of sound frames indicate a physical change within the environment(100), and causing a predefined action to be performed when the evaluating indicates the physical change.
Need to check novelty before this filing date? Find Prior Art

Description

PHYSICAL CHANGE DETECTION BASED ON ACTIVE ACOUSTIC WAVE SENSING RELATED APPLICATION DATATECHNICAL FIELD

[0001] The present application generally relates to methods, systems and computer media related to physical change detection based on active acoustic wave sensing.BACKGROUND

[0002] Presence detection and activity recognition typically involves identifying and classifying movements and activities from data that is gathered using various sensors. Desirable features of a presence detection / activity recognition system can include: (1) ubiquity, meaning that the system can be implemented on a wide range of electronic devices, (2) contactless, meaning that the system does not require the use of sensors that are attached to the objects of interest, (3) efficiency, meaning that the system efficiently consumes computational / memory resources, while being cost efficient, (4) accuracy, meaning that the system provides a high detection accuracy, (5) flexibility, meaning that the system can easily adapt to different detection ranges and environments, (6) robustness, meaning that the system can mitigate interference, and (7) privacy, meaning that the system does not introduce privacy concerns.

[0003] Considering the prevalence of electronic devices within people’s personal spaces such as homes and vehicles, presence detection is becoming increasingly common in smart homes and smart vehicles. Furthermore, hand, gesture and motion type recognition (e.g., activity recognition) has also gained a lot of interest in smart home applications including energy efficient smart homes, fitness tracking, and health care of elderly / sick / impaired people that can among other things, include fall detection and respiration monitoring. Furthermore, hand and body gestures can be considered for in-vehicle and smart home interactions.

[0004] Increasing demand for presence detection / activity recognition systems has led to several proposed solutions for such systems in recent years. Proposed solutions have included systems based on different types of sensors, including for example inertial sensors, radio  frequency (RF) sensors, acoustic sensors, infrared (IR) , electromagnetic, camera, pressure and biological sensors.

[0005] In the case of inertial sensors, systems have employed electronic devices equipped with inertial measuring units (IMUs) to perform human activity detection and occupancy detection. Although IMU-based solutions are generally cost effective and less resource demanding, they can be invasive in the sense that the sensor is to be attached to the object of interest (e.g., a human) and therefore provide a solution that is not contactless. Infrared (IR) and electric field sensor based solutions have also been used to detect occupancy. However, IR sensors need to be deployed in Line-of-Sight (LoS) of the targets to be detected and both IR and electric field sensors have a limited range. Furthermore, such sensors are often not present in electronic devices that are commercial-off-the-shelf (COTS) devices, and thus fail to meet the ubiquity feature noted above. Camera based solutions are commonly used for human activity detection and occupancy detection. However, camera sensors only work in LoS mode with proper lighting conditions, and more importantly, using a camera sensor can also give rise to privacy concerns.

[0006] Pressure sensor based solutions for detecting occupancy in indoor environments are typically not able to classify activity types, and also require specialized hardware to be installed in or provided over flooring. CO2 detector based solutions have been used to perform occupancy detection, which shares the same drawbacks with the pressure sensor based approach, requiring the dedicated sensors to be deployed in the environment to be monitored, and not being suitable for activity detection. Biological sensor based systems such as those used to detect electromyogram (EMG) and photoplythesmogram (PPG) signals have been applied for human activity recognition. However, such approaches are invasive (e.g., not contactless) , need dedicated hardware and are not suitable for presence detection applications.

[0007] As any movement in a static (non-changing) environment changes the reflection properties of the environment, these changes can be measured using either electromagnetic (EM) or mechanical (acoustic) waves. Currently, EM wave-based transmit and receive technologies have gained a lot of interest for presence / activity detection purposes. WiFi-based Channel State Information (CSI) can be used to perform human activity detection (such as fall recognition) within indoor environments in which a WiFi network is available. WiFi  solutions rely on continuously gathering CSI from an indoor environment in order to detect any change in the reflection properties. More clearly, CSI indicates how EM signals propagate through a specific channel within a certain carrier frequency. Thus, CSI can be considered as the channel frequency response between a transmitter and receiver, which characterizes the multipath properties of the channel. In order to calculate CSI, a WiFi transmitter sends training sequence as preamble. At the receiver side, a CSI matrix is estimated based on the original training sequence and its reflection through the channel. This procedure is repeated frequently in time in order to have images (CSI matrix) of the environment versus time, which are then de-noised and fed to a principal component analysis to extract relevant features. Finally, a Support Vector Machine (SVM) classifier is trained using the extracted CSI and then used to detect specific activities in the environment (fall detection) .

[0008] Although EM wave-based techniques are contactless, they can raise privacy concerns as collected data can be easily leaked into a network. Furthermore, EM signals can easily penetrate through walls which makes it difficult to distinguish and discard CSI resulting from the activities within neighboring units / rooms, which means that the range of the detection is not controllable. Bluetooth (BT) technology can also be used for EM wave-based occupancy detection, however BT is not present in many COTS devices and can have a relatively small range. Furthermore, the narrow bandwidth of WiFi and BT technologies leads to less time resolutions.

[0009] Acoustic-wave based solutions use sound waves to measure changing reflectance properties in a static environment. In this sense, either CSI measurements or Doppler shift based features can be used to detect presence or recognize activity in the environment. Acoustic waves have a much lower velocity than EM waves and are thus suitable for implementation on limited-resource commodity-type COTS devices. Further, microphones and speakers are ubiquitous in many COTS devices. Passive acoustic measurements thorough COTS microphones are described for presence detection in: Khan, M.A.A.H., Hossain, H.S., &Roy, N. (2015, June) . Sensepresence: Infrastructure-less occupancy detection for opportunistic sensing applications. In 2015 16th IEEE International Conference on Mobile Data Management (Vol. 2, pp. 56-61) . IEEE. However, continuous recording of indoor acoustic data raises privacy issues for most of the users. Surface acoustic waves are exploited  to detect near device gestures and taps in: Gong, J., Gupta, A., &Benko, H. (2020, October) . Acustico: surface tap detection and localization using wrist-based acoustic TDOA sensing. In Proceedings of the 33rd Annual ACM Symposium on User Interface Software and Technology (pp. 406-419) . However, such a solution is not suitable for presence detection or remote gesture recognition.

[0010] Further, due to limited bandwidth of microphones and speakers (especially in the inaudible ultrasonic range) , acoustic-wave based solutions face some difficulties in detecting and tracking fine-grained and fast movements. Solutions that attempt to address these shortcomings include AURES [Shih, O., Lazik, P., &Rowe, A. (2016, November) . Aures: A wide-band ultrasonic occupancy sensing platform. In Proceedings of the 3rd ACM international conference on systems for energy-efficient built environments (pp. 157-166) ] and UltraSense [Hammoud, A., Deriaz, M., &Konstantas, D. (2017) . Ultrasense: A self-calibrating ultrasound-based room occupancy sensing system. Procedia Computer Science, 109, 75-83] .

[0011] AURES uses dedicated hardware which works at a sampling rate of 196 KHz, and plays Ultrasonic tones of duration 300 msec in order to perform presence detection. Three classifiers are trained based on Doppler-shift features using fast Fourier transform, which makes the solution a supervised presence detection method (needs labeled data for training) . The 196 KHz sampling rate is too high for most COTS devices (typical sampling rates of acoustic modules in COTS devices are 44.1 and 48 kHz) . On the other hand, UltraSense uses long tone segments of duration 3 seconds in ultrasonic frequency range in order to compensate for a smaller sampling rate of 44.1 kHz. Doppler-shift based feature extraction using fast Fourier transform is used to detect occupancy in the environment, but UltraSence uses an inference stage that is not supervised (as no labeled data is used to train a classifier) . Instead it uses a self-calibration technique, in which certain number of frames are collected to form a training set and then used to get the still frequency spectrum and set a threshold value, according to the given room environment.

[0012] Finally, the system uses unsupervised learning to classify the frames into still and motion frames. The time resolution of UltraSense is too low (as it uses recording chucks of length 3 seconds) , therefore, it is not capable of detecting short-time / fine-grained movements. Further, neither of AURES and UltraSense is capable of activity recognition.

[0013] Accordingly, known presence detection / activity recognition solutions each lack one or more of the desirable features noted above. Among other things, limitations of existing systems can include one or more of: failure to provide no information about the distance / range of the motion / activity, therefore; inability to provide controllable effective range / coverage; requirement for dedicated hardware and not suitable to be implemented on COTS devices; requirement that a signal transmitter and receiver be on the same device (synchronized detection) ; acoustic based solutions are either for presence detection or activity recognition and do not provide a unified solution; discontinuous detection, which is not suitable for activity detection and tracking (fine-grained and fast movements are not detectable) ; and high detection latency which makes it unsuitable for detection and tracking of rapid events.

[0014] The concept of a “Smart Home” that includes several computer-enabled smart devices connected to a wireless home network has been popular for several years. Examples of common smart devices that can be included in a smart home environment include smart TVs, virtual assistant-enabled smart speakers, smart lights, smart electrical switches and smart appliances such as smart fridges, which support audio notifications and voice commands.

[0015] Accordingly, there is a need for methods and systems that can address at least some of the shortcomings noted above.

[0016] SUMMARY

[0017] According to an example aspect, a computer implemented method for detecting a physical change within an environment is disclosed. The method includes: acquiring a first set of sound frames representing sound waves reflected in the environment; applying frequency filtering to extract, from the first set of sound frames, a corresponding first set of filtered frames limited to a predefined frequency band that corresponds to a frequency band of a predefined cyclic acoustic probing signal; computing instantaneous amplitudes for the filtered frames in the first set of filtered frames; evaluating, based on the computed instantaneous amplitudes and predefined reference parameters that correspond to a reference state of the environment, whether the first set of sound frames indicate a physical change within the environment; and causing a predefined action to be performed when the evaluating indicates the physical change.

[0018] In some examples of the preceding aspect, the predefined cyclic acoustic probing signal comprises a sequence of sound segments that each comprise a first duration of a constant frequency sound tone and a second duration of varying frequency sound chirp.

[0019] In some examples of one or more of the preceding aspects, the method comprises, contemporaneously with acquiring the first set of sound frames, outputting via a speaker, the predefined cyclic acoustic probing signal into the environment.

[0020] In some examples of one or more of the preceding aspects, for each sound segment, the constant frequency sound tone transitions with a continuous phase to the varying frequency sound chirp, and the varying frequency sound chirp linearly changes in frequency until an end of the second duration.

[0021] In some examples of one or more of the preceding aspects, for each sound segment, the constant frequency sound tone is within a range of approximately 17.5KHz to 23.5KHz and the varying frequency sound chirp is within a range of approximately 17KHz to 23.5KHz.

[0022] In some examples of one or more of the preceding aspects, for each sound segment, the first duration has a range of approximately 0.5ms to 50ms and the second duration has a range of approximately 10ms to 500ms.

[0023] In some examples of one or more of the preceding aspects, each sound segment has a total cycle duration of approximately 70ms, and for each sound segment the first duration is approximately 20ms, the second duration is approximately 50ms, the constant frequency sound tone has a frequency of approximately 17.8KHz and the varying frequency sound chirp varies linearly from approximately 17.8KHz to approximately 21KHz during the second duration.

[0024] In some examples of one or more of the preceding aspects, applying frequency filtering comprises applying a filter that excludes sound frequencies below 17.5KHz.

[0025] In some examples of one or more of the preceding aspects, acquiring the first set of sound frames comprises: digitizing and storing, for a sampling duration, a sound signal representing sound waves as captured by a microphone located in the environment, generating the first set of sound frames from the sound signal such that the sound frames overlap in time, with each sound frame having: (i) a first half overlapping in time with a  second half of a preceding sound frame, and (ii) a second half overlapping in time with a first half of a following sound frame.

[0026] In some examples of one or more of the preceding aspects, the predefined reference parameters include a predefined amplitude threshold and the evaluating comprises: comparing the instantaneous amplitude of each of the filtered frames of the first set of filtered frames to the predefined amplitude threshold to determine a number of outlier frames of the first set of filtered frames having instantaneous amplitudes that fall outside of the predefined amplitude threshold, and inferring that the physical change has occurred within the environment when the number of the outlier frames meets an outlier threshold.

[0027] In some examples of one or more of the preceding aspects, the sampling duration has a duration of at least 0.1 seconds.

[0028] In some examples of one or more of the preceding aspects, the sampling duration is in the range of approximately 0.3 to 0.5 seconds, and the outlier threshold is at least 10%of the total number of the filtered frames corresponding to the sampling duration.

[0029] In some examples of one or more of the preceding aspects, the method includes performing a reference state calibration process to establish the predefined reference parameters, the reference state calibration process comprising: acquiring a calibration set of sound frames representing sound waves in the environment for a calibration duration; applying frequency filtering to extract, from the calibration set of sound frames, a corresponding calibration set of filtered frames limited to the predefined frequency band; computing instantaneous amplitudes for the calibration set of filtered frames; computing the predefined reference parameters based on the instantaneous amplitudes computed for calibration set of filtered frames.

[0030] In some examples of one or more of the preceding aspects, the method is performed on a first stationary smart device and the environment is an indoor environment.

[0031] In some examples of one or more of the preceding aspects, the first stationary smart device comprises one of a smart TV, an interactive smart speaker, a smart appliance, and a smart sound system.

[0032] In some examples of one or more of the preceding aspects, the method comprises, at a second smart device that is located in the environment and spaced apart from the first stationary smart device, acquiring, through a respective microphone of the second smart  device, a second set of sound frames from the environment. At one of the first or second smart devices, the following actions are performed: applying frequency filtering to extract, from the second set of sound frames, a corresponding second set of filtered frames limited to the predefined frequency band; computing instantaneous amplitudes for the sound frames of the second set of filtered frames. The evaluating is further based on comparing the instantaneous amplitudes for the second set of filtered frames to the predefined amplitude criteria.

[0033] In some examples of one or more of the preceding aspects, the first set of sound frames acquired by the first stationary smart device correspond to a first spatial volume of the environment and the second set of sound frames acquired by the second stationary smart device correspond to a second spatial volume of the environment, wherein the evaluating comprises: inferring, based on comparing the instantaneous amplitudes for the first set filtered frames to predefined amplitude criteria, whether the first set of sound frames indicates physical change within the first spatial volume of the environment; and inferring, based on comparing the instantaneous amplitudes for the second set of filtered frames to predefined amplitude criteria, whether the second set of sound frames indicates physical change within the second spatial volume of the environment.

[0034] In some examples of one or more of the preceding aspects, causing the predefined action to be performed comprises causing a first predefined action to be performed when physical change is inferred to be within the first spatial volume and a second predefined action to be performed when physical change is inferred to be within the second spatial volume of the environment.

[0035] In some examples of one or more of the preceding aspects, the physical change corresponds to a moving object in the environment.

[0036] In some examples of one or more of the preceding aspects, the predefined action comprises computing, based on the first set of sound frames, a range of the moving object to a reference location.

[0037] In some examples of one or more of the preceding aspects, computing the range of the moving object comprise extracting base-band features from the first set of sound frames.

[0038] In some examples of one or more of the preceding aspects, the predefined action comprises computing, based on the first set of sound frames, a velocity of the predefined  action comprises predicting, based on the first set of sound frames, a type of motion of the moving object.

[0039] In some examples of one or more of the preceding aspects, the repeating motion corresponds to a breathing motion and the physical change corresponds to a cessation of the breathing for a threshold duration.

[0040] In some examples of one or more of the preceding aspects, the method comprises computing instantaneous frequencies for the filtered frames in the first set of filtered frames, wherein the evaluating is also based on the computed instantaneous frequencies.

[0041] According to a further example aspect, a system is disclosed that includes or more processors, and one or more memories storing machine-executable instructions thereon which, when executed by the one or more processors, cause the system to perform the method of any one of the preceding methods.

[0042] According to a further example aspect, a non-transitory processor-readable medium is disclosed having machine-executable instructions stored thereon which, when executed by one or more processors, cause the one or more processors to perform the method of any one of the preceding methods.

[0043] According to a further example aspect, computer program is disclosed that configures a computer system to perform the method of any one of the preceding methods.

[0044] According to a further example aspect, an apparatus is disclosed that is configured to perform the method of any one of the preceding methods. The examples disclosed herein may provide various advantages. The disclosed solutions can, in some scenarios, provide a presence detection system and method that can be implemented using the speakers and microphones of ciommonly available electronic devices without required specialized hardware.Brief Description of the Drawings

[0045] Reference will now be made, by way of example, to the accompanying drawings which show example embodiments of the present application, and in which:

[0046] FIG. 1 is a block diagram illustrating an example of an interior environment to which example embodiments of active acoustic sensing methods and systems of the present disclosure can be applied;

[0047] FIG. 2 is a block diagram of the environment of FIG. 1 in a different physical state than shown in FIG. 1;

[0048] FIG. 3 is a block diagram of a processor system that can be used to configure an electronic device to implement an acoustic sensing system in the environment of FIG. 1, according to example embodiments;

[0049] FIG. 4A is a block diagram of an acoustic sensing process that the electronic device is configured to perform according to a first example embodiment;

[0050] FIG. 4B is a block diagram of a sound wave reflection measuring operation of the acoustic sensing process of FIG. 4A;

[0051] FIG. 5 is a plot illustrating an acoustic probing signal and resulting reflected sound frames;

[0052] FIG. 6 is a block diagram of a feature extraction operation of the sound wave reflection measuring operation of FIG. 4B;

[0053] FIG. 7 illustrates plots of instantaneous amplitude and instantaneous frequency generated for reflected sound waves in the environment of FIG. 1 and FIG. 2, respectively.

[0054] FIG. 8 is a block diagram illustrating a further example of an interior environment to which example embodiments of multiple-microphone active acoustic sensing methods and systems of the present disclosure can be applied;

[0055] FIG. 9 is a block diagram of the environment of FIG. 8 in a different physical state than shown in FIG. 8;

[0056] FIG. 10 is a block diagram of a sound wave reflection measuring operation of the acoustic sensing process of FIG. 4A in a multiple-microphone active acoustic sensing configuration; and

[0057] FIG. 11 is a block diagram of a feature extraction operation of a range and velocity computation action that can be performed as part of the acoustic sensing process of FIG. 4A.

[0058] Similar reference numerals may have been used in different figures to denote similar components.DETAILED DESCRIPTION

[0059] This disclosure describes systems and methods for detecting physical change in an environment by using active acoustic wave sensing. In example implementations, standard electronic devices (e.g., COTS devices) are configured with software that enables such  devices to perform active acoustic wave sensing without requiring additional hardware and without requiring a further device to be attached to an object-of–interest. In at least some examples, acoustic sensing systems and methods for detecting physical change can be used for one or both of presence detection and activity recognition. In at least some examples, once a moving presence is detected, a distance of the motion from the device is also estimated and further actions will be taken in accordance with the distance. For example, if a motion is within or beyond a specific distance, an action will be taken.

[0060] FIGs. 1 and 2 are block diagrams illustrating an example of an active acoustic wave sensing system 90 located within an interior environment 100. In illustrated examples the environment 100 is an enclosed environment that includes an interior region that is defined by a set of barriers 104 that are static relative to the interior region and can be at least partially sound reflecting. By way of example, in some scenarios, environment 100 can be an indoor environment of a home or office or other structure in which the barriers include walls, floors, ceilings, closed windows and closed doors. In some example, environment 100 can be the interior of a vehicle, with barriers 104 including the structural elements that define a cabin or interior space of the vehicle. Further, environment 100 can include a number of objects (not shown) that are typically unmoving within the environment, such as furniture, entertainment devices, decorations and the like. FIG. 1 illustrates the environment 100 in a reference state (for example, a static state) . FIG. 2 illustrates the environment 100 in a physically changed state relative to FIG. 1. The physical change can, for example, be due to the introduction of a moving object-of-interest 124 (hereinafter object 124, which may, for example, be a human) into the environment 100.

[0061] In the example of FIGs. 1 and 2, the environment 100 includes at least one electronic device 102. Electronic device 102 is a processor-enabled device that includes a processor system 110, a speaker 112 for converting an input audio signal into output sound waves that are propagated into the environment 100, and a microphone 114 for capturing sound that is propagating within the environment 100 and converting that sound into an input audio signal. In at least some examples, electronic device 102 is a COTS device that has been provisioned with specialized software instructions that configure the processor system 110 with an acoustic sensing module 116 that enables the electronic device 102 to function as active acoustic wave sensing system 90. For example, electronic device 102 can be, among  other things, a COTS device such as a smart TV, an interactive smart speaker, a smart appliance, or a smart sound system. As will be explained in greater detail below, the acoustic sensing module 116 can, among other things, configure the electronic device 102 to detect changes in acoustic properties of the environment 100 that indicate the presence of a moving object into the environment 100.

[0062] FIG. 3 illustrates an example of a processor system 110 that can be used to implement the electronic device 102. Processor system 110 includes one or more processors 202, such as a central processing unit, a microprocessor, an application-specific integrated circuit (ASIC) , a field-programmable gate array (FPGA) , a dedicated logic circuitry, a tensor processing unit, a neural processing unit, a dedicated artificial intelligence processing unit, or combinations thereof. The one or more processors 202 may collectively be referred to as a “processor device” . The processor system 200 also includes one or more input / output (I / O) interfaces 204, which interfaces with input devices (e.g., microphone 114) and output devices (e.g., speaker 112) .

[0063] The processor system 110 can include one or more network interfaces 206 that may, for example, enable the processor system 110 to communicate with one or more further devices through a network such as a local area wireless network.

[0064] The processor system 110 includes one or more memories 208, which may include a volatile or non-volatile memory (e.g., a flash memory, a random access memory (RAM) , and / or a read-only memory (ROM) ) . The non-transitory memory (ies) 208 may store instructions for execution by the processor (s) 202, such as to carry out examples described in the present disclosure. The memory (ies) 208 may include other software instructions, such as for implementing an operating system and other applications / functions. In the illustrated example, the memory 208 includes specialized software instructions 116I for implementing acoustic sensing module 116.

[0065] In some examples, the processor system 110 may also include one or more electronic storage units (not shown) , such as a solid state drive, a hard disk drive, a magnetic disk drive and / or an optical disk drive. In some examples, one or more data sets and / or modules may be provided by an external memory (e.g., an external drive in wired or wireless communication with the processor system 110) or may be provided by a transitory or non-transitory computer-readable medium. Examples of non-transitory computer readable media  include a RAM, a ROM, an erasable programmable ROM (EPROM) , an electrically erasable programmable ROM (EEPROM) , a flash memory, a CD-ROM, or other portable memory storage. The components of the processor system 200 may communicate with each other via a bus, for example.

[0066] As used here, a “module” can refer to a combination of a hardware processing circuit (e.g. processor 202) and machine-readable instructions (software (e.g., smart device interaction instructions 160I) and / or firmware) executable on the hardware processing circuit.

[0067] FIG. 4A illustrates an example of an acoustic sensing process 400 that can be perfumed by active acoustic wave sensing system 90. The process 400, which is based on detecting changes in sound wave reflections within the environment 100, can in some example implementations be used to detect a presence of a moving object in the environment 100. In the illustrated example, process 400 includes a calibration operation 402 that is used to establish reference sound wave reflection parameters for a reference state of the environment 100 that corresponds to background activity level. As part of calibration operation 402, the active acoustic wave sensing system 90 performs an operation 404 to measure characteristics of sound wave reflections within the environment 100 while the environment 100 is in reference state such as the static state illustrated in FIG. 1. The reference sound wave reflection parameters can represent background activity that is present in the environment 100.

[0068] Sound wave reflection measurement operation 404 is illustrated in greater detail in FIG. 4B. During operation 404, the acoustic sensing module 116 causes the speaker 112 to output a predefined cyclic acoustic probing signal PS (operation 420) , while simultaneously causing a sound signal to be recorded (e.g., sampled and stored) via the microphone 114 (operation 422) to capture sound waves within the environment 100. As indicated in FIG. 1, the probing signal PS generated by the speaker 112 results in output sound waves 120 that result in reflected sound waves 122 that are recorded by processor system 110 via microphone 114.

[0069] An example of predefined cyclic acoustic probing signal PS is shown in FIG. 5. The probing signal PS is configured to allow physical changes within the environment 100 to be identified through modulations of the acoustic probing signal PS that can be detected on reflected sound waves 122. Further, the acoustic probing signal PS is configured so to  minimize interference with normally audible human hearing sounds, but at the same time fall within a range of sound frequencies that can be generated by a speaker of a typical COTS electronic device. Acoustic probing signal PS is further configured to have properties that will allow the effects of normal motions to be modulated onto reflected sound waves 122.

[0070] In this regard, in an illustrated example, the acoustic probing signal PS falls within or close to an ultrasonic range that is at or above an upper end of human audible sounds. In an illustrated example, the acoustic probing signal PS is a sequence of sound segments 502 that each include a first duration (Tt) of a constant frequency sound tone 504 and a second duration (Tc) of varying frequency sound chirp 506. For each sound segment, the constant frequency sound tone 504 transitions with a continuous phase to the varying frequency sound chirp 506, and the varying frequency sound chirp 506 linearly changes in frequency until an end of the second duration (Tc) . In a particular example, each sound segment 502 has a total cycle duration (Ts) of approximately 70ms with the constant frequency sound tone 504 having a duration (Tt) of approximately 20ms and a frequency of approximately 17.8KHz, and the varying frequency sound chirp 506 having a duration of approximately 50ms and a linear frequency variation from approximately 17.8KHz to approximately 21KHz.

[0071] In various examples, different sound frequencies can be used. For example, in the case where the receiving microphone can support a 48KHz sampling rate, the constant frequency sound tone 504 can be a tone that is within a range of approximately 17.5KHz to 23.5Khz and the varying frequency sound chirp 506 can be a chirp having a range of approximately 17.5KHz to 23.5KHz. Higher ultrasonic frequencies could be used if supported by the transmission and receiving devices. Furthermore, different cyclic sound durations could be used in different implementations embodiments. For example, the duration Tt of constant frequency sound tone 504 could be within a range of approximately 0.5ms to 50ms in different examples, and the duration of sound chirp 506 within a range of approximately 10ms to 500ms.

[0072] In example embodiments the probing signal PS is outputted (operation 420) and the resulting environmental sound recorded (operation 422) for a predefined sampling duration (e.g., duration Tsd) . In some examples, the duration Tsd that is applied during  calibration can be in the range of approximately 0.3 to 0.5 seconds. In some examples, time durations of 0.1 seconds or higher could be used for the duration Tsd.

[0073] As indicated in FIG. 4B, a pre-processing operation 424 is applied to the recorded sound signal to provide a pre-processed sound signal x (t) . In an example embodiment, pre-processing operation 424 organizes the recorded sound signal into a set of overlapping sound frames (illustrated as frames F1 to Fn in FIG. 5) that make up the pre-processed sound signal x (t) . For example, the recorded sound signal can be sampled to provide pre-processed sound signal x (t) for the duration Tsd, with the pre-processed sound signal x (t) including the set of sound frames F1 to Fn that overlap in time by 50%, with each sound frame Fi: (i) having a first half overlapping in time with a second half of a preceding sound frame Fi-1, and (ii) a second half overlapping in time with a first half of a following sound frame Fi+1, where i indicates a frame index value. In some examples, the duration (Tf) of each sound frame Fi is selected to be the same as the time duration (Ts) of a sound segment 502 of the probing signal PS. Thus, each sound frame Fi comprises a plurality of sound samples, with the sound samples that form the first half of sound frame Fi also forming the second half of overlapping preceding sound frame Fi-1, and the sound samples that form the second half of sound frame Fi also forming the first half of overlapping subsequent sound frame Fi+1.

[0074] The resulting pre-processed sound signal x (t) (comprising sound frames F1 to Fn) is then processed by a feature extraction operation 426 to extract sound wave reflection characteristics. In the illustrated example, the reference sound wave reflection characteristics includes an instantaneous amplitude |x’a (t) | (also referred to as the envelope) that is extracted from the pre-processed sound signal x (t) . In some examples, the reference sound wave reflection characteristics can also include an instantaneous frequency f’ (t) extracted from the pre-processed sound signal x (t) .

[0075] FIG. 6 illustrates a set of functions that are performed as part of the feature extraction operation 426 according to example implementations. The sound signal x (t) is first subjected to frequency filtering to remove any frequencies below 17.5KHz (Block 602) , resulting in filtered sound signal x’ (t) . The effect of such filtering is to remove captured sound elements represented in the sound signal x (t) that are not caused by reflections of the acoustic probing signal PS in the environment 100.

[0076] An instantaneous amplitude |xa (t) | is then computed for the filtered sound signal x’ (t) . As known in the art, the instantaneous amplitude (or envelope) of a signal x’ (t) can be defined as the magnitude of the complex valued analytic signal xa (t) associated with the sound signal x’ (t) , where the analytic signal xa (t) can be represented by the equation: xa (t) =x’ (t) +jH {x’ (t) } ,

[0077] and where H {x’ (t) } denotes the Hilbert transform of the signal x’ (t) . In the example of FIG. 6, functions 603, 604 and 606 collectively generate the analytic signal xa (t) that is associated with filtered sound signal x’ (t) . Function 605 computes the instantaneous amplitude |xa (t) | of the analytic signal xa (t) . The resulting instantaneous amplitude |xa (t) | (which is comprised of a series of samples) is subjected to a Gaussian smoothing (Block 607) , and the resulting smoothed signal is overlapped and added (Block 612) to output instantaneous amplitude |x’a (t) | that corresponds to the filtered sound signal x’ (t) .

[0078] In some examples, an instantaneous frequency f (t) can also be computed for the filtered sound signal x’ (t) . In this regard, as known in the art, analytic signal xa (t) can also be represented by the equation: xa (t) =|x’ (t) |ejΨ (t)

[0079] where Ψ (t) is the instantaneous phase of analytic signal xa (t) , and the instantaneous frequency f (t) can be computed using the equation: f (t) = (1 / 2Π) (dΨ (t)  / dt) .

[0080] In the example of FIG. 6, functions 608 and 609 collectively generate the instantaneous frequency f (t) of the analytic signal xa (t) that is associated with filtered sound signal x’ (t) . The resulting instantaneous frequency f (t) (which is comprised of a series of samples) is subjected to a Guassian smoothing (Block 610) , and the resulting smoothed signal is overlapped and added (Block 612) to output instantaneous frequency f’ (t) that corresponds to the filtered sound signal x’ (t) .

[0081] By way of illustrated example, the left side 702 of FIG. 7 shows respective plots for the instantaneous amplitude |x’a (t) |and instantaneous frequency f’ (t) characteristics that  have been extracted from the sound signal x (t) that results from the probing signal PS being played in the reference environment 100 of FIG. 1. The instantaneous amplitude |x’a (t) | and instantaneous frequency f’ (t) maintain approximately constant values during the represented sampling duration Tsd as no movement has occurred to modulate reflections of the probing signal PS.

[0082] With reference to FIG. 4A, the instantaneous amplitude |x’a (t) | computed during the sound wave reflection measuring operation 404 of calibration operation 402 is used as, or to generate, one or more reflected sound wave reference parameters that are stored for future use (block 406) . In one example, a mean μ and standard deviation σ of the instantaneous amplitude |x’a (t) | are computed for the set of frames gathered during sampling duration Tsd, and are stored as reference parameters.

[0083] In some examples, parameters of the instantaneous frequency f’ (t) computed during the sound wave reflection measuring operation 404 of calibration operation 402 are also stored as reference parameters for future use. For example, the mean and standard deviation of the instantaneous frequency for the sampling duration Tsd can be computed and stored as reference frequency parameters.

[0084] In example embodiments, the process 400 continually monitors for the occurrence of a trigger event (Decision Block 408) that will trigger the active acoustic wave sensing system 90 to perform an acoustic probing operation 410 in the environment 100. By way of the example, the trigger event includes one or more predefined events including: (i) expiration of a countdown timer; (ii) activation of an input device such as a keyboard or mouse; (iii) detection of an audible sound (e.g., sound of a closing door, a voice) ; and (iv) triggering of a switch (e.g., a door operated switch) . In some examples, in the absence of a trigger event, the process will periodically re-perform calibration operation 402 to update the stored reflected sound wave reference parameters for the environment 100. In this manner, the background reflecting sound wave properties of the environment 100 can be regularly updated pending the occurrence of a trigger event.

[0085] In an illustrative example relating to the environment 100 of FIG. 1 and FIG. 2, a trigger event occurs that coincides with the object 124 (e.g. a human person) that is not in the environment 100 in FIG. 1 moving into the environment 100 as shown in FIG. 2. By way of  example, the trigger event may be that the microphone 114 captures the voice of the person or other audible sound caused by entry of the person.

[0086] Upon detection of the predefined trigger event, the electronic device 102 performs probe operation 410. As illustrated in FIG. 4A, as part of probe operation 410, the electronic device 102 performs the above described sound wave reflection measuring operation 404 to measure sound wave characteristics of reflections within the environment 100. In particular, the acoustic sensing module 116 causes the speaker 112 to output the predefined cyclic acoustic probing signal PS of FIG. 5 (operation 420) for a defined sampling duration (TSD) , while simultaneously causing a sound signal to be recorded (e.g., sampled and stored) via the microphone 114 (operation 422) to capture sound waves within the environment 100. As indicated in FIG. 2, the probing signal PS generated by the speaker 112 results in output sound waves 120 that result in reflected sound waves 122 that are recorded by processor system 110 via microphone 114.

[0087] As indicated in FIG. 4B, pre-processing operation 424 is applied to the recorded sound signal to provide a pre-processed sound signal x (t) . As described above, in the illustrated example, processing operation 424 organizes the recorded sound signal into a set of overlapping sound frames (illustrated as F1 to Fn in FIG. 5) that make up the pre-processed sound signal x (t) .

[0088] The resulting pre-processed sound signal x (t) (comprising the set of sound frames F1 to Fn, where n is a total number of sound frames collected during the sampling duration Tsd, with Fi used herein to denote a generic frame within the set) is then processed by feature extraction operation 426 to generate sound wave reflection characteristics. In the illustrated example, the sound wave reflection characteristics includes an instantaneous amplitude |x’a (t) |(also referred to as the envelope) that is extracted from the pre-processed sound signal x (t) . In some examples, the reference sound wave reflection characteristics can also include an instantaneous frequency f’ (t) extracted from the pre-processed sound signal x (t) . As described above, feature extraction operation 426 applies filtering to extract sound waves that are restricted to the bandwidth of the probing signal PS from the sound signal x (t) , resulting in filtered sound signal x’ (t) . Instantaneous amplitude |xa (t) | is then computed for the filtered sound signal x’ (t) . Gaussian smoothing and overlap and add operations are applied to output instantaneous amplitude |x’a (t) | that corresponds to the filtered sound signal x’ (t) .

[0089] The same process described above in respect of the calibration operation 402 can also be applied in the probe operation 410 to provide an instantaneous frequency f’ (t) that corresponds to the filtered sound signal x’ (t) .

[0090] In some examples, the sampling duration Tsd used for the probing operation 410 could be the same as used for calibration operation 402, for example in range of 0.3 to 0.5 seconds. In some examples, the sampling duration Tsd used for the probing operation 410 could be different than that used for calibration operation 402.

[0091] With reference to FIG. 7, relative to the reflected sound waves 122 of FIG. 1, the newly added presence of the object in 124 in FIG. 2 provides a physical change that results in different properties for reflected sound waves 122. As noted above, the left side 702 of FIG. 7 shows respective plots for the instantaneous amplitude |x’a (t) |and instantaneous frequency f’(t) that have been extracted from the sound signal x (t) that results from the probing signal PS being played in the reference environment 100 of FIG. 1. By way of contrast, the right side 704 of FIG. 7 shows respective plots for the instantaneous amplitude |x’a (t) |and instantaneous frequency f’ (t) that have been extracted from the sound signal x (t) that results from the probing signal PS being played in the reference environment 100 of FIG. 2. The introduction and movement of the object 124 can be detected by comparing the right side 704 plots relative to the reference plots of left side 702.

[0092] With reference to FIG. 4A, as part of the probe operation 410, an evaluation operation 412 is performed in respect of the instantaneous amplitude |x’a (t) | (and in some examples, the instantaneous frequency f’ (t) ) computed during the sound wave reflection measuring operation 404 of probe operation 410 in order to infer is a physical change has or has not occurred in the environment 100.

[0093] In some examples, evaluation operation 412 can apply determinative rules. For example, during the evaluation operation 412, the instantaneous amplitudes |x’a (t) | computed in respect of the sound frames that make up sound signal x (t) are evaluated based on predefined amplitude criteria that corresponds to the reference state of the environment 100. Although different predefined amplitude criteria can be applied in different implementations, in an illustrative example the predefined amplitude criteria is based on a comparison of the instantaneous amplitude |x’a (t) | computed for each of the sound frames Fi collected during a sampling duration Tsd of the probe operation 410 relative to the mean μ and standard  deviation σ of the reference instantaneous amplitude |x’a (t) | that were previously stored as reference parameters during calibration operation 402.

[0094] In one example, for each sound frame Fi in the set of sound frames that make up sound signal x (t) corresponding to probe operation 410, an average instantaneous amplitude is computed for the set of samples that make up the sound frame Fi. The average instantaneous amplitude for each sound frame Fi is then respectively compared to the mean μand standard deviation σ of the reference instantaneous amplitude from calibration operation 402. For example, a predefined amplitude threshold can be set that corresponds to an amplitude range that falls within a threshold deviation σt from the reference mean μ. The average instantaneous amplitude computed for each sound frame Fi is compared to the range of the predefined amplitude threshold to identify the number of conforming sound frames from sound signal x (t) that fall within the threshold and the number of outlier sound frames sound signal x (t) that fall outside the predefined amplitude threshold. The number of outlier sound frames is then compared to an outlier threshold. For example, the outlier threshold can be set at a value of 10% (e.g., n / 10) (or another predefined value) of the total number n of sound frames. If the number of outlier sound frames meets or exceeds the outlier threshold, then an inference is made that a physical change has occurred within the environment 100. If the number of outlier sound frames does not meet the outlier threshold, then an inference is made that a physical change has not occurred within the environment 100.

[0095] In the presently described example corresponding to environment 100 as shown in FIG. 1 and 2, the occurrence of a physical change is associated with an object 124 moving in the environment; e.g., the process 400 corresponds to a detecting a moving presence in environment 100 that was not present during calibration operation 402.

[0096] As indicated by Decision Block 413, when evaluation operation 412 infers or generates a determination that a physical change has not occurred, then the process 400 returns to monitoring for a further trigger event (Decision Block 408) .

[0097] However, in the event that evaluation operation 412 infers or generates a determination that a physical change has occurred, then the process 400 causes a predefined action 414 to be performed in response to the detected physical change. By way of example, the action 414 could be causing a controller (e.g., a smart home control system) to turn on a light, activate a heater or air conditioner, start a sequence of smart device operations  including streaming via a smart speaker, cause an alarm to be sounded or a notification message to be sent to a remote location, etc.

[0098] In some examples, instantaneous frequency can be used in a similar fashion as instantaneous amplitude (envelope) to detect motion. More clearly, whenever there is a motion in the environment, there is a perturbation in both the envelope and frequency. However, in general, instantaneous amplitude envelope is more indicative of a motion than instantaneous frequency as long as the direction of motion is not aligned with the speaker / mic. Accordingly, in some examples, both instantaneous envelop and frequency are used to complementarily while detecting motion (e.g., physical change) . In such examples, reference parameters are determined for both the instantaneous amplitude and instantaneous frequency. In alternative examples, evaluation operation 412 can apply a machine-learning based module that has been trained to output an inference value indicating a probability that a physical change has occurred based inputs that include the predefined reference parameters and the computed instantaneous amplitude acquired by the probe operation 410. In some examples, the instantaneous frequency acquired by probe operation 410 can also be used as an input to the machine learning based model.

[0099] The active acoustic wave sensing system 90 shown in the context of FIGs. 1 and 2 is implemented using a single electronic device 102 and a single microphone. However, in example embodiments, the active acoustic wave sensing system 90 can also be implemented using multiple microphones, at least some of which may be associated with different electronic devices. In example embodiments, multiple microphones can be located in an environment for recording reflected sound waves (represented as dashed lines) corresponding to the probing signal PS generated by speaker 112, and the sound collected by each of the multiple microphones processed in order to detect physical changes within the environment 100. By way of example FIGs. 8 and 9 show an example of a multiple microphone implementation of active acoustic wave sensing system 90 in an environment 800 that includes three microphones 114A, 114B, 114C that can collectively be used to detect a physical change caused by a presence of a moving object 124 in environment 800 of FIG. 9 relative to a background static reference state of the environment 800 in FIG. 8. Although the present example will be described in the context of three microphones, the multiple microphones could include as few as two or more than three in other examples. Each  microphone corresponds to a different respective sound wave recording channel (referred to hereafter as channels A, B and C) . As with environment 100, the environment 800 can in some examples correspond to an interior space in a building, or can correspond to an interior space of a vehicle.

[0100] In some examples, the multiple microphones 114A, 114B, 114C may each be connected to respective microphone inputs of a common electronic device. In some alternative examples, the multiple microphones 114A, 114B, 114C may each be associated with respective electronic devices, for example respective COTS electronic devices 102A, 102B and 102C that are configured to communicate through a common network (e.g., a WiFi network) , are each provisioned with software instructions 116I to implement respective acoustic sensing modules 116, and collectively implement active acoustic wave sensing system 90. In some examples, the microphones 114A, 114B, 114C may be associated with respective spatial zones in the environment 800. For example, in FIGs. 8 and 9, electronic devices 102B and 102C (and their respective microphones 114B, 114C) are pre-associated with a spatial zone 2 of the environment 800 and electronic device 102A (and its respective microphone 114A) is pre-associated with a spatial zone 1 of the environment 800. In the illustrated example, spatial zone 1 is partially separated from zone 2 by barrier 802 such as a wall.

[0101] In the context of a multiple microphone implementation such as illustrated in FIGs. 8 and 9, the acoustic sensing process 400 of FIG. 4A can be performed in a manner similar to that described above in respect of the single microphone solution of FIG. s1 and 2, except that the channels A, B and C can be processed independently during sound wave reflection operation 404, and the resulting sound wave characteristics evaluated in evaluation operation 412, either individually or collaboratively or both individually and collaboratively, to detect physical changes in the environment 800. In this regard, FIG. 10 illustrates an example of a sound wave reflection measurement operation 404’ that can be used in place of the sound wave reflection measurement operation 404 of FIG. 4B when multiple microphones are used to perform acoustic sensing process 400.

[0102] In the illustrated example, wave reflection measurement operation 404’ includes a respective processing channel 430A, 440B, 430C for sound waves captured by each of the respective microphones 114A, 114B, 114C. Each processing channel 430A, 430B, 430C  includes a respective set of recording, pre-processing and feature extraction operations 422, 424 and 426 operations (numeric references 422, 424 and 426 have been appended respectively with A, B and C in FIG. 10 for processing channels 430A, 430B, 430C, respectively) that operate in the same manner as described above in respect of FIG. 4B. In some examples, the performance of the operations of the respective processing channels 430A, 430B, 430C can be performed at the respective electronic devices 102A, 102B, 102C for sounds received by the respective microphones 114A, 114B, 114C of such devices and the resulting channel specific sound characteristics (e.g., instantaneous amplitude and, in some examples, instantaneous frequency) communicated to a single electronic device (e.g., electronic device 102C) for generating reference parameters (operation 406) during calibration operation 402 and for evaluating (operation 412) during probe operation 410. However, in other example embodiments, recorded sound signals from the respective electronic devices 102A, 102B, 102C can be communicated to a common device (e.g. electronic device 102C) that can be configured to perform at least some of the operations of the respective processing channels 430A, 430B, 430C for can be performed at electronic devices 102A, 102B, 102C other than the recording device.

[0103] In an illustrative example, during the sound wave reflection measurement operation 404’ of calibration operation 402, the probing signal PS is played by speaker 112C of electronic device 102C. Resulting channel specific sound waves reflected within the environment 800 are respectively sampled and recorded via each of the microphones 114A, 114B and 114C. The processing channels 430A, 430B and 430C each output respective channel specific reference state characteristics, e.g. instantaneous amplitudes 428A, 428B and 428C (and respective instantaneous frequencies in some examples) respectively for channels A, B and C. These respective reference state characteristics are then processed by reference parameter generation operation 406 to generate channel specific reference parameters. For example, the reference parameters can include: (i) Channel A mean μ and standard deviation σ of the reference instantaneous amplitude |x’a (t) | for sound waves received at microphone 114A; (ii) Channel B mean μ and standard deviation σ of the reference instantaneous amplitude |x’a (t) | for sound waves received at microphone 114B; and (iii) Channel C mean μand standard deviation σ of the reference instantaneous amplitude |x’a (t) | for sound waves received at microphone 114C. In some implementations, the corresponding channel specific  instantaneous frequency parameters can also be computed. In some examples, the reference parameter generation operation 406 for each respective channel can be performed at the respective electronic device 102A, 102B and 102C at which the corresponding sound wave samples were recorded, and the resulting channel specific parameters communicated to a common electronic device (e.g., electronic device 102C) for storage. In other examples, the recorded sound signals from the respective electronic devices 102A, 102B, 102C can be communicated to a common device (e.g. electronic device 102C) that performs the reference parameter generation operation 406 for each of the different channels.

[0104] In example embodiments, the sound wave reflection measurement operation 404’ of probe operation 410 operates in a similar manner as described above in respect of calibration operation 402, with the resulting channel specific sound characteristics (e.g., instantaneous amplitude and, in some examples, instantaneous frequency) communicated to a single electronic device (e.g., electronic device 102C) for evaluating (operation 412) . In some examples, the evaluation results for each channel are used collaboratively, and in some examples the evaluation results for each channel are used independently of each other.

[0105] In one example, evaluation operation 412 is performed to generate a respective channel-specific inference for each of the channels A, B and C. Each channel specific evaluation is performed in the same manner as described above in respect of FIG. s1 and 2, with the evaluation for each channel A, B and C being based on its respective stored channel specific reference parameters (determined during calibration operation 402) and the channel specific sound wave characteristics computed during probe operation 410.

[0106] In one example, as part of a second step of evaluation operation 412, the channel-specific inferences are then collectively evaluated to make physical change determinations in respect of the entire environment, or in some examples, respective zones of the environment. For example, in one scenario, a positive inference of a physical change based on the channel A reflected sound wave characteristics causes evaluation operation 412 to conclude that a moving object is present in Zone 1; positive inference of a physical change for BOTH channel B AND channel C causes evaluation operation 412 to conclude that a moving object is present in Zone 2. In an alternative example, positive inference of a physical change for EITHER of channel B OR channel C causes evaluation operation 412 to conclude that a moving object is present in Zone 2.

[0107] The action 414 that is taken can be dependent on which channel or combination of channels have indicated a physical change. For example, Zone 1 lights turned on if a physical change is inferred in respect channel A.

[0108] The example implementations disclosed above have focused on detection of physical change in the context of presence detection. In at least some example implementations, the detection of a physical change can be used to trigger one or more actions 414 that use the results of the acoustic sensing process to provide additional dynamic properties of the detected object 124, including for example detection or classification of a type of motion of the object, and one or more of a range, velocity and motion direction of the object relative to a reference location.

[0109] In this regard, FIG. 11 shows a block diagram of a range and velocity computation operation 450 that can be performed as part of an action 414. In the example of FIG. 11, the range and velocity computation operation 450 receives as input the pre-processed sound signal x (t) (comprising overlapping frames F1 to Fn) . By way of example, the pre-processed sound signal x (t) can be a signal generated by sound signal pre-processing operation 424 (FIG. 4B) , corresponding to sound waves by recorded via microphone 114 by sound signal recording operation 422 (FIG. 4B) simultaneously with playing of an acoustic probing signal 420 by speaker 112 in environment 100 as part of probe operation 410. In the illustrated example, evaluation operation 412 has determined that the pre-processed sound signal x (t) indicates a physical change corresponding to a moving object in the environment 100. The determination of the physical change triggers the acoustic sensing process 400 to perform range and velocity computation operation 450 to further process the sound signal x (t) to determine a range r and velocity v of the detected moving object relative to the location of the microphone 114.

[0110] In the illustrated example of FIG. 11, the pre-processed sound signal x (t) is subjected to base-band processing. A high pass filter operation 452 is applied to signal x (t) , followed by mixing operation 456 (mixing local oscillator signal 454 with high pass filter output) , and a low pass filter operation 458. The resulting signal is down sampled (operation 460) . Onset detection (operation 426) is performed to detect probing signal chirps 506, followed by mixing (operation 466) with the original base-band probe signal PS 464, resulting in a base-band signal. A low-pass filter is then applied (operation 468) . Operations  462-468 collectively remove the base-band probe signal from the down-sampled signal, leaving a signal that represents motion (i.e., range and velocity) of a detected object of interest. The flowing operations are performed to extract the range and velocity values. Fast-Fourier Transform (FFT) (operation 770) is applied. A beat matrix is generated (operation 472) . In example embodiments, a static cutter removal operation 473 is applied to remove background noise. For example static cutter removal operation 473 can apply a baseband version of noise parameters collected as part of the calibration operation 402. Range bins of the beat matrix are subjected to a further FFT (operation 474) , resulting in a vibration matrix 476. The range r and velocity v of the object 124 relative to microphone 114 can then be extracted from the vibration matrix 476.

[0111] FIG. 11 illustrates a range and velocity computation operation 450 for one sound wave channel corresponding to one microphone 114. Additional processing channels can be replicated for each microphone recorded signal a (t) in a multi-microphone / multi-channel system. For example, in the three channel system of FIGs. 8 and 9, the operations shown in FIG. 11 can be independently performed in respect of all three channels A, B, C, providing independent range and velocity information for object 124 relative to each of the microphones 114A, 114B and 114C. This information can in turn be processed using known triangulation techniques to further resolve data about the location and motion of the object 124.

[0112] The range and velocity data generated by range and velocity computation operation 450 may then be used to determine further possible actions. In at least some examples, delaying the performance of range and velocity computation operation 450 until after a determination of physical change has been inferred ensures that the range and velocity computation operation 450, which can be computationally intense, is not performed unless presence of a moving object is actually confirmed. This can avoid unnecessary use of the computational resources (e.g., memory, processor capability, power consumption) of a COTS electronic device that could have negative performance impact on other operations that the COTS electronic device is responsible for performing.

[0113] Examples of some further use cases of the above disclosed implementations will now be described.

[0114] Multi-Zone Presence Detection System: In a first example, active acoustic wave sensing system 90 is configured to function as a multi-room presence detection system using the process 400 as represented in FIGs. 4A and 10 comprising a plurality of microphones (e.g., microphones 114A, 114B and 114C) located in different zones (e.g., rooms) that are all within a probe signal reflection listening range of a speaker (e.g., speaker 112) . The microphones 114A, 114B and 114C and speaker 112 may each be associated with a single electronic device 102 that is provisioned with software to implement acoustic sensing module 116. Alternatively, at least some of the microphones 114A, 114B and 114C and speaker 112 may each be associated with different electronic devices (e.g., electronic devices 102A, 102B and 102C) that are each provisioned with software to implement acoustic sensing module 116 and are configured to communicate with each other to enable the presently described presence detection system.

[0115] The calibration and probe operations 402, 410 as described above can be performed (using multiple microphone sound wave reflection measuring operation 404’ ) to detect when a moving object becomes present in one of the zones. An action 414 that is specific to that zone can then be taken (e.g., cause a light in the zone to turn on via a smart home system) .

[0116] Auto-Sleep System: In a further use case example, active acoustic wave sensing system 90 is configured to implement an auto-sleep system using calibration and probe operations 402, 410 as described above in the context of single microphone sound wave reflection measuring operation 404. The auto-sleep system can implement an auto-sleep function for an electronic device 102 such as a laptop, personal computer or TV by detecting the absence of a moving object within a defined physical range of the microphone 114 of the electronic device 102, and subsequently turning the device off or putting the device into a power save mode.

[0117] In this use case example, calibration operation 402 is applied to gather and store reference parameters representing a static state environment within a monitoring range of the physical change sensing system (as used here, monitoring range refers to the region of an environment in which the reflection measuring operation 404 can be applied to discern a physical change therein, e.g., microphone is within effective probe signal reflection listening range) . In the case of an auto-sleep system, the system is intended to cause an action 414 to  be taken when no physical change exists between the time at which probe operation 410 is taken and the reference parameters were generated. Thus, the trigger event (decision block 408) that is used to trigger probe operation 410 corresponds to detecting that a user (i.e., object 124) that was previously interacting with or using the electronic device 102 may no longer be interacting with or using the electronic device 102. For example, such a trigger could be caused by the electronic device 102 detecting that a user input device (e.g., a mouse or keyboard or wireless remote control) has not received any inputs for a defined time duration. In such case the electronic device 102 performs probe operation 410 to determine, via evaluation operation 412, whether or not there is a physical change within the monitoring range of the electronic device 102. In the event that a physical change is not inferred (thereby indicating absence of a user within the environment 100) , then the active acoustic wave sensing system 90 determines that the user is not present, and action is taken to automatically turn electronic device 100 off or put it into a power save mode.

[0118] In example embodiments, if evaluation operation 412 infers that a physical change has occurred, then a further action is performed to determine the range (r) of the detected user presence to the microphone 114. For example, range and velocity computation operation 450 can be performed to determine if the detected user presence is within a defined physical range of the microphone 114. If the detected user presence is computed to be within the defined physical range of the microphone 114, a determination is made that the user is present and interacting with the electronic device 102 and no further action is required. However, if the detected user presence is computed to be outside of the defined physical range of the microphone 114, a determination is made that the user is NOT interacting with the electronic device 102 and action is taken to automatically turn electronic device 100 off or put it into a power save mode.

[0119] Accordingly, in such use case a three step decision process can be applied to confirm that an electronic device 102 should be turned off or put into a power saving mode: (1) First, a trigger event is detected (decision block 408) that indicates probable absence of a user (e.g., no keyboard or mouse activity for a defined duration) ; (2) Second, probe operation 410 is performed and captured sound wave characteristics evaluated to determine if a user presence can be detected within a monitoring range of the electronic device 102 (e.g., within effective probe signal reflection listening range) , and (3) Third, when the second step does  indicates a user presence, the reflected sound wave characteristics are further processed by range and velocity computation operation 450 to determine how close (range (r) ) the user presence actually is to microphone 114. If the range (r) is too far in step 3, or no presence is detected in step two, a power saving action is performed, otherwise the electronic device 102 is allowed to continue in its previous state.

[0120] Although the Auto-Sleep System use case has been described above in the context of a single-microphone system, it can also be applied in a multiple channel / multiple-microphone system, with each microphone channel used to control a power saving operation for a respective electronic device 102.

[0121] In-Vehicle Occupancy Detection: In a further use case example, an audio output / microphone input system of a vehicle is configured to implement active acoustic wave sensing system 90 as in-vehicle occupancy detection system. The hazards of mistakenly leaving a child or pet unattended in a vehicle are well known. An in-vehicle occupancy detection system can be used to warn a user when a moving object has been left behind in a vehicle.

[0122] In such an example, calibration operation 402 is applied to gather and store reference parameters representing the vehicle interior in a static state. During post calibration operation, the trigger event (decision block 408) that is used to trigger probe operation 410 corresponds to detecting that the vehicle has been exited by a user and left unattended. For example, a trigger event can be detected in the case where, after a period of use: (i) vehicle doors are locked; (ii) vehicle engine is turned off; and (iii) key fob is not detected within the vehicle. In such case, the probe operation 410 can be activated either once or multiple times for a predetermined period. In the case when evaluation operation 412 infers that a physical change is detected within the vehicle, action 414 can include sending a wireless signal via a cellular network (or other network) to a central monitoring station that in turn relays a warning message to one or more pre-registered electronic contact addresses (e.g., text number or email addresses) . Such a system will enable a user to be warned that a moving object has been left behind (or has entered) a vehicle.

[0123] Contactless Respiration Detection: In a further use case example, active acoustic wave sensing system 90 is configured to detect the respiration patterns of living object. In such an example, sound wave reflection measurement operation 404 (or 404’ ) can be used to  extract features from reflected sound waves recorded by one or more microphones located in the vicinity of a patient’s chest in response to a speaker generated probing signal PS. The resulting computed sound wave characteristics (e.g., instantaneous amplitude and / or instantaneous frequency) can be provided to range and velocity computation operation 450. The output range (r) and velocity (v) can then be analyzed by one or both of a rules based and / or machine learning based model to determine if the computed sound wave characteristics are indicative of an acceptable respiration pattern or not. If not, an alarm can be generated.

[0124] Gesture Recognition: In a further use case example, active acoustic wave sensing system 90 is configured to detect specific user gestures. In such an example, sound wave reflection measurement operation 404 (or 404’ ) can be used to extract features from reflected sound waves recorded by one or more microphones located in the vicinity of a user in response to a speaker generated probing signal PS. The resulting computed sound wave characteristics (e.g., instantaneous amplitude and / or instantaneous frequency) can be provided to range and velocity computation operation 450 to determine a range of the motion from the smartphone. Additionally, the computed sound wave characteristics (e.g., instantaneous amplitude and / or instantaneous frequency) can be provided to a rules based and / or machine learning based model to classify the computed sound wave characteristics as corresponding to a certain type of gestures from a set of candidate gestures. For example, a trigger event (decision block 408) can be triggered by a smartphone generating an audible alarm. Following the trigger event, probe operation 410 is performed using a speaker and microphone of the device. If a user presence is detected (operations 412, 413) , range and velocity computation operation 450 is applied to output a range (r) and velocity (v) . Furthermore, the computed sound wave characteristics are provided to a classification model (e.g., a support vector machine (SVM) ) to predict motion type. In the event that the range (r) and motion type classification collectively indicate that a hand waving gesture has occurred within a defined range of the smartphone, a snooze function action is performed to silence the audible alarm for a defined snooze duration.

[0125] In various embodiments, the active acoustic sensing system 90 can provide one or more of the following features or advantages.

[0126] (1) The use of a continuous sound wave measurement in combination with a continuous ultrasonic probing signal during a probe operation enables physical changes to be continually monitored throughout a sampling duration (Tsd) . This enables the active acoustic sensing system 90 to detect a moving object presence based on features extracted from high pass filter data (e.g., evaluation 412 based on reflected sound wave characteristics provided by extraction operation 426) and also to determine a range and motion of the detected object based on features extracted from base band data (e.g., range and velocity computation operation 450) .

[0127] (2) Evaluation operation 412 can apply a simple and computationally inexpensive inference algorithm to detect physical changes that correspond to presence of a moving object. This enables the presence detection functionality of the active acoustic sensing system 90 to be implemented in electronic devices that have relatively limited computational resources and without requiring a model to be trained using supervised data.

[0128] (3) The sampling duration Tsd and the smoothing factor used for Gaussian smoothing during feature extraction operation 426 impact the sensitivity of the evaluation operation 412. In example embodiments, the sampling duration Tsd and the smoothing factor can be adjusted to provide sensitivity adjustment based on the intended use case.

[0129] (4) The use of base-band features in range and velocity computation operation 450 enables the acoustic sensing system 90 to be tuned to discern motion to within defined ranges in the environment. Furthermore, for physical change events where motion type (e.g., gesture) classification is desired, an SVM model can be trained to infer motion type within the defined range.

[0130] (5) Calibration operation 402 can be executed regularly to update reference parameters, providing adaptive background activity estimation.

[0131] (6) There is no requirement for synchronization between microphones and speakers; the recorded sound from each microphone can be independently processed to infer a physical change for a coverage area of that microphone.

[0132] (7) Audible sounds can be used as trigger events for instigating a probe operation 410. The same microphones used for detecting audible event triggers can then be used to also record ultrasonic (or near ultrasonic) sounds resulting from the probing signal PS.

[0133] In summary, the active acoustic sensing system 90 can be implemented in both single electronic device and multiple electronic device configurations. In a single electronic device configuration architecture, both microphone and speaker are on the same device, while in multiple electronic device configurations, multiple microphones can exist on separate devices, while still relying on a single speaker in the environment. The acoustic sensing system 90 can be calibrated / initialized to acquire reference parameters that represent a background or static state of an environment. An active sensing probe operation by the system 90 is triggered by a trigger event, for example, an opening door or an audible sound in the environment, among other things. Once the probe operation is triggered, a continuous probing signal is played through the system speaker and the system microphone / microphones will start recording at the same time. High pass or band-pass feature extraction can be used to generate sound wave characteristics that can be used to detect a physical change in an environment and base-band feature extraction can be used to generate range and velocity estimates for the physical change. In a normal operating mode, once a trigger event occurs, the probing signal is played continuously and the received signal is recorded and stored for a sampling duration. The sound wave characteristics extracted via high-pass or band-pass feature extraction are used to search for physical change indicative of motion / activity in the environment. In some examples, an action is taken if the physical change is detected (based on comparison with the reference parameters) . In some actions, this action can include processing the base-band features of the recorded sound waves to extracted range and velocity information corresponding to the physical change. In some examples, an action will be taken based on the detected physical change and a range of the physical change.

[0134] General: Although the present disclosure describes methods and processes with steps in a certain order, one or more steps of the methods and processes may be omitted or altered as appropriate. One or more steps may take place in an order other than that in which they are described, as appropriate.

[0135] Although the present disclosure is described, at least in part, in terms of methods, a person of ordinary skill in the art will understand that the present disclosure is also directed to the various components for performing at least some of the aspects and features of the described methods, be it by way of hardware components, software or any combination of the two. Accordingly, the technical solution of the present disclosure may be embodied in the  form of a software product. A suitable software product may be stored in a pre-recorded storage device or other similar non-volatile or non-transitory computer readable medium, including DVDs, CD-ROMs, USB flash disk, a removable hard disk, or other storage media, for example. The software product includes instructions tangibly stored thereon that enable a processing device (e.g., a personal computer, a server, or a network device) to execute examples of the methods disclosed herein.

[0136] The present disclosure may be embodied in other specific forms without departing from the subject matter of the claims. The described example embodiments are to be considered in all respects as being only illustrative and not restrictive. Selected features from one or more of the above-described embodiments may be combined to create alternative embodiments not explicitly described, features suitable for such combinations being understood within the scope of this disclosure.

[0137] All values and sub-ranges within disclosed ranges are also disclosed. Also, although the systems, devices and processes disclosed and shown herein may comprise a specific number of elements / components, the systems, devices and assemblies could be modified to include additional or fewer of such elements / components. For example, although any of the elements / components disclosed may be referenced as being singular, the embodiments disclosed herein could be modified to include a plurality of such elements / components. The subject matter described herein intends to cover and embrace all suitable changes in technology.

[0138] The terms “substantially” and “approximately” as used in this disclosure can mean that the recited characteristic, parameter, or value need not be achieved exactly, but that deviations or variations including for example, tolerances, measurement error measurement accuracy limitations and other factors known to those skilled in the art, may occur in amounts that do not preclude the effect the characteristic was intended to provide. By way of illustration, in some examples, the terms “substantially” and “approximately” , can mean a range of within 5%of the stated characteristic.

[0139] As used herein, statements that a second item is “based on” a first item can mean that properties of the second item are affected or determined at least in part by properties of the first item. The first item can be considered an input to an operation or calculation, or a  series of operations or calculations that produces the second item as an output that is not independent from the first item.

Claims

1.A computer implemented method for detecting a physical change within an environment:acquiring a first set of sound frames representing sound waves reflected in the environment;applying frequency filtering to extract, from the first set of sound frames, a corresponding first set of filtered frames limited to a predefined frequency band that corresponds to a frequency band of a predefined cyclic acoustic probing signal;computing instantaneous amplitudes for the filtered frames in the first set of filtered frames;evaluating, based on the computed instantaneous amplitudes and predefined reference parameters that correspond to a reference state of the environment, whether the first set of sound frames indicate a physical change within the environment; andcausing a predefined action to be performed when the evaluating indicates the physical change.2.The method of claim 1 wherein the predefined cyclic acoustic probing signal comprises a sequence of sound segments that each comprise a first duration of a constant frequency sound tone and a second duration of varying frequency sound chirp.3.The method of claim 2 comprising, contemporaneously with acquiring the first set of sound frames, outputting via a speaker, the predefined cyclic acoustic probing signal into the environment.4.The method of claim 3 wherein, for each sound segment, the constant frequency sound tone transitions with a continuous phase to the varying frequency sound chirp, and the varying frequency sound chirp linearly changes in frequency until an end of the second duration.5.The method of claim 4 wherein, for each sound segment, the constant frequency sound tone is within a range of approximately 17.5KHz to 23.5KHz and the varying frequency sound chirp is within a range of approximately 17KHz to 23.5KHz.6.The method of claim 4 or 5 wherein, for each sound segment, the first duration has a range of approximately 0.5ms to 50ms and the second duration has a range of approximately 10ms to 500ms.7.The method of claim 4 wherein each sound segment has a total cycle duration of approximately 70ms, and for each sound segment the first duration is approximately 20ms, the second duration is approximately 50ms, the constant frequency sound tone has a frequency of approximately 17.8KHz and the varying frequency sound chirp varies linearly from approximately 17.8KHz to approximately 21KHz during the second duration.8.The method of anyone of claims 1 to 7 wherein applying frequency filtering comprises applying a filter that excludes sound frequencies below 17.5KHz.9.The method of any one of claims 1 to 8 wherein acquiring the first set of sound frames comprises:digitizing and storing, for a sampling duration, a sound signal representing sound waves as captured by a microphone located in the environment,generating the first set of sound frames from the sound signal such that the sound frames overlap in time, with each sound frame having: (i) a first half overlapping in time with a second half of a preceding sound frame, and (ii) a second half overlapping in time with a first half of a following sound frame.10.The method of claim 9, wherein the predefined reference parameters include a predefined amplitude threshold and the evaluating comprises:comparing the instantaneous amplitude of each of the filtered frames of the first set of filtered frames to the predefined amplitude threshold to determine a number of outlier frames of the first set of filtered frames having instantaneous amplitudes that fall outside of the predefined amplitude threshold, andinferring that the physical change has occurred within the environment when the number of the outlier frames meets an outlier threshold.11.The method of claim 10 wherein the sampling duration has a duration of at least 0.1 seconds.12.The method of claim 10 wherein the sampling duration is in the range of approximately 0.3 to 0.5 seconds, and the outlier threshold is at least 10%of the total number of the filtered frames corresponding to the sampling duration.13.The method of any one of claims 1 to 12 comprising performing a reference state calibration process to establish the predefined reference parameters, the reference state calibration process comprising:acquiring a calibration set of sound frames representing sound waves in the environment for a calibration duration;applying frequency filtering to extract, from the calibration set of sound frames, a corresponding calibration set of filtered frames limited to the predefined frequency band;computing instantaneous amplitudes for the calibration set of filtered frames;computing the predefined reference parameters based on the instantaneous amplitudes computed for calibration set of filtered frames.14.The method of any one of claims 1 to 13 wherein the method is performed on a first stationary smart device and the environment is an indoor environment.15.The method of claim 14 wherein the first stationary smart device comprises one of a smart TV, an interactive smart speaker, a smart appliance, and a smart sound system.16.The method of claim 14 or 15, further comprising:at a second smart device that is located in the environment and spaced apart from the first stationary smart device, acquiring, through a respective microphone of the second smart device, a second set of sound frames from the environment;at one of the first or second smart devices:applying frequency filtering to extract, from the second set of sound frames, a corresponding second set of filtered frames limited to the predefined frequency band;computing instantaneous amplitudes for the sound frames of the second set of filtered frames,wherein the evaluating is further based on comparing the instantaneous amplitudes for the second set of filtered frames to the predefined amplitude criteria.17.The method of claim 16 wherein the first set of sound frames acquired by the first stationary smart device correspond to a first spatial volume of the environment and the second set of sound frames acquired by the second stationary smart device correspond to a second spatial volume of the environment, wherein the evaluating comprises:inferring, based on comparing the instantaneous amplitudes for the first set filtered frames to predefined amplitude criteria, whether the first set of sound frames indicates physical change within the first spatial volume of the environment; andinferring, based on comparing the instantaneous amplitudes for the second set of filtered frames to predefined amplitude criteria, whether the second set of sound frames indicates physical change within the second spatial volume of the environment.18.The method of claim 17 wherein causing the predefined action to be performed comprises causing a first predefined action to be performed when physical change is inferred to be within the first spatial volume and a second predefined action to be performed when physical change is inferred to be within the second spatial volume of the environment.19.The method of any one of claims 1 to 18 wherein the physical change corresponds to a moving object in the environment.20.The method of claim 19 wherein the predefined action comprises computing, based on the first set of sound frames, a range of the moving object to a reference location.21.The method of claim 20 wherein computing the range of the moving object comprise extracting base-band features from the first set of sound frames.22.The method of any one of claims 19, 20 or 21 wherein the predefined action comprises computing, based on the first set of sound frames, a velocity of the moving object.23.The method of any one of claims 19 to 22 wherein the predefined action comprises predicting, based on the first set of sound frames, a type of motion of the moving object.24.The method of claim 23 wherein the repeating motion corresponds to a breathing motion and the physical change corresponds to a cessation of the breathing for a threshold duration.25.The method of any one of claims 1 to 24, comprising:computing instantaneous frequencies for the filtered frames in the first set of filtered frames,wherein the evaluating is also based on the computed instantaneous frequencies.26.A system comprising:one or more processors; andone or more memories storing machine-executable instructions thereon which, when executed by the one or more processors, cause the system to perform the method of any one of claims 1 to 25.27.A non-transitory processor-readable medium having machine-executable instructions stored thereon which, when executed by one or more processors, cause the one or more processors to perform the method of any one of claims 1 to 25.28.A computer program that configures a computer system to perform the method of any one of claims 1 to 25.29.An apparatus that is configured to perform the method of any one of claims 1 to 25.

Citation Information

Patent Citations

  • Gesture recognition method, electronic equipment and storage medium

    CN115494935A

  • A method for generating output data

    EP1898786B1

  • Apparatus, system, and method for detecting physiological movement from audio and multimodal signals

    WO2018050913A1